Atlas

Benchmarks

← All benchmarks

ExploitGym

Safety · 2026-05-11

ExploitGym evaluates whether an AI agent can turn a supplied real-world vulnerability and proof-of-vulnerability input into unauthorized code execution. The canonical pass@1 score is the percentage of the official 869-task userspace, V8, and Linux-kernel suite exploited through the intended vulnerability with standard mitigations disabled.

Top models (higher is better)

ModelScore
GPT-5.6 Sol33.7
Claude Opus 522.0
Claude Mythos Preview18.1
GPT-5.514.8
GPT-5.47.0
Opus 4.61.8
Opus 4.71.4
Gemini 3.1 Pro Preview1.4
Muse Spark 1.10.8
GLM-5.10.5
Loading Atlas data…