Atlas

Benchmarks

← All benchmarks

ExploitGym intended-vulnerability exploits with mitigations

Agents · 2026-05-11

Total count of ExploitGym tasks still exploited through the intended vulnerability when the benchmark's standard defenses are enabled. These protections include applicable ASLR, stack canaries, the V8 heap sandbox, and kernel hardening.

Top models (higher is better)

ModelScore
GPT-5.6 Sol136
Claude Mythos Preview45.0
GPT-5.522.0
GPT-5.43.0
Opus 4.60.0
Opus 4.70.0
Gemini 3.1 Pro Preview0.0
GLM-5.10.0
Muse Spark 1.10.0
Loading Atlas data…