ExploitGym intended-vulnerability exploits with mitigations
Agents · 2026-05-11
Total count of ExploitGym tasks still exploited through the intended vulnerability when the benchmark's standard defenses are enabled. These protections include applicable ASLR, stack canaries, the V8 heap sandbox, and kernel hardening.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol | 136 |
| Claude Mythos Preview | 45.0 |
| GPT-5.5 | 22.0 |
| GPT-5.4 | 3.0 |
| Opus 4.6 | 0.0 |
| Opus 4.7 | 0.0 |
| Gemini 3.1 Pro Preview | 0.0 |
| GLM-5.1 | 0.0 |
| Muse Spark 1.1 | 0.0 |