ExploitGym
Safety · 2026-05-11
ExploitGym evaluates whether an AI agent can turn a supplied real-world vulnerability and proof-of-vulnerability input into unauthorized code execution. The canonical pass@1 score is the percentage of the official 869-task userspace, V8, and Linux-kernel suite exploited through the intended vulnerability with standard mitigations disabled.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol | 33.7 |
| Claude Opus 5 | 22.0 |
| Claude Mythos Preview | 18.1 |
| GPT-5.5 | 14.8 |
| GPT-5.4 | 7.0 |
| Opus 4.6 | 1.8 |
| Opus 4.7 | 1.4 |
| Gemini 3.1 Pro Preview | 1.4 |
| Muse Spark 1.1 | 0.8 |
| GLM-5.1 | 0.5 |