ExploitBench Leaderboard (Capability coverage)
Agents · 2026-05-13
Capability coverage on ExploitBench v8-bench: for each model and run regime, the official leaderboard unions the 16 mechanically graded exploit-capability flags reached across the published successful episodes for each of 41 V8 environments, then reports the share of the 656 environment-capability cells that were reached. The default all-runs view uses each regime's best available experiment, so seed counts and turn budgets can vary and are recorded on each score observation.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos Preview | 78.0 |
| Claude Mythos 5 | 78.0 |
| GPT-5.5 | 72.0 |
| Claude Opus 5 | 70.0 |
| Opus 4.8 | 40.0 |
| Sonnet 5 | 31.0 |
| Opus 4.7 | 28.0 |
| Sonnet 4.6 | 26.0 |
| Gemini 3.1 Pro Preview | 26.0 |
| GLM-5.1 | 18.0 |
| Kimi K2.6 | 18.0 |
| Haiku 4.5 | 14.0 |
| MiniMax M2.7 | 13.0 |