ExploitBench Leaderboard (Mean capability score, out of 16)
Agents · 2026-05-13
Mean number of ExploitBench's 16 mechanically graded exploit-capability flags reached per successful episode across the 41-environment V8 matrix. Unlike capability coverage, this metric averages episode scores instead of unioning capabilities across repeated seeds.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos 5 | 10.8 |
| Claude Opus 5 | 10.1 |
| Claude Mythos Preview | 10.0 |
| GPT-5.5 | 9.8 |
| Opus 4.8 | 5.6 |
| Sonnet 5 | 4.2 |
| Gemini 3.1 Pro Preview | 3.7 |
| Opus 4.7 | 3.6 |
| Sonnet 4.6 | 3.2 |
| Kimi K2.6 | 2.6 |
| GLM-5.1 | 2.6 |
| Haiku 4.5 | 2.1 |
| MiniMax M2.7 | 2.1 |