Atlas

Benchmarks

← All benchmarks

ExploitBench Leaderboard (Capability coverage)

Agents · 2026-05-13

Capability coverage on ExploitBench v8-bench: for each model and run regime, the official leaderboard unions the 16 mechanically graded exploit-capability flags reached across the published successful episodes for each of 41 V8 environments, then reports the share of the 656 environment-capability cells that were reached. The default all-runs view uses each regime's best available experiment, so seed counts and turn budgets can vary and are recorded on each score observation.

Top models (higher is better)

ModelScore
Claude Mythos Preview78.0
Claude Mythos 578.0
GPT-5.572.0
Claude Opus 570.0
Opus 4.840.0
Sonnet 531.0
Opus 4.728.0
Sonnet 4.626.0
Gemini 3.1 Pro Preview26.0
GLM-5.118.0
Kimi K2.618.0
Haiku 4.514.0
MiniMax M2.713.0
Loading Atlas data…