Atlas

Benchmarks

← All benchmarks

ExploitBench Leaderboard (Mean capability score, out of 16)

Agents · 2026-05-13

Mean number of ExploitBench's 16 mechanically graded exploit-capability flags reached per successful episode across the 41-environment V8 matrix. Unlike capability coverage, this metric averages episode scores instead of unioning capabilities across repeated seeds.

Top models (higher is better)

ModelScore
Claude Mythos 510.8
Claude Opus 510.1
Claude Mythos Preview10.0
GPT-5.59.8
Opus 4.85.6
Sonnet 54.2
Gemini 3.1 Pro Preview3.7
Opus 4.73.6
Sonnet 4.63.2
Kimi K2.62.6
GLM-5.12.6
Haiku 4.52.1
MiniMax M2.72.1
Loading Atlas data…