Atlas

Benchmarks

← All benchmarks

ExploitBench (CAISI best of 3)

Code · 2026-09-17

CAISI's 41 V8 exploit-development tasks, each graded on a 16-point capability ladder. The reported percentage uses the best of three attempts per task, with a 300-turn ReAct agent, benchmark MCP tools, planning checklist, compaction and refusal retries.

Top models (higher is better)

ModelScore
GLM-5.361.1
Loading Atlas data…