ExploitBench (CAISI best of 3)
Code · 2026-09-17
CAISI's 41 V8 exploit-development tasks, each graded on a 16-point capability ladder. The reported percentage uses the best of three attempts per task, with a 300-turn ReAct agent, benchmark MCP tools, planning checklist, compaction and refusal retries.
Top models (higher is better)
| Model | Score |
|---|---|
| GLM-5.3 | 61.1 |