ExploitBench (CAISI Inspect + AutoNudge)
Agents · 2026-07-17
CAISI's ExploitBench implementation evaluates 41 recent V8 vulnerabilities using Inspect, a ReAct scaffold tuned to the public benchmark, AutoNudge, and a 300-assistant-message budget. Scores report the percentage of the 16 available exploit-capability flags captured across environments.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos Preview | 57.2 |
| GPT-5.5 | 40.7 |
| Opus 4.8 | 38.1 |
| GLM-5.2 | 21.4 |