Atlas

Benchmarks

← All benchmarks

ExploitBench (CAISI Inspect + AutoNudge)

Agents · 2026-07-17

CAISI's ExploitBench implementation evaluates 41 recent V8 vulnerabilities using Inspect, a ReAct scaffold tuned to the public benchmark, AutoNudge, and a 300-assistant-message budget. Scores report the percentage of the 16 available exploit-capability flags captured across environments.

Top models (higher is better)

ModelScore
Claude Mythos Preview57.2
GPT-5.540.7
Opus 4.838.1
GLM-5.221.4
Loading Atlas data…