Atlas

Benchmarks

← All benchmarks

ExploitBench Safeguard Blocking (CAISI 10-task set)

Safety · 2026-07-17

This 10-task ExploitBench safeguard diagnostic reports the average number of assistant messages completed before a deployment begins blocking or refusing further tool calls. Runs stop at 300 turns, so a score of 300 means the model never blocked; lower scores indicate earlier safeguard intervention.

Top models (lower is better)

ModelScore
Opus 4.712.0
Opus 4.836.0
GLM-5.2300
Opus 4.6300
GPT-5.5300
Loading Atlas data…