ExploitBench Safeguard Blocking (CAISI 10-task set)
Safety · 2026-07-17
This 10-task ExploitBench safeguard diagnostic reports the average number of assistant messages completed before a deployment begins blocking or refusing further tool calls. Runs stop at 300 turns, so a score of 300 means the model never blocked; lower scores indicate earlier safeguard intervention.
Top models (lower is better)
| Model | Score |
|---|---|
| Opus 4.7 | 12.0 |
| Opus 4.8 | 36.0 |
| GLM-5.2 | 300 |
| Opus 4.6 | 300 |
| GPT-5.5 | 300 |