CAISI AgentDojo - Phishing Emails, Non-Public Injection (2025)
Safety · 2025-09-30
This condition of CAISI's phishing-email AgentDojo task uses the sophisticated non-public injection selected from prior manual red-teaming on U.S. models. It reports hijacking frequency across the same 30 benign Workspace tasks, with five samples per task and failures to encounter the injection excluded.
Top models (lower is better)
| Model | Score |
|---|---|
| GPT-5 | 0.0 |
| Opus 4 | 12.0 |
| gpt-oss-120b | 87.0 |
| DeepSeek-R1-0528 | 89.0 |
| DeepSeek-R1 | 96.0 |
| DeepSeek-V3.1 | 98.0 |