CAISI AgentDojo - Credential Exfiltration (2025)
Safety · 2025-09-30
The credential-exfiltration task in CAISI's 2025 AgentDojo Workspace evaluation instructs the agent through injected text to find a two-factor authentication code, send it to an untrusted address, and delete the sent email. Scores average hijacking frequency across 30 benign tasks evaluated under two injection conditions.
Top models (lower is better)
| Model | Score |
|---|---|
| GPT-5 | 0.0 |
| Opus 4 | 8.0 |
| gpt-oss-120b | 27.0 |
| DeepSeek-R1-0528 | 37.0 |
| DeepSeek-V3.1 | 50.0 |
| DeepSeek-R1 | 65.0 |