Atlas

Benchmarks

← All benchmarks

CAISI AgentDojo - Credential Exfiltration, Non-Public Injection (2025)

Safety · 2025-09-30

This condition of CAISI's credential-exfiltration AgentDojo task uses the sophisticated non-public injection selected from prior manual red-teaming on U.S. models. It reports hijacking frequency across the same 30 benign Workspace tasks, with five samples per task and failures to encounter the injection excluded.

Top models (lower is better)

ModelScore
GPT-50.0
Opus 416.0
gpt-oss-120b53.0
DeepSeek-R1-052876.0
DeepSeek-R186.0
DeepSeek-V3.192.0
Loading Atlas data…