Atlas

Benchmarks

← All benchmarks

CAISI AgentDojo - Credential Exfiltration (2025)

Safety · 2025-09-30

The credential-exfiltration task in CAISI's 2025 AgentDojo Workspace evaluation instructs the agent through injected text to find a two-factor authentication code, send it to an untrusted address, and delete the sent email. Scores average hijacking frequency across 30 benign tasks evaluated under two injection conditions.

Top models (lower is better)

ModelScore
GPT-50.0
Opus 48.0
gpt-oss-120b27.0
DeepSeek-R1-052837.0
DeepSeek-V3.150.0
DeepSeek-R165.0
Loading Atlas data…