Atlas

Benchmarks

← All benchmarks

CAISI AgentDojo - Malware Execution, Non-Public Injection (2025)

Safety · 2025-09-30

This condition of CAISI's malware-execution AgentDojo task uses the sophisticated non-public injection selected from prior manual red-teaming on U.S. models. It reports hijacking frequency across the same 30 benign Workspace tasks, with five samples per task and failures to encounter the injection excluded.

Top models (lower is better)

ModelScore
GPT-50.0
Opus 416.0
gpt-oss-120b56.0
DeepSeek-R1-052894.0
DeepSeek-R197.0
DeepSeek-V3.198.0
Loading Atlas data…