Atlas

Benchmarks

← All benchmarks

Agents' Last Exam (ALE)

Agents · 2026-06-03

Agents' Last Exam evaluates generalist computer-use agents on 152 long-horizon, economically valuable professional workflows with hidden references and verifiable outcome-based grading. This headline metric is the pass rate on the full task set: the share of tasks receiving a perfect rubric score.

Top models (higher is better)

ModelScore
GPT-5.6 Sol30.6
GPT-5.6 Luna29.6
Kimi K328.3
GPT-5.6 Terra28.0
Opus 4.827.0
GPT-5.526.6
Claude Fable 525.7
GPT-5.420.5
Opus 4.720.4
Composer 2.520.4
GLM-5.220.4
Doubao Seed 2.1 Pro19.5
Gemini 3.1 Pro Preview15.8
DeepSeek-V4-Pro12.4
Qwen3.7-Max11.8
GLM-5.111.5
Kimi K2.69.2
MiMo-V2.58.6
Qwen3.6 Plus (2026-04-02)8.6
Grok 4.36.6
MiniMax M2.75.9
Loading Atlas data…