Atlas

Benchmarks

← All benchmarks

Agents' Last Exam ALE-CLI Score

Agents · 2026-06-03

Average partial-credit rubric score on the 105-task ALE-CLI Linux-only subset of Agents' Last Exam.

Top models (higher is better)

ModelScore
GPT-5.6 Sol53.6
Kimi K352.8
GPT-5.6 Luna50.7
GPT-5.6 Terra49.8
GPT-5.549.2
Claude Fable 548.5
Opus 4.744.7
Opus 4.844.6
GLM-5.243.4
GPT-5.441.8
Doubao Seed 2.1 Pro41.4
Composer 2.540.8
Gemini 3.1 Pro Preview36.2
Qwen3.7-Max34.0
Sonnet 4.632.0
DeepSeek-V4-Pro30.8
GLM-5.130.8
Qwen3.6 Plus (2026-04-02)28.6
MiMo-V2.526.5
Grok 4.324.3
Kimi K2.621.2
MiniMax M2.714.6
Loading Atlas data…