Atlas

Benchmarks

← All benchmarks

Agents' Last Exam Last-Exam Pass Rate

Agents · 2026-06-03

Perfect-score pass rate on the 38-task Last-Exam subset containing the benchmark's frontier-difficulty workflows.

Top models (higher is better)

ModelScore
Kimi K310.5
Claude Fable 57.9
GPT-5.6 Sol5.3
Opus 4.72.6
Opus 4.82.6
GLM-5.22.6
GPT-5.52.6
GPT-5.6 Luna2.6
GPT-5.6 Terra2.6
Composer 2.50.0
DeepSeek-V4-Pro0.0
Doubao Seed 2.1 Pro0.0
Gemini 3.1 Pro Preview0.0
GLM-5.10.0
GPT-5.40.0
Grok 4.30.0
Kimi K2.60.0
MiMo-V2.50.0
MiniMax M2.70.0
Qwen3.6 Plus (2026-04-02)0.0
Qwen3.7-Max0.0
Loading Atlas data…