Atlas

Benchmarks

← All benchmarks

Humanity's Last Exam (Text-Only)

General QA · 2025-01-24

The 2,158-question text-only subset of Humanity's Last Exam. Tool availability is a run-level setting recorded on each score, not a separate benchmark identity.

Top models (higher is better)

ModelScore
Claude Fable 553.3
Claude Opus 552.6
GPT-5.6 Sol47.2
Opus 4.845.7
Muse Spark 1.145.1
Gemini 3.1 Pro Preview44.7
Kimi K344.3
GPT-5.544.3
GPT-5.6 Terra41.8
GPT-5.441.6
Gemini 3.5 Flash41.0
Grok 4.540.3
GLM-5.240.1
Muse Spark39.9
GPT-5.3-Codex39.9
Opus 4.739.6
Sonnet 539.6
Gemini 3.6 Flash38.3
Motif-3-Beta38.2
Qwen3.7-Max38.1
Gemini 3 Pro Preview37.2
GPT-5.6 Luna37.2
MiniMax M337.1
DeepSeek-V4-Flash-073136.8
Opus 4.636.7
Grok Build 0.136.0
Kimi K2.635.9
DeepSeek-V4-Pro35.9
GPT-5.235.4
Grok 4.335.0
Gemini 3 Flash Preview34.7
MiMo-V2.5-Pro33.8
GPT-5.2-Codex33.5
KAT-Coder-Pro V133.4
Qwen3.7-Plus33.4
Kimi K2.7 Code32.8
Nex-N2-Pro32.4
Grok 4.2032.2
DeepSeek-V4-Flash32.1
LongCat-2.032.1
Loading Atlas data…