Atlas

Benchmarks

← All benchmarks

Humanity's Last Exam

General QA · 2025-01-24

A set of 2,500 expert-authored questions spanning over 100 academic subjects, designed to test the limits of frontier AI models on problems that require deep, specialized knowledge.

Top models (higher is better)

ModelScore
Claude Opus 564.8
Claude Mythos Preview64.7
Claude Mythos 564.5
Claude Fable 563.9
Opus 4.857.6
Kimi K356.0
GPT-5.6 Sol55.0
GLM-5.254.7
Kimi K2.654.0
Opus 4.653.1
GPT-5.552.2
GPT-5.452.1
Gemini 3.1 Pro Preview51.4
Kimi K2.550.2
Muse Spark 1.149.4
GPT-5.6 Luna48.9
Qwen3.5 397B A17B48.3
DeepSeek-V4-Pro48.2
Inkling-Small47.8
Opus 4.746.9
Inkling-Small Preview46.6
Inkling46.0
DeepSeek-V4-Flash45.1
GPT-5.4 Pro44.3
Sonnet 544.1
Opus 4.543.2
Gemini 3.5 Flash-Lite42.5
Muse Spark40.6
MiniMax M2.740.3
MiMo-V2.540.0
Grok 4.539.3
Nemotron 3 Ultra 550B A55B37.4
Gemini 3 Flash Preview36.6
GPT-5.6 Terra35.9
Gemini 2.5 Deep Think34.8
Sonnet 4.633.2
Grok 4.333.1
GPT-5 Pro31.6
Grok 4.2030.2
GPT-5.229.9
Loading Atlas data…