Atlas
/
Benchmarks
137 sources · 855 models × 1278 benchmarks · 74,806 scores
☾
☀
Capabilities
Models
Compare
Benchmarks
Coverage
Methodology
← All benchmarks
Humanity's Last Exam Text Only (Preview)
General QA · 2025-01-24
Preview score on the text-only Humanity's Last Exam split.
Top models
(higher is better)
Model
Score
Gemini 2.5 Pro Preview 03-25
18.6
o3-mini
14.0
Claude 3.7 Sonnet
8.6
DeepSeek-R1
8.6
O1 Pro
8.4
O1
8.3
Gemini 2.0 Flash Thinking Experimental 01-21
7.0
Gemini 2.0 Pro Experimental 02-05
6.7
GPT-4.5
6.6
Llama 4 Maverick Instruct
6.2
Llama 3.2 90B Vision Instruct
5.5
DeepSeek-V3-0324
5.2
Gemini 1.5 Pro 002
5.2
Gemini 2.0 Flash 001
4.9
Gemini 2.0 Flash Experimental
4.9
Claude 3.5 Sonnet (Oct 2024)
4.8
Qwen2 VL 72B Instruct
4.7
Nova Micro
4.6
Gemini 2.0 Flash-Lite
4.4
o1-mini
4.0
Opus 3
4.0
Gemini 1.5 Flash 002
3.8
GPT-4o (2024-11-20)
2.6
Loading Atlas data…