Atlas

Benchmarks

← All benchmarks

Humanity's Last Exam (Text Only)

General QA · 2025-01-24

This Humanity's Last Exam subtrack covers expert-authored questions that use text rather than image inputs. Scores report answer accuracy in percentage points.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview47.3
GPT-5.4 Pro45.3
Muse Spark40.9
Gemini 3 Pro Preview37.7
GPT-5.436.5
Opus 4.636.2
GPT-5 Pro33.3
GPT-5.228.5
Opus 4.526.3
GPT-526.3
GPT-5.124.6
Gemini 2.5 Pro Preview 06-0522.1
o320.6
GPT-5 Mini19.7
O4 Mini18.9
Gemini 2.5 Pro Preview 05-0618.4
Gemini 2.5 Pro Preview 03-2518.4
gpt-oss-120b15.5
Qwen3 235B A22B Thinking 250715.4
Sonnet 4.514.1
DeepSeek-R1-052814.0
o3-mini13.4
DeepSeek-V3.112.9
Gemini 2.5 Flash Preview 04-1712.6
Qwen3-235B-A22B11.8
Opus 4.111.3
Opus 410.8
Gemini 2.5 Flash Preview 05-2010.7
gpt-oss-20b9.7
GLM-4.59.6
GLM 4.5 Air9.4
DeepSeek-R18.5
Gemini 3.1 Flash-Lite Preview8.0
Claude 3.7 Sonnet7.9
O17.8
O1 Pro7.7
Sonnet 47.6
Gemini 2.0 Flash Thinking Experimental 01-216.5
GPT-5.1 Instant6.5
GPT-4.55.8
Loading Atlas data…