Atlas

Benchmarks

← All benchmarks

Humanity's Last Exam Diamond

General QA

Humanity's Last Exam Diamond is a fixed 1,000-question revision created by Scale AI and the Center for AI Safety, with 500 reasoning and 500 knowledge questions drawn from HLE and its held-out reserve. The leaderboard reports overall accuracy under closed-book evaluation with automatic answer extraction and judging.

Top models (higher is better)

ModelScore
GPT-6 Astra60.6
Claude Opus 5.555.0
Claude Fable 5.151.3
Claude Opus 538.6
Gemini 3.8 Flash34.3
GPT-6 Sol33.8
GPT-5.6 Sol31.2
Muse Spark 1.325.4
Grok 4.723.4
Kimi K322.2
GLM-5.316.4
Loading Atlas data…