Atlas

Benchmarks

← All benchmarks

GPQA Diamond

General QA · 2023-11-20

The hardest GPQA subset, with graduate-level Google-proof questions in biology, physics, and chemistry that require deep domain reasoning. Scores are reported as accuracy.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview96.0
GPT-5.595.5
GPT-5.6 Sol94.9
Opus 4.794.2
Claude Mythos 594.1
GPT-5.5 Pro93.9
Claude Opus 593.7
Kimi K393.4
GPT-5.2 Pro93.2
MiniMax M392.9
Grok 4.392.9
Claude Fable 592.9
Grok 4.592.9
GPT-5.6 Terra92.9
Claude Mythos Preview92.9
GPT-5.492.8
GPT-5.6 Luna92.4
Gemini 3.6 Flash92.4
Interfaze Beta92.4
Gemini 3.5 Flash92.1
Opus 4.892.0
GPT-5.291.9
GLM-5.291.4
Opus 4.691.3
Grok 4.2091.1
Kimi K2.691.1
Gemini 3 Pro Preview90.9
DeepSeek-V4-Flash-073190.8
Sonnet 590.5
DeepSeek-V4-Pro90.5
Muse Spark 1.190.4
Qwen3.7-Plus90.0
Sonnet 4.689.9
Qwen3.7-Max89.9
GPT-5.2-Codex89.9
Gemini 3 Flash Preview89.8
Hy389.7
Grok Build 0.189.5
DeepSeek-V4-Flash89.4
Qwen3.5 397B A17B89.3
Loading Atlas data…