Atlas

Benchmarks

← All benchmarks

GPQA (Open LLM Leaderboard v2)

General QA · 2024-06-26

Graduate-level expert-written science questions. The Open LLM Leaderboard v2 figure pools the main, extended, and diamond splits into one 1,192-question set (confirmed by the reported score lattice), which is a broader and easier set than the 198-question diamond split reported elsewhere. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.

Top models (higher is better)

ModelScore
Mistral Large 2.1 (Instruct 2411)43.7
Qwen2.5 32B41.2
Phi-440.6
Qwen2.5 72B40.5
Qwen2-72B39.4
DeepSeek R1 Distill Qwen 14B38.8
Llama 3.1 70B38.8
Qwen2 VL 72B Instruct38.8
Llama 3.1 Tülu 3 70B DPO37.6
Mixtral 8x22B37.6
Gemma 2 27B IT37.5
Mixtral 8x22B Instruct37.3
Qwen2 72B Instruct37.2
Solar Pro Instruct Preview37.1
Yi-34B36.7
Yi 1.5 34B Chat36.5
Aria36.2
Hermes 3 Llama 3.1 70B36.2
Gemma 2 9B IT36.1
Llama 3.1 70B Instruct35.7
Phi-3.5-MoE-instruct35.6
C4AI Command R+ 08-202435.1
Gemma 2 27B35.1
Qwen2.5 Coder 32B Instruct34.9
DBRX Instruct34.1
Qwen1.5 110B Chat34.1
Phi-3.5 Mini Instruct34.0
Qwen2.5 Instruct 32B33.8
Yi 34B Chat33.8
Phi-3 Medium 128K Instruct33.6
Phi-3 Medium 4K Instruct33.6
Yi 1.5 9B Chat33.5
Qwen1.5 32B33.0
Gemma 2 9B32.9
Llama 3.3 70B Instruct32.9
Aya Expanse 32B32.6
Mistral Small Instruct 240932.4
Smaug-72B-v0.132.4
Qwen2.5 14B Instruct32.2
Phi-3-mini-4k-instruct32.0
Loading Atlas data…