Atlas

Benchmarks

← All benchmarks

MMLU-Pro (Open LLM Leaderboard v2)

General QA · 2024-06-26

MMLU-Pro extends MMLU to ten answer options with expert review to reduce noise. Reported scores land on the 12,032-question lattice, confirming a single-run proportion over the full set. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.

Top models (higher is better)

ModelScore
Qwen2.5 72B59.7
Qwen2.5 32B58.1
Qwen2-72B57.3
Qwen2 VL 72B Instruct57.2
QwQ-32B-Preview56.8
Qwen2.5 Instruct 32B56.7
Mistral Large 2.1 (Instruct 2411)55.6
Qwen2 72B Instruct54.0
Llama 3.3 70B Instruct53.3
Llama 3.1 70B Instruct53.1
Phi-452.9
Solar Pro Instruct Preview52.7
Llama 3.1 Nemotron 70B Instruct HF49.2
Qwen2.5 14B Instruct49.0
Qwen1.5 110B Chat48.2
DeepSeek-R1-Distill-Llama-70B47.5
Hermes 3 Llama 3.1 70B47.3
Phi-3 Medium 128K Instruct47.1
DeepSeek-R1-Distill-Qwen-32B46.9
Phi-3 Medium 4K Instruct46.8
DeepSeek R1 Distill Qwen 14B46.7
Phi-3.5-MoE-instruct46.6
Llama 3.1 70B46.5
Mixtral 8x22B46.4
Llama 3.1 Tülu 3 70B DPO46.3
Smaug-72B-v0.146.2
Qwen2.5-Coder-14B45.2
Yi 1.5 34B Chat45.2
Phi-3-Small-8K-Instruct45.1
Qwen1.5 32B45.0
Mixtral 8x22B Instruct44.8
Qwen1.5 32B Chat44.6
Gemma 2 27B IT44.5
C4AI Command R+ 08-202444.2
Qwen2.5 Coder 32B Instruct44.1
Yi-34B44.1
Aria44.0
Gemma 2 27B43.7
Qwen2.5 7B Instruct42.9
Aya Expanse 32B41.3
Loading Atlas data…