Atlas

Benchmarks

← All benchmarks

MATH Level 5 (Open LLM Leaderboard v2)

Math · 2024-06-26

The hardest tier of the MATH competition set, scored by exact match on the final answer with no partial credit and no guessing floor. Reported scores land on the 1,324-problem lattice, confirming a single-run proportion. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.

Top models (higher is better)

ModelScore
Qwen2.5 Instruct 32B62.5
DeepSeek R1 Distill Qwen 14B57.0
Qwen2.5 14B Instruct54.8
Qwen2.5 7B Instruct50.0
Mistral Large 2.1 (Instruct 2411)49.5
Qwen2.5 Coder 32B Instruct49.5
Llama 3.3 70B Instruct48.3
Llama 3.1 Tülu 3 70B DPO44.9
QwQ-32B-Preview44.9
Llama 3.1 Nemotron 70B Instruct HF42.7
Qwen2 72B Instruct41.8
Qwen2.5 72B39.1
Llama 3.1 70B Instruct38.1
Qwen2.5-Coder-7B-Instruct37.2
Qwen2.5 3B Instruct36.8
Qwen2.5 32B35.6
Qwen2 VL 72B Instruct34.4
Phi-431.6
Phi-3.5-MoE-instruct31.2
Qwen2-72B31.1
DeepSeek-R1-Distill-Llama-70B30.7
Qwen1.5 32B30.3
Command R7B (Dec 2024)29.9
Yi 1.5 34B Chat27.7
Qwen2 7B Instruct27.6
Gemma 2 27B IT23.9
Qwen1.5 110B Chat23.4
Yi 1.5 9B Chat22.6
Qwen2.5-Coder-14B22.5
Solar Pro Instruct Preview22.1
DeepSeek R1 Distill Llama 8B22.0
Hermes 3 Llama 3.1 70B21.0
Qwen1.5 14B20.2
Qwen2 VL 7B Instruct19.9
Llama-3.1-Tülu-3-8B19.6
Phi-3.5 Mini Instruct19.6
Ministral 8B Instruct 241019.6
Phi-3 Medium 4K Instruct19.6
Qwen1.5 32B Chat19.6
Gemma 2 9B IT19.5
Loading Atlas data…