Atlas

Benchmarks

← All benchmarks

MATH Level 5

Math · 2021-03-05

The hardest tier of problems from the MATH dataset, drawn from competitions like the AMC 10, AMC 12, and AIME.

Top models (higher is better)

ModelScore
GPT-598.1
GPT-5 Mini97.8
O4 Mini97.8
o397.8
Sonnet 4.597.7
Qwen3-Max (2025-09-23)97.1
DeepSeek-R1-052896.6
o3-mini96.5
Haiku 4.596.4
Gemini 2.5 Pro Preview 05-0695.9
Gemini 2.5 Pro Preview 03-2595.6
GPT-5 Nano95.2
O194.7
DeepSeek-R193.1
Claude 3.7 Sonnet91.2
Grok 3 Mini90.9
DeepSeek-R1-Distill-Llama-70B89.9
o1-mini89.2
Grok 388.7
GPT-4.1 Mini87.3
DeepSeek R1 Distill Qwen 14B87.1
Opus 485.0
Sonnet 484.4
Gemini 2.0 Pro Experimental 02-0583.5
GPT-4.183.0
Gemini 2.0 Flash 00182.2
o1 Preview81.6
Mistral Medium 381.6
GPT-4.578.6
DeepSeek-V3-032475.5
Gemma 3 27B IT74.0
Llama 4 Maverick Instruct FP873.0
Gemini 1.5 Pro 00270.4
GPT-4.1 Nano70.0
Qwen3-235B-A22B68.9
Qwen2.5-Max67.2
Qwen-Plus (2025-01-25)65.3
Phi-464.9
DeepSeek-V364.9
Grok 2 121263.5
Loading Atlas data…