Atlas

Benchmarks

← All benchmarks

MATH-500 (Artificial Analysis source)

Math · 2023-05-31

MATH-500 is a 500-problem subset of the MATH competition math dataset spanning algebra, geometry, intermediate algebra, number theory, precalculus, and probability. It measures exact final-answer accuracy on hard math problems; higher scores mean more problems solved.

Top models (higher is better)

ModelScore
GPT-599.4
Grok 3 Mini99.2
o399.2
Sonnet 499.1
Grok 499.0
O4 Mini98.9
Gemini 2.5 Pro Preview 05-0698.6
o3-mini98.5
Qwen3 235B A22B Thinking 250798.4
Llama 3.3 Nemotron Super 49B V1.598.3
DeepSeek-R1-052898.3
Opus 498.2
Gemini 2.5 Flash98.1
Gemini 2.5 Flash Preview 04-1798.1
Gemini 2.5 Pro98.0
MiniMax M1 80k98.0
Qwen3 235B A22B Instruct 250798.0
GLM-4.597.9
EXAONE 4.0 32B97.7
Qwen3 30B A3B Thinking 250797.6
Qwen3 30B A3B Instruct 250797.5
MiniMax M1 40k97.2
Kimi K2 Instruct97.1
O197.0
Gemini 2.5 Flash-Lite96.9
Solar Pro 296.7
DeepSeek-R196.6
GLM 4.5 Air96.5
Magistral Small 1.096.3
Qwen3 14B96.1
Qwen3 32B96.1
Qwen3-30B-A3B95.9
Llama 3.3 Nemotron Super 49B V195.9
QwQ-32B95.7
Sonar Reasoning Pro95.7
R1 177695.4
Llama 3.1 Nemotron Ultra 253B V195.2
DeepSeek R1 Distill Qwen 14B94.9
Claude 3.7 Sonnet94.7
Llama 3.1 Nemotron Nano 4B V1.194.7
Loading Atlas data…