Atlas

Benchmarks

← All benchmarks

MATH-500

Math · 2024-11-15

MATH-500 is a 500-problem subset of the MATH competition mathematics dataset used for evaluating multi-step mathematical problem solving. Scores are typically reported as percentage accuracy.

Top models (higher is better)

ModelScore
Gemini 3 Pro Preview96.4
Grok 496.2
GPT-596.0
Opus 4.195.4
Gemini 2.5 Pro Experimental 03-2595.2
GPT-5 Mini94.8
gpt-oss-120b94.8
o394.6
Qwen3-235B-A22B94.6
gpt-oss-20b94.2
Grok 3 Mini94.2
Kimi K2 Instruct94.2
O4 Mini94.2
GLM-4.594.0
Sonnet 493.8
GPT-5 Nano93.8
DeepSeek-R192.2
Gemini 2.5 Flash Preview 04-1791.8
o3-mini91.8
Claude 3.7 Sonnet91.6
Llama 3.3 Nemotron Super 49B V191.4
Opus 490.4
O190.4
Grok 389.8
Gemini 2.0 Flash Experimental89.0
MiniMax M2.189.0
DeepSeek-V3-032488.6
Gemini 2.0 Flash 00188.0
GPT-4.1 Mini88.0
GPT-4.187.2
Mistral Medium 387.0
Llama 4 Maverick Instruct85.2
Gemini 2.0 Flash Thinking Experimental 01-2184.6
Gemini 1.5 Pro 00282.8
DeepSeek-V380.4
GPT-4.1 Nano80.2
Llama 4 Scout Instruct79.2
Gemini 1.5 Flash 00278.8
Grok 278.4
Command A76.2
Loading Atlas data…