MATH-500 (Artificial Analysis source)
Math · 2023-05-31
MATH-500 is a 500-problem subset of the MATH competition math dataset spanning algebra, geometry, intermediate algebra, number theory, precalculus, and probability. It measures exact final-answer accuracy on hard math problems; higher scores mean more problems solved.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5 | 99.4 |
| Grok 3 Mini | 99.2 |
| o3 | 99.2 |
| Sonnet 4 | 99.1 |
| Grok 4 | 99.0 |
| O4 Mini | 98.9 |
| Gemini 2.5 Pro Preview 05-06 | 98.6 |
| o3-mini | 98.5 |
| Qwen3 235B A22B Thinking 2507 | 98.4 |
| Llama 3.3 Nemotron Super 49B V1.5 | 98.3 |
| DeepSeek-R1-0528 | 98.3 |
| Opus 4 | 98.2 |
| Gemini 2.5 Flash | 98.1 |
| Gemini 2.5 Flash Preview 04-17 | 98.1 |
| Gemini 2.5 Pro | 98.0 |
| MiniMax M1 80k | 98.0 |
| Qwen3 235B A22B Instruct 2507 | 98.0 |
| GLM-4.5 | 97.9 |
| EXAONE 4.0 32B | 97.7 |
| Qwen3 30B A3B Thinking 2507 | 97.6 |
| Qwen3 30B A3B Instruct 2507 | 97.5 |
| MiniMax M1 40k | 97.2 |
| Kimi K2 Instruct | 97.1 |
| O1 | 97.0 |
| Gemini 2.5 Flash-Lite | 96.9 |
| Solar Pro 2 | 96.7 |
| DeepSeek-R1 | 96.6 |
| GLM 4.5 Air | 96.5 |
| Magistral Small 1.0 | 96.3 |
| Qwen3 14B | 96.1 |
| Qwen3 32B | 96.1 |
| Qwen3-30B-A3B | 95.9 |
| Llama 3.3 Nemotron Super 49B V1 | 95.9 |
| QwQ-32B | 95.7 |
| Sonar Reasoning Pro | 95.7 |
| R1 1776 | 95.4 |
| Llama 3.1 Nemotron Ultra 253B V1 | 95.2 |
| DeepSeek R1 Distill Qwen 14B | 94.9 |
| Claude 3.7 Sonnet | 94.7 |
| Llama 3.1 Nemotron Nano 4B V1.1 | 94.7 |