MATH-500
Math · 2024-11-15
MATH-500 is a 500-problem subset of the MATH competition mathematics dataset used for evaluating multi-step mathematical problem solving. Scores are typically reported as percentage accuracy.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3 Pro Preview | 96.4 |
| Grok 4 | 96.2 |
| GPT-5 | 96.0 |
| Opus 4.1 | 95.4 |
| Gemini 2.5 Pro Experimental 03-25 | 95.2 |
| GPT-5 Mini | 94.8 |
| gpt-oss-120b | 94.8 |
| o3 | 94.6 |
| Qwen3-235B-A22B | 94.6 |
| gpt-oss-20b | 94.2 |
| Grok 3 Mini | 94.2 |
| Kimi K2 Instruct | 94.2 |
| O4 Mini | 94.2 |
| GLM-4.5 | 94.0 |
| Sonnet 4 | 93.8 |
| GPT-5 Nano | 93.8 |
| DeepSeek-R1 | 92.2 |
| Gemini 2.5 Flash Preview 04-17 | 91.8 |
| o3-mini | 91.8 |
| Claude 3.7 Sonnet | 91.6 |
| Llama 3.3 Nemotron Super 49B V1 | 91.4 |
| Opus 4 | 90.4 |
| O1 | 90.4 |
| Grok 3 | 89.8 |
| Gemini 2.0 Flash Experimental | 89.0 |
| MiniMax M2.1 | 89.0 |
| DeepSeek-V3-0324 | 88.6 |
| Gemini 2.0 Flash 001 | 88.0 |
| GPT-4.1 Mini | 88.0 |
| GPT-4.1 | 87.2 |
| Mistral Medium 3 | 87.0 |
| Llama 4 Maverick Instruct | 85.2 |
| Gemini 2.0 Flash Thinking Experimental 01-21 | 84.6 |
| Gemini 1.5 Pro 002 | 82.8 |
| DeepSeek-V3 | 80.4 |
| GPT-4.1 Nano | 80.2 |
| Llama 4 Scout Instruct | 79.2 |
| Gemini 1.5 Flash 002 | 78.8 |
| Grok 2 | 78.4 |
| Command A | 76.2 |