SMT 2025
Math · 2025-05-29
This MathArena track evaluates models on the SMT 2025 problem set. Scores report the percentage of problems answered correctly.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.2 | 96.2 |
| Gemini 3 Pro Preview | 93.4 |
| Gemini 3 Flash Preview | 92.9 |
| GPT-5 | 92.0 |
| Step 3.5 Flash | 91.5 |
| GLM-5 | 91.0 |
| GPT-5.1 | 91.0 |
| Kimi K2 Thinking | 91.0 |
| GLM 4.6 | 90.6 |
| Kimi K2.5 | 90.6 |
| DeepSeek-V3.2-Speciale | 89.2 |
| GPT-5 Mini | 89.2 |
| O4 Mini | 88.7 |
| DeepSeek-V3.2 | 87.7 |
| gpt-oss-120b | 87.7 |
| o3 | 87.7 |
| Falcon H1R 7B | 85.8 |
| Grok 4 | 85.8 |
| GPT-5 Nano | 85.4 |
| DeepSeek-V3.2-Exp | 84.9 |
| Gemini 2.5 Pro | 84.9 |
| Grok 4.1 Fast | 84.6 |
| Grok 4 Fast | 84.4 |
| Sonnet 4.5 | 84.0 |
| DeepSeek-V3.1 | 84.0 |
| DeepSeek-R1-0528 | 83.0 |
| gpt-oss-20b | 83.0 |
| GLM-4.5 | 82.1 |
| K2-Think | 79.7 |
| QED-Nano | 79.7 |
| Grok 3 Mini | 78.8 |
| GLM 4.5 Air | 77.4 |
| Qwen3-235B-A22B | 76.9 |
| Gemini 2.5 Flash | 75.5 |
| Qwen3-30B-A3B | 67.9 |
| DeepSeek-R1 | 67.0 |
| DeepSeek-R1-Distill-Llama-70B | 60.9 |
| DeepSeek-R1-Distill-Qwen-32B | 60.4 |
| Claude 3.7 Sonnet | 56.6 |
| DeepSeek R1 Distill Qwen 14B | 54.7 |