USAMO 2025
Math · 2025-03-27
This MathArena track evaluates models on the six proof-based problems of USAMO 2025. Unlike MathArena's final-answer competitions, which are graded automatically, the proofs are graded by human verification, so scores report awarded partial credit rather than the share of problems answered correctly.
Top models (higher is better)
| Model | Score |
|---|---|
| DeepSeek-R1-0528 | 30.1 |
| Gemini 2.5 Pro | 24.4 |
| O4 Mini | 19.1 |
| DeepSeek-R1 | 4.8 |
| Grok 3 | 4.8 |
| Gemini 2.0 Flash Thinking Experimental (unspecified checkpoint) | 4.2 |
| Claude 3.7 Sonnet | 3.6 |
| QwQ-32B | 3.0 |
| O1 Pro | 2.8 |
| o3-mini | 2.1 |