Project Euler (MathArena IRT estimate)
Math
MathArena Project Euler estimated accuracy includes item-response-theory predictions for unanswered problems. This is a diagnostic estimate, distinct from the measured panel, and does not enter the capability fit.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.6 | 87.7 |
| GPT-5.2 | 81.6 |
| Gemini 3 Pro Preview | 62.4 |
| Gemini 3 Flash Preview | 61.8 |
| GPT-5.1 | 61.6 |
| Kimi K2.5 | 59.9 |
| GPT-5 | 57.0 |
| DeepSeek-V3.2 | 50.3 |
| Kimi K2 Thinking | 50.1 |
| O4 Mini | 48.4 |
| Grok 4 | 46.5 |
| Grok 4 Fast | 46.2 |
| Grok 4.1 Fast | 45.4 |
| Gemini 2.5 Pro | 26.9 |