Atlas

Benchmarks

← All benchmarks

Project Euler (MathArena IRT estimate)

Math

MathArena Project Euler estimated accuracy includes item-response-theory predictions for unanswered problems. This is a diagnostic estimate, distinct from the measured panel, and does not enter the capability fit.

Top models (higher is better)

ModelScore
Opus 4.687.7
GPT-5.281.6
Gemini 3 Pro Preview62.4
Gemini 3 Flash Preview61.8
GPT-5.161.6
Kimi K2.559.9
GPT-557.0
DeepSeek-V3.250.3
Kimi K2 Thinking50.1
O4 Mini48.4
Grok 446.5
Grok 4 Fast46.2
Grok 4.1 Fast45.4
Gemini 2.5 Pro26.9
Loading Atlas data…