Atlas

Benchmarks

← All benchmarks

AIME 2024-2025

Math

AIME 2024-2025 is an aggregate reporting variant combining AIME problems from the 2024 and 2025 competitions. Scores are percentage accuracy under the reporting source's AVG@8 sampling protocol, averaging eight sampled solutions per problem.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview98.1
GPT-5.296.9
Muse Spark96.9
Gemini 3 Pro Preview96.7
GPT-5.496.7
Grok 4.2096.5
Opus 4.796.3
Opus 4.695.6
Gemini 3 Flash Preview95.6
GPT-5.4 Mini95.6
Kimi K2.595.6
Opus 4.595.4
Qwen3.6 Plus (2026-04-02)94.6
GPT-593.4
GLM-4.793.3
GPT-5.193.3
GLM 4.692.7
gpt-oss-120b92.6
Qwen3.5-Flash92.5
Sonnet 4.692.3
GLM-5.191.9
Grok 4.1 Fast91.9
GLM-591.7
GPT-5 Mini91.5
Grok 4 Fast91.3
MiniMax M2.791.0
Grok 490.6
GPT-5.4 Nano88.8
MiniMax M2.588.8
Sonnet 4.588.2
GLM-4.586.7
o3-mini86.5
gpt-oss-20b86.0
Qwen3.5 Plus (2026-02-15)86.0
Gemini 2.5 Pro Experimental 03-2585.8
Kimi K2 Thinking85.4
o385.3
Grok 3 Mini85.0
DeepSeek-V3.284.6
Qwen3-235B-A22B84.0
Loading Atlas data…