Atlas

Benchmarks

← All benchmarks

AIME 2025

Math · 2025-02-12

AIME 2025 combines problems from the 2025 American Invitational Mathematics Examination contests. Scores are reported as answer accuracy.

Top models (higher is better)

ModelScore
GPT-5.2100.0
Opus 4.699.8
Gemini 2.5 Deep Think99.2
Step 3.5 Flash98.3
Gemini 3 Flash Preview97.5
GLM-596.7
DeepSeek-V3.2-Speciale95.8
Kimi K2.595.8
Sonnet 4.695.6
Gemini 3 Pro Preview95.0
GPT-595.0
DeepSeek-V3.294.2
GLM-4.593.3
Opus 4.592.8
O4 Mini92.7
Grok 492.5
Kimi K2 Thinking92.5
DeepSeek-V3.2-Exp91.7
GLM 4.691.7
DeepSeek-V3.190.8
Grok 4 Fast90.8
gpt-oss-120b90.0
DeepSeek-R1-052889.2
gpt-oss-20b89.2
o389.2
Gemini 2.5 Pro88.3
GPT-5 Mini87.5
Gemini 2.5 Pro Experimental 03-2586.7
Falcon H1R 7B86.7
o3-mini86.7
GPT-5 Nano85.0
Sonnet 4.584.2
GLM 4.5 Air83.3
K2-Think83.3
Grok 3 Mini81.7
O181.7
Qwen3-235B-A22B80.8
Haiku 4.580.7
Gemini 2.5 Flash Preview 04-1778.0
QED-Nano77.5
Loading Atlas data…