Atlas

Benchmarks

← All benchmarks

AIME 2025

Math · 2025-02-12

AIME 2025 combines problems from the 2025 American Invitational Mathematics Examination contests. Scores are reported as answer accuracy.

Top models (higher is better)

ModelScore
GPT-5.2100.0
Opus 4.699.8
Gemini 2.5 Deep Think99.2
GPT-5-Codex98.7
Step 3.5 Flash98.3
Gemini 3 Flash Preview97.5
GLM-596.7
MiMo-V2-Flash96.3
Kimi K2.595.8
Gemini 3 Pro Preview95.7
GPT-5.1-Codex95.7
Sonnet 4.695.6
GLM-4.795.0
KAT-Coder-Pro V194.7
Nova 2 Lite94.3
Opus 4.592.8
O4 Mini92.7
Grok 491.7
GPT-5.1-Codex mini91.7
GPT-591.7
Nemotron 3 Nano 30B A3B91.0
Qwen3 235B A22B Thinking 250791.0
K-EXAONE 236B-A23B90.3
DeepSeek-V3.1-Terminus89.7
Nova 2 Omni (Preview)89.7
Ring-1T89.3
DeepSeek-R1-052889.2
o389.2
Nova 2 Pro (Preview)89.0
Qwen3 VL 235B A22B Thinking88.3
Apriel-1.6-15B-Thinker88.0
INTELLECT-388.0
Gemini 2.5 Pro88.0
Apriel-1.5-15B-Thinker87.5
Gemini 2.5 Pro Experimental 03-2586.7
o3-mini86.7
GLM-4.6V85.3
ERNIE 5.085.0
GPT-5 Mini85.0
Qwen3 VL 32B Thinking84.7
Loading Atlas data…