Atlas

Benchmarks

← All benchmarks

AIME 2025 (Artificial Analysis avg@10)

Math · 2025-02-12

Artificial Analysis' AIME 2025 protocol averages pass@1 over ten independent attempts for each of 30 questions (300 scored attempts). It is separated from single-attempt AIME observations so its uncertainty is not modeled as a 30-item binomial proportion.

Top models (higher is better)

ModelScore
GPT-5.299.0
GPT-5-Codex98.7
Gemini 3 Flash Preview97.0
DeepSeek-V3.2-Speciale96.7
MiMo-V2-Flash96.3
Gemini 3 Pro Preview95.7
GPT-5.1-Codex95.7
GLM-4.795.0
KAT-Coder-Pro V194.7
Kimi K2 Thinking94.7
GPT-594.3
Nova 2 Lite94.3
GPT-5.194.0
gpt-oss-120b93.4
Grok 492.7
DeepSeek-V3.292.0
GPT-5.1-Codex mini91.7
Opus 4.591.3
Nemotron 3 Nano 30B A3B91.0
Qwen3 235B A22B Thinking 250791.0
GPT-5 Mini90.7
O4 Mini90.7
K-EXAONE 236B-A23B90.3
DeepSeek-V3.189.7
DeepSeek-V3.1-Terminus89.7
Grok 4 Fast89.7
Nova 2 Omni (Preview)89.7
gpt-oss-20b89.3
Grok 4.1 Fast89.3
Ring-1T89.3
Nova 2 Pro (Preview)89.0
o388.3
Qwen3 VL 235B A22B Thinking88.3
Apriel-1.6-15B-Thinker88.0
Sonnet 4.588.0
INTELLECT-388.0
DeepSeek-V3.2-Exp87.7
Gemini 2.5 Pro87.7
Apriel-1.5-15B-Thinker87.5
GLM 4.686.0
Loading Atlas data…