Atlas

Benchmarks

← All benchmarks

AIME 2024 (avg@8)

Math · 2024-02-07

AIME 2024 accuracy averaged across eight sampled solutions per problem, separated from single-attempt results.

Top models (higher is better)

ModelScore
Opus 4.599.6
Gemini 3.1 Pro Preview99.2
Gemini 3 Pro Preview99.2
Grok 4.2098.8
Opus 4.697.5
Kimi K2.597.5
GPT-5.297.1
GPT-5.497.1
Opus 4.796.7
Muse Spark96.7
Gemini 3 Flash Preview95.8
GPT-5.4 Mini95.8
Qwen3.6 Plus (2026-04-02)95.4
GLM-4.795.0
GLM 4.694.6
GPT-594.2
Grok 4 Fast94.2
GPT-5.193.8
Sonnet 4.693.3
GLM-593.3
Grok 493.3
Grok 4.1 Fast93.3
Qwen3.5-Flash93.3
gpt-oss-120b93.1
GPT-5 Mini92.9
GLM-5.192.1
MiniMax M2.791.7
Sonnet 4.590.7
MiniMax M2.589.2
o3-mini89.2
Grok 3 Mini88.8
Gemini 3.1 Flash-Lite Preview88.3
Qwen3.5 Plus (2026-02-15)88.3
DeepSeek-V3.287.5
Gemini 2.5 Pro Experimental 03-2587.5
o387.2
GLM-4.587.1
Magistral Medium 1.287.1
gpt-oss-20b86.7
Kimi K2 Thinking85.8
Loading Atlas data…