Atlas

Benchmarks

← All benchmarks

AIME 2026

Math · 2026-02-13

This MathArena track evaluates models on the AIME 2026 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
Opus 4.8100.0
GPT-5.5100.0
Claude Fable 599.9
GPT-5.6 Sol99.9
GPT-5.499.2
Gemini 3.1 Pro Preview98.3
GPT-5.298.3
GPT-5.6 Luna97.6
Inkling97.1
Opus 4.696.7
DeepSeek-V4-Pro96.7
Gemini 3.6 Flash96.7
Gemini 3 Flash Preview96.7
GLM-596.7
Kimi K396.7
Step 3.5 Flash96.7
Opus 4.795.8
DeepSeek-V4-Flash95.8
GLM-5.195.8
Kimi K2.595.8
Kimi K2.695.8
Inkling-Small95.5
Inkling-Small Preview95.1
Gemini 3.5 Flash95.0
Step 3.7 Flash95.0
Nemotron 3 Ultra 550B A55B94.2
DeepSeek-V3.294.2
Grok 4.1 Fast94.2
Qwen3.5 397B A17B94.2
MiMo-V2.593.6
Qwen3.5 35B A3B93.3
Qwen3.5-9B92.5
Gemini 3 Pro Preview91.7
Nemotron 3 Super 120B A12B91.7
Qwen3.5-27B91.7
GLM-5.290.0
Qwen3.5-4B89.7
Qwen3 30B A3B Thinking 250788.3
MiniMax M2.787.7
Haiku 4.585.1
Loading Atlas data…