Atlas

Benchmarks

← All benchmarks

Apex (MathArena)

Math · 2025-08-18

This MathArena track evaluates models on the Apex (MathArena) problem set. Scores report the percentage of problems answered correctly. The leaderboard reports run-averaged accuracy (decimal percentages that land off the 12-item lattice), so the reported number is a mean of per-item pass rates rather than a single-run proportion, and a binomial noise model is not licensed.

Top models (higher is better)

ModelScore
Opus 4.881.3
GPT-5.580.2
GPT-5.4 Pro69.8
Kimi K365.6
Gemini 3.1 Pro Preview60.9
GPT-5.454.2
Opus 4.740.6
Opus 4.634.5
Gemini 3.5 Flash32.3
DeepSeek-V4-Pro28.1
DeepSeek-V4-Flash27.1
Gemini 3.6 Flash26.0
Kimi K2.624.0
Gemini 3 Pro Preview23.4
GLM-5.219.8
Gemini 3 Flash Preview15.6
Step 3.7 Flash14.6
GPT-5.213.5
Step 3.5 Flash13.5
GLM-5.111.5
GLM-510.9
DeepSeek-V3.2-Speciale9.4
Kimi K2.58.8
Nemotron 3 Super 120B A12B7.8
Grok 4.1 Fast5.2
Grok 4 Fast5.2
Qwen3 235B A22B Thinking 25075.2
Qwen3.5 35B A3B4.2
DeepSeek-V3.22.1
GPT-52.1
Grok 42.1
Qwen3.5-27B2.1
Qwen3 4B Thinking 25072.1
Sonnet 4.51.6
QED-Nano1.6
DeepSeek-R1-05281.0
GLM-4.51.0
GPT-5 Mini1.0
gpt-oss-120b1.0
DeepSeek-V3.10.5
Loading Atlas data…