Atlas

Benchmarks

← All benchmarks

ArXivMath 02/2026

Math · 2026-03-13

This MathArena track evaluates models on the ArXivMath 02/2026 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
GPT-5.4 Pro75.8
GPT-5.573.4
Gemini 3.1 Pro Preview62.5
Opus 4.860.2
GPT-5.459.4
DeepSeek-V4-Pro51.6
Gemini 3.5 Flash50.8
GLM-5.246.1
DeepSeek-V4-Flash43.8
Kimi K2.643.0
GLM-541.4
Opus 4.640.6
Opus 4.740.6
GLM-5.139.1
GPT-5.237.5
Step 3.5 Flash32.8
Grok 4.1 Fast32.0
Step 3.7 Flash32.0
Qwen3.5-27B31.3
Nemotron 3 Super 120B A12B30.5
Qwen3.5 35B A3B30.5
Qwen3.5-9B27.3
Qwen3 30B A3B Thinking 250722.7
Qwen3.5-4B19.4
Qwen3 4B Thinking 250718.0
QED-Nano14.1
Loading Atlas data…