Atlas

Benchmarks

← All benchmarks

ArXivMath 01/2026

Math · 2026-02-04

This MathArena track evaluates models on the ArXivMath 01/2026 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
GPT-5.476.1
DeepSeek-V4-Pro73.9
GPT-5.573.9
Opus 4.672.8
Kimi K2.671.7
Gemini 3.1 Pro Preview70.7
GPT-5.267.9
GLM-5.165.2
Kimi K2.562.5
Gemini 3 Pro Preview62.0
Step 3.5 Flash60.3
Gemini 3 Flash Preview58.1
DeepSeek-V3.257.1
Qwen3.5 397B A17B54.9
GLM-554.4
Grok 4.1 Fast53.3
Qwen3.5-27B53.3
Opus 4.752.2
Qwen3.5 35B A3B50.0
Nemotron 3 Super 120B A12B48.9
Qwen3.5-9B44.6
Qwen3 30B A3B Thinking 250739.1
Qwen3.5-4B38.1
QED-Nano37.0
Qwen3 4B Thinking 250723.9
Loading Atlas data…