Atlas

Benchmarks

← All benchmarks

ArXivMath 06/2026

Math

This MathArena track evaluates models on 49 research-level mathematics problems sourced from arXiv papers submitted in June 2026. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
Claude Opus 591.3
GPT-5.6 Sol86.7
Claude Fable 583.7
GPT-5.582.2
Kimi K372.1
Opus 4.868.5
Gemini 3.1 Pro Preview66.0
Grok 4.566.0
Gemini 3.6 Flash57.1
Muse Spark 1.156.5
Gemini 3.5 Flash53.1
GLM-5.246.9
DeepSeek-V4-Flash42.9
Step 3.7 Flash29.9
Qwen3.6 35B A3B26.5
Qwen3.5-2B4.1
Loading Atlas data…