Atlas

Benchmarks

← All benchmarks

BrokenArXiv 02/2026

Math · 2026-03-13

This MathArena track evaluates models on the BrokenArXiv 02/2026 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
GPT-5.568.2
GPT-5.437.9
Opus 4.835.5
GPT-5.225.8
Gemini 3.1 Pro Preview18.6
DeepSeek-V4-Flash17.7
DeepSeek-V4-Pro13.3
Step 3.7 Flash12.1
GLM-511.7
Kimi K2.611.7
Step 3.5 Flash11.3
GLM-5.210.5
GLM-5.19.3
Gemini 3.5 Flash6.5
Opus 4.74.0
Opus 4.63.2
Kimi K2.52.8
Loading Atlas data…