Atlas

Benchmarks

← All benchmarks

BrokenArXiv 03/2026

Math · 2026-05-01

This MathArena track evaluates models on the BrokenArXiv 03/2026 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
GPT-5.573.7
GPT-5.436.6
Opus 4.835.7
DeepSeek-V4-Flash19.6
Gemini 3.5 Flash16.1
Kimi K2.615.6
DeepSeek-V4-Pro15.2
Gemini 3.1 Pro Preview13.8
GLM-5.212.1
GLM-512.1
Step 3.7 Flash10.7
Step 3.5 Flash7.1
GLM-5.16.5
Opus 4.65.8
Opus 4.75.8
Loading Atlas data…