Atlas

Benchmarks

← All benchmarks

BrokenArXiv 04/2026

Math · 2026-05-01

This MathArena track evaluates models on the BrokenArXiv 04/2026 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
GPT-5.572.1
Kimi K359.4
Claude Fable 554.1
Opus 4.834.4
Gemini 3.1 Pro Preview25.8
DeepSeek-V4-Flash23.8
DeepSeek-V4-Pro22.1
Gemini 3.5 Flash21.3
Step 3.7 Flash18.9
GLM-5.218.4
Gemini 3.6 Flash16.0
Step 3.5 Flash13.9
GLM-5.18.2
Opus 4.74.1
Loading Atlas data…