Atlas

Benchmarks

← All benchmarks

BrokenArXiv 06/2026

Math

This MathArena track evaluates whether models recognize and refuse to prove 54 plausible but false mathematical statements extracted from recent arXiv papers. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
Claude Opus 590.7
GPT-5.569.4
GPT-5.6 Sol67.3
Kimi K351.9
Muse Spark 1.150.6
Claude Fable 547.8
Opus 4.841.7
Grok 4.529.3
Gemini 3.1 Pro Preview17.6
DeepSeek-V4-Flash16.7
Step 3.7 Flash15.7
GLM-5.214.2
Gemini 3.5 Flash12.3
Gemini 3.6 Flash12.0
Qwen3.6 35B A3B3.4
Qwen3.5-2B2.8
Loading Atlas data…