Atlas

Benchmarks

← All benchmarks

BrokenArXiv — Overall

Math · 2026-03-13

BrokenArXiv presents plausible but false statements perturbed from recent arXiv papers. A model is rewarded for declining to prove the claim and for explicitly recognising that it is false as written, rather than for producing a correct answer. Responses are graded by an LLM judge, and this aggregate covers the full set of BrokenArXiv competitions.

Top models (higher is better)

ModelScore
GPT-5.563.9
Kimi K351.3
Claude Fable 548.8
Opus 4.833.0
Gemini 3.1 Pro Preview20.1
DeepSeek-V4-Flash20.0
Gemini 3.5 Flash14.9
GLM-5.214.4
Step 3.7 Flash14.4
Gemini 3.6 Flash10.7
Loading Atlas data…