BrokenArXiv — Overall
Math · 2026-03-13
BrokenArXiv presents plausible but false statements perturbed from recent arXiv papers. A model is rewarded for declining to prove the claim and for explicitly recognising that it is false as written, rather than for producing a correct answer. Responses are graded by an LLM judge, and this aggregate covers the full set of BrokenArXiv competitions.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 63.9 |
| Kimi K3 | 51.3 |
| Claude Fable 5 | 48.8 |
| Opus 4.8 | 33.0 |
| Gemini 3.1 Pro Preview | 20.1 |
| DeepSeek-V4-Flash | 20.0 |
| Gemini 3.5 Flash | 14.9 |
| GLM-5.2 | 14.4 |
| Step 3.7 Flash | 14.4 |
| Gemini 3.6 Flash | 10.7 |