BrokenArXiv 08/2026 (Native Harnesses)
Math · 2026-09-15
MathArena's August 2026 false-statement track has 56 questions and native model harnesses with Python and SageMath without internet. Revised grading awards 0–3 points per question, reserving full credit for identifying falsity and capping contradictory repairs at one point. Scores normalize mean points to 0–100. This protocol is distinct from earlier releases; time and cost limits were introduced partway through the published runs.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 81.9 |
| Claude Fable 5.1 | 79.8 |
| Kimi K3 | 61.9 |
| Gemini 3.8 Flash | 27.4 |
| Muse Spark 1.3 | 20.2 |