QED-Nano
LM-Provers · 2026-02-12 · 4B parameters
QED-Nano is a dense 4-billion-parameter theorem-proving checkpoint based on Qwen3-4B-Thinking-2507 and post-trained with supervised fine-tuning and reinforcement learning on Olympiad-level proof problems. Its separately evaluated agent scaffold can scale inference-time computation up to two million tokens per problem, but the released model itself is the checkpoint recorded here.
Benchmark scores
| Benchmark | Score |
|---|---|
| AIME 2025 | 77.5 |
| AIME 2026 | 82.5 |
| Apex (MathArena) | 1.6 |
| Apex Shortlist | 22.3 |
| ArXivMath 01/2026 | 37.0 |
| ArXivMath 02/2026 | 14.1 |
| ArXivMath 12/2025 | 26.5 |
| BRUMO 2025 | 87.5 |
| Capability | 138.5 |
| CMIMC 2025 | 59.4 |
| Final-Answer Comps — Overall | 42.7 |
| HMMT Feb 2025 | 76.7 |
| HMMT Feb 2026 | 64.4 |
| HMMT Nov 2025 | 75.0 |
| SMT 2025 | 79.7 |