RiemannBench
Math · 2026-04-08
RiemannBench evaluates unconstrained AI research agents on 25 research-level math problems with full access to coding tools, search, and open-ended reasoning. Scores are reported as task success rates, so higher is better.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 79.0 |
| GPT-5.6 Sol | 74.4 |
| Claude Fable 5 | 60.0 |
| GPT-5.5 | 55.2 |
| Claude Mythos 5 | 55.0 |
| Claude Mythos Preview | 43.0 |
| GPT-5.4 | 41.6 |
| Grok 4.5 | 38.4 |
| GPT-5.2 | 37.6 |
| Kimi K3 | 37.6 |
| Gemini 3.5 Flash | 36.8 |
| Opus 4.8 | 34.0 |
| Gemini 3.1 Pro Preview | 33.6 |
| Inkling | 15.2 |
| Qwen3.7-Max | 15.2 |
| Kimi K2.5 | 12.0 |
| DeepSeek-V4-Flash | 10.4 |
| GLM-5.2 | 10.4 |
| DeepSeek-V3.2 | 8.0 |
| Kimi K2.6 | 7.2 |
| DeepSeek-V4-Pro | 5.6 |