Atlas

Benchmarks

← All benchmarks

RiemannBench

Math · 2026-04-08

RiemannBench evaluates unconstrained AI research agents on 25 research-level math problems with full access to coding tools, search, and open-ended reasoning. Scores are reported as task success rates, so higher is better.

Top models (higher is better)

ModelScore
Claude Opus 579.0
GPT-5.6 Sol74.4
Claude Fable 560.0
GPT-5.555.2
Claude Mythos 555.0
Claude Mythos Preview43.0
GPT-5.441.6
Grok 4.538.4
GPT-5.237.6
Kimi K337.6
Gemini 3.5 Flash36.8
Opus 4.834.0
Gemini 3.1 Pro Preview33.6
Inkling15.2
Qwen3.7-Max15.2
Kimi K2.512.0
DeepSeek-V4-Flash10.4
GLM-5.210.4
DeepSeek-V3.28.0
Kimi K2.67.2
DeepSeek-V4-Pro5.6
Loading Atlas data…