Atlas

Benchmarks

← All benchmarks

ProofBench

Math · 2026-03-27

Automated theorem-proving benchmark for formal mathematical proof tasks. Scores measure the share of proof obligations solved.

Top models (higher is better)

ModelScore
Claude Opus 578.0
Claude Fable 577.0
GPT-5.6 Sol77.0
Aristotle71.0
GPT-5.6 Terra71.0
Kimi K370.0
Opus 4.869.0
Sonnet 566.0
GPT-5.456.0
Opus 4.754.0
GPT-5.6 Luna54.0
Opus 4.650.0
GPT-5.550.0
Sonnet 4.645.0
Muse Spark 1.139.0
Opus 4.536.0
Gemini 3.6 Flash36.0
GLM-5.235.0
Grok 4.530.0
Gemini 3.5 Flash29.0
Gemini 3.1 Pro Preview26.0
Qwen3.7-Max26.0
MiMo-V2.5-Pro24.0
GLM-5.122.2
GPT-5.4 Mini21.0
Gemini 3 Pro Preview20.0
Sonnet 4.519.0
MiniMax M319.0
GPT-518.0
Muse Spark17.0
Kimi K2.616.0
MiMo-V2.516.0
Gemini 3 Flash Preview15.0
GPT-5.215.0
Grok 4.2014.0
Gemini 3.5 Flash-Lite13.0
GPT-5 Nano12.0
Grok 4.311.0
DeepSeek-V4-Pro10.0
Mistral Medium 3.510.0
Loading Atlas data…