Atlas

Benchmarks

← All benchmarks

ProofBench v1.1

Math

Vals ProofBench v1.1 measures success on 100 private formal-mathematics problems using Lean 4.25.2 and the pinned Mathlib revision. General models receive Mathlib search, Lean execution and a final proof-submission tool, with up to 40 interaction turns and no internet. Vals explicitly states that v1.1 scores are not comparable with earlier revisions.

Top models (higher is better)

ModelScore
Claude Fable 5.1100.0
Claude Opus 5.5100.0
Claude Opus 599.0
GPT-6 Astra99.0
Claude Fable 595.0
Kimi K387.0
GPT-5.6 Sol83.0
GPT-6 Sol83.0
Sonnet 577.0
Hy4 Preview75.0
GPT-5.6 Terra74.0
MiMo-V2.6-Pro70.0
GPT-6 Luna64.0
MiMo-V2.6-Flash63.0
GPT-5.6 Luna60.0
Gemini 3.7 Flash58.0
Muse Spark 1.358.0
Qwen3.8-Max58.0
DeepSeek-V4-Flash-073156.0
DeepSeek V4.1 Flash54.0
Grok 4.651.0
DeepSeek V4 Pro 081350.0
GLM-5.349.0
Gemini 3.8 Flash48.0
Muse Spark 1.243.0
Gemini 3.5 Flash31.0
Grok 4.531.0
Gemini 3.1 Pro Preview26.0
Grok 4.726.0
MiMo-V2.5-Pro22.0
GLM-5.3 Flash21.0
MiniMax M318.0
DeepSeek-V4-Pro16.0
MiMo-V2.516.0
Qwen3.8 27B16.0
Mistral Medium 3.59.0
Inkling-Small6.0
Mercury 2.53.0
Inkling0.0
Loading Atlas data…