Atlas

Benchmarks

← All benchmarks

USAMO 2025

Math · 2025-03-27

This MathArena track evaluates models on the six proof-based problems of USAMO 2025. Unlike MathArena's final-answer competitions, which are graded automatically, the proofs are graded by human verification, so scores report awarded partial credit rather than the share of problems answered correctly.

Top models (higher is better)

ModelScore
DeepSeek-R1-052830.1
Gemini 2.5 Pro24.4
O4 Mini19.1
DeepSeek-R14.8
Grok 34.8
Gemini 2.0 Flash Thinking Experimental (unspecified checkpoint)4.2
Claude 3.7 Sonnet3.6
QwQ-32B3.0
O1 Pro2.8
o3-mini2.1
Loading Atlas data…