Atlas

Benchmarks

← All benchmarks

BRUMO 2025

Math · 2025-05-29

This MathArena track evaluates models on the BRUMO 2025 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
Gemini 3 Flash Preview100.0
Step 3.5 Flash100.0
DeepSeek-V3.2-Speciale99.2
GLM-599.2
Gemini 3 Pro Preview98.3
GPT-5.298.3
Kimi K2.598.3
Grok 4.1 Fast97.5
DeepSeek-V3.296.7
DeepSeek-V3.2-Exp95.8
Grok 4 Fast95.8
o395.8
Grok 495.0
GLM 4.694.2
GPT-5.193.3
Kimi K2 Thinking93.3
DeepSeek-R1-052892.5
GLM-4.592.5
gpt-oss-120b92.5
GPT-591.7
Sonnet 4.590.8
DeepSeek-V3.190.0
Gemini 2.5 Pro90.0
GLM 4.5 Air90.0
GPT-5 Mini90.0
Gemini 2.5 Pro Preview 05-0689.2
QED-Nano87.5
gpt-oss-20b86.7
O4 Mini86.7
Qwen3-235B-A22B86.7
Falcon H1R 7B85.8
Grok 3 Mini85.0
Gemini 2.5 Flash83.3
K2-Think83.3
Opus 481.7
DeepSeek-R180.8
GPT-5 Nano80.8
Qwen3-30B-A3B77.5
DeepSeek R1 Distill Qwen 14B68.3
DeepSeek-R1-Distill-Qwen-32B68.3
Loading Atlas data…