Atlas

Benchmarks

← All benchmarks

CMIMC 2025

Math · 2025-05-29

This MathArena track evaluates models on the CMIMC 2025 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
DeepSeek-V3.2-Speciale94.4
Step 3.5 Flash93.8
GLM-592.5
GPT-5.191.9
Kimi K2 Thinking91.9
GPT-5.291.3
Kimi K2.591.3
Gemini 3 Flash Preview90.6
Gemini 3 Pro Preview90.0
GPT-590.0
GLM 4.688.8
gpt-oss-120b85.6
Grok 4 Fast85.6
GPT-5 Mini84.4
Grok 4.1 Fast84.4
O4 Mini84.4
DeepSeek-V3.283.8
Grok 483.8
DeepSeek-V3.181.3
o379.4
DeepSeek-V3.2-Exp75.6
GPT-5 Nano73.8
gpt-oss-20b72.5
GLM-4.571.3
GLM 4.5 Air70.6
Falcon H1R 7B70.0
DeepSeek-R1-052869.4
Sonnet 4.566.9
Grok 3 Mini66.3
K2-Think65.6
QED-Nano59.4
Gemini 2.5 Pro58.1
Gemini 2.5 Flash51.9
Loading Atlas data…