Atlas

Benchmarks

← All benchmarks

Final-Answer Comps — Overall

Math · 2025-02-15

Accuracy on MathArena's Final-Answer Comps overall set.

Top models (higher is better)

ModelScore
GPT-5.594.3
Opus 4.891.8
Kimi K387.8
Gemini 3.1 Pro Preview86.5
GPT-5.483.1
Opus 4.678.5
DeepSeek-V4-Pro76.6
DeepSeek-V4-Flash76.5
Gemini 3.5 Flash76.3
Opus 4.773.6
Kimi K2.672.9
GPT-5.272.0
Gemini 3.6 Flash70.8
Step 3.7 Flash68.5
Gemini 3 Flash Preview67.6
GLM-5.267.6
GLM-5.167.1
Gemini 3 Pro Preview67.0
Step 3.5 Flash66.8
GLM-565.7
Kimi K2.562.3
Grok 4.1 Fast60.9
Nemotron 3 Super 120B A12B60.4
DeepSeek-V3.257.7
Qwen3.5-27B56.7
Qwen3.5 35B A3B56.0
Qwen3.5-9B48.5
Qwen3 30B A3B Thinking 250747.8
QED-Nano42.7
Qwen3 4B Thinking 250738.5
Loading Atlas data…