Atlas

Benchmarks

← All benchmarks

SMT 2025

Math · 2025-05-29

This MathArena track evaluates models on the SMT 2025 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
GPT-5.296.2
Gemini 3 Pro Preview93.4
Gemini 3 Flash Preview92.9
GPT-592.0
Step 3.5 Flash91.5
GLM-591.0
GPT-5.191.0
Kimi K2 Thinking91.0
GLM 4.690.6
Kimi K2.590.6
DeepSeek-V3.2-Speciale89.2
GPT-5 Mini89.2
O4 Mini88.7
DeepSeek-V3.287.7
gpt-oss-120b87.7
o387.7
Falcon H1R 7B85.8
Grok 485.8
GPT-5 Nano85.4
DeepSeek-V3.2-Exp84.9
Gemini 2.5 Pro84.9
Grok 4.1 Fast84.6
Grok 4 Fast84.4
Sonnet 4.584.0
DeepSeek-V3.184.0
DeepSeek-R1-052883.0
gpt-oss-20b83.0
GLM-4.582.1
K2-Think79.7
QED-Nano79.7
Grok 3 Mini78.8
GLM 4.5 Air77.4
Qwen3-235B-A22B76.9
Gemini 2.5 Flash75.5
Qwen3-30B-A3B67.9
DeepSeek-R167.0
DeepSeek-R1-Distill-Llama-70B60.9
DeepSeek-R1-Distill-Qwen-32B60.4
Claude 3.7 Sonnet56.6
DeepSeek R1 Distill Qwen 14B54.7
Loading Atlas data…