Atlas

Benchmarks

← All benchmarks

Apex Shortlist

Math · 2025-10-25

This MathArena track evaluates models on the Apex Shortlist problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
GPT-5.598.4
Gemini 3.1 Pro Preview92.0
Kimi K392.0
Opus 4.890.4
DeepSeek-V4-Flash89.4
DeepSeek-V4-Pro87.8
Opus 4.686.7
Gemini 3.5 Flash82.5
GPT-5.481.4
GPT-5.279.3
Kimi K2.677.1
Step 3.7 Flash76.6
GLM-5.171.8
Gemini 3.6 Flash71.3
Step 3.5 Flash70.7
DeepSeek-V3.2-Speciale69.7
Gemini 3 Flash Preview68.6
GLM-568.6
GLM-5.268.1
Gemini 3 Pro Preview66.5
Opus 4.763.8
Qwen3.5 397B A17B60.1
Grok 4.1 Fast57.9
Grok 457.5
Kimi K2.557.5
Nemotron 3 Super 120B A12B57.5
GPT-5.155.9
Grok 4 Fast55.3
Qwen3.5-27B52.1
DeepSeek-V3.250.5
Kimi K2 Thinking47.3
gpt-oss-120b45.2
Qwen3.5 35B A3B44.7
GPT-5 Mini39.9
Qwen3.5-4B32.5
Qwen3.5-9B29.8
Qwen3 30B A3B Thinking 250723.4
QED-Nano22.3
Qwen3 4B Thinking 250716.5
Qwen3.5-2B9.0
Loading Atlas data…