Atlas

Benchmarks

← All benchmarks

HMMT Feb 2025

Math · 2025-02-15

This MathArena track evaluates models on the HMMT Feb 2025 problem set. Scores report the percentage of problems answered correctly.

Top models (higher is better)

ModelScore
GPT-5.2100.0
Step 3.5 Flash98.3
DeepSeek-V3.2-Speciale97.5
Gemini 3 Flash Preview97.5
Gemini 3 Pro Preview97.5
GLM-597.5
Grok 495.0
GLM 4.693.3
GPT-5.193.3
Kimi K2.593.3
Kimi K2 Thinking93.3
DeepSeek-V3.292.5
Grok 4 Fast91.7
DeepSeek-V3.2-Exp90.0
gpt-oss-120b90.0
Grok 4.1 Fast90.0
GPT-5 Mini89.2
GPT-588.3
DeepSeek-V3.185.8
Falcon H1R 7B84.2
O4 Mini83.3
Gemini 2.5 Pro82.5
Gemini 2.5 Pro Preview 05-0680.8
GLM-4.577.5
o377.5
DeepSeek-R1-052876.7
gpt-oss-20b76.7
QED-Nano76.7
GPT-5 Nano74.2
Grok 3 Mini74.2
GLM 4.5 Air69.2
Sonnet 4.567.5
o3-mini67.5
K2-Think65.0
Gemini 2.5 Flash64.2
Qwen3-235B-A22B62.5
Opus 460.0
Qwen3-30B-A3B50.8
O148.3
QwQ-32B47.5
Loading Atlas data…