Atlas

Benchmarks

← All benchmarks

LiveBench Mathematics Average

Math · 2024-06-12

LiveBench leaderboard aggregate for livebench mathematics average. LiveBench is refreshed over time to reduce contamination and covers reasoning, coding, mathematics, language, data analysis, instruction following, and agentic coding.

Top models (higher is better)

ModelScore
GPT-5.595.9
Claude Fable 595.7
Opus 4.895.3
GPT-5.494.2
GPT-5.293.2
Sonnet 592.9
Opus 4.792.8
Gemini 3.1 Pro Preview91.0
GPT-5.4 Nano91.0
DeepSeek-V4-Pro90.7
Opus 4.590.4
GLM-5.289.8
Opus 4.689.3
GPT-5.2-Codex88.8
Gemini 3.5 Flash88.2
Sonnet 4.687.0
Qwen3.7-Max85.3
Grok 4.384.3
Kimi K2.684.3
Qwen3.6 Plus (2026-04-02)83.7
Qwen3.6 27B79.9
DeepSeek-V4-Flash79.7
Kimi K2.7 Code79.6
GPT-5.4 Mini78.5
Grok Build 0.178.4
MiniMax M377.0
Loading Atlas data…