Atlas

Benchmarks

← All benchmarks

LiveBench

General QA · 2024-06-12

LiveBench is a frequently updated, contamination-resistant benchmark with objective tasks spanning mathematics, coding, reasoning, language, instruction following, data analysis, and agentic coding. Scores are reported as overall or category averages.

Top models (higher is better)

ModelScore
GPT-5.579.9
Claude Fable 579.5
Opus 4.878.9
GPT-5.478.0
Gemini 3.1 Pro Preview77.1
Opus 4.776.5
Sonnet 574.8
Gemini 3.5 Flash74.6
GPT-5.274.6
Opus 4.674.5
GPT-5.2-Codex74.0
GLM-5.273.2
Qwen3.7-Max73.1
Sonnet 4.673.0
Opus 4.572.6
DeepSeek-V4-Pro71.6
Kimi K2.670.5
GPT-5.4 Nano69.6
Qwen3.6 Plus (2026-04-02)68.9
Kimi K2.7 Code68.4
Grok Build 0.167.8
MiniMax M367.3
GPT-5.4 Mini66.4
DeepSeek-V4-Flash65.5
Qwen3.6 27B64.0
Grok 4.362.3
Loading Atlas data…