Atlas

Benchmarks

← All benchmarks

LiveBench

General QA · 2024-06-12

LiveBench is a frequently updated, contamination-resistant benchmark with objective tasks spanning mathematics, coding, reasoning, language, instruction following, data analysis, and agentic coding. Scores are reported as overall or category averages.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro Experimental 03-2582.3
GPT-5.579.9
Claude Fable 579.5
Opus 4.878.9
GPT-5.178.8
GPT-5.478.0
Gemini 3.1 Pro Preview77.1
Opus 4.776.5
Claude 3.7 Sonnet76.1
o3-mini75.9
O175.7
Sonnet 574.8
Gemini 3.5 Flash74.6
GPT-5.274.6
Opus 4.674.5
GPT-5.2-Codex74.0
GLM-5.273.2
Qwen3.7-Max73.1
Sonnet 4.673.0
Opus 4.572.6
QwQ-32B72.0
DeepSeek-V4-Pro71.6
DeepSeek-R171.6
Kimi K2.670.5
GPT-5.4 Nano69.6
GPT-4.569.0
Qwen3.6 Plus (2026-04-02)68.9
Kimi K2.7 Code68.4
Grok Build 0.167.8
MiniMax M367.3
Gemini 2.0 Flash Thinking Experimental 01-2166.9
DeepSeek-V3-032466.9
GPT-5.4 Mini66.4
DeepSeek-V4-Flash65.5
Gemini 2.0 Pro Experimental 02-0565.1
Gemini Exp-120664.1
Qwen3.6 27B64.0
Qwen2.5-Max62.3
Grok 4.362.3
Gemini 2.0 Flash 00161.5
Loading Atlas data…