Atlas

Benchmarks

← All benchmarks

LiveBench Language Average

Classic NLP · 2024-06-12

LiveBench leaderboard aggregate for livebench language average. LiveBench is refreshed over time to reduce contamination and covers reasoning, coding, mathematics, language, data analysis, instruction following, and agentic coding.

Top models (higher is better)

ModelScore
Claude Fable 589.5
GPT-5.587.8
Gemini 3.1 Pro Preview85.4
Gemini 3.5 Flash84.6
Opus 4.683.3
GPT-5.482.6
Opus 4.881.4
Opus 4.581.3
GPT-5.279.8
Qwen3.7-Max79.7
DeepSeek-V4-Pro78.1
Opus 4.777.9
Kimi K2.7 Code77.9
MiniMax M376.8
GLM-5.276.2
Sonnet 4.676.1
Kimi K2.675.1
Qwen3.6 Plus (2026-04-02)75.0
Sonnet 575.0
GPT-5.2-Codex73.7
Grok 4.373.6
Grok Build 0.172.5
GPT-5.4 Mini71.0
DeepSeek-V4-Flash70.1
Qwen3.6 27B63.3
GPT-5.4 Nano62.5
Loading Atlas data…