Atlas

Benchmarks

← All benchmarks

LiveBench Instruction Following Average

General QA · 2024-06-12

LiveBench leaderboard aggregate for livebench instruction following average. LiveBench is refreshed over time to reduce contamination and covers reasoning, coding, mathematics, language, data analysis, instruction following, and agentic coding.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview79.1
Gemini 3.5 Flash75.6
Qwen3.7-Max74.0
Opus 4.872.4
Claude Fable 572.0
GPT-5.571.4
GPT-5.470.2
GPT-5.4 Nano67.2
Opus 4.766.7
GPT-5.2-Codex66.5
Grok Build 0.165.2
Kimi K2.664.4
Sonnet 563.9
Opus 4.663.3
Sonnet 4.663.2
DeepSeek-V4-Flash63.1
Grok 4.362.8
Opus 4.562.5
DeepSeek-V4-Pro62.4
GLM-5.262.3
GPT-5.261.8
GPT-5.4 Mini59.8
Qwen3.6 Plus (2026-04-02)58.3
MiniMax M357.5
Kimi K2.7 Code56.3
Qwen3.6 27B53.2
Loading Atlas data…