LiveBench Instruction Following Average
General QA · 2024-06-12
LiveBench leaderboard aggregate for livebench instruction following average. LiveBench is refreshed over time to reduce contamination and covers reasoning, coding, mathematics, language, data analysis, instruction following, and agentic coding.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3.1 Pro Preview | 79.1 |
| Gemini 3.5 Flash | 75.6 |
| Qwen3.7-Max | 74.0 |
| Opus 4.8 | 72.4 |
| Claude Fable 5 | 72.0 |
| GPT-5.5 | 71.4 |
| GPT-5.4 | 70.2 |
| GPT-5.4 Nano | 67.2 |
| Opus 4.7 | 66.7 |
| GPT-5.2-Codex | 66.5 |
| Grok Build 0.1 | 65.2 |
| Kimi K2.6 | 64.4 |
| Sonnet 5 | 63.9 |
| Opus 4.6 | 63.3 |
| Sonnet 4.6 | 63.2 |
| DeepSeek-V4-Flash | 63.1 |
| Grok 4.3 | 62.8 |
| Opus 4.5 | 62.5 |
| DeepSeek-V4-Pro | 62.4 |
| GLM-5.2 | 62.3 |
| GPT-5.2 | 61.8 |
| GPT-5.4 Mini | 59.8 |
| Qwen3.6 Plus (2026-04-02) | 58.3 |
| MiniMax M3 | 57.5 |
| Kimi K2.7 Code | 56.3 |
| Qwen3.6 27B | 53.2 |