LiveBench
General QA · 2024-06-12
LiveBench is a frequently updated, contamination-resistant benchmark with objective tasks spanning mathematics, coding, reasoning, language, instruction following, data analysis, and agentic coding. Scores are reported as overall or category averages.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 2.5 Pro Experimental 03-25 | 82.3 |
| GPT-5.5 | 79.9 |
| Claude Fable 5 | 79.5 |
| Opus 4.8 | 78.9 |
| GPT-5.1 | 78.8 |
| GPT-5.4 | 78.0 |
| Gemini 3.1 Pro Preview | 77.1 |
| Opus 4.7 | 76.5 |
| Claude 3.7 Sonnet | 76.1 |
| o3-mini | 75.9 |
| O1 | 75.7 |
| Sonnet 5 | 74.8 |
| Gemini 3.5 Flash | 74.6 |
| GPT-5.2 | 74.6 |
| Opus 4.6 | 74.5 |
| GPT-5.2-Codex | 74.0 |
| GLM-5.2 | 73.2 |
| Qwen3.7-Max | 73.1 |
| Sonnet 4.6 | 73.0 |
| Opus 4.5 | 72.6 |
| QwQ-32B | 72.0 |
| DeepSeek-V4-Pro | 71.6 |
| DeepSeek-R1 | 71.6 |
| Kimi K2.6 | 70.5 |
| GPT-5.4 Nano | 69.6 |
| GPT-4.5 | 69.0 |
| Qwen3.6 Plus (2026-04-02) | 68.9 |
| Kimi K2.7 Code | 68.4 |
| Grok Build 0.1 | 67.8 |
| MiniMax M3 | 67.3 |
| Gemini 2.0 Flash Thinking Experimental 01-21 | 66.9 |
| DeepSeek-V3-0324 | 66.9 |
| GPT-5.4 Mini | 66.4 |
| DeepSeek-V4-Flash | 65.5 |
| Gemini 2.0 Pro Experimental 02-05 | 65.1 |
| Gemini Exp-1206 | 64.1 |
| Qwen3.6 27B | 64.0 |
| Qwen2.5-Max | 62.3 |
| Grok 4.3 | 62.3 |
| Gemini 2.0 Flash 001 | 61.5 |