LiveBench Data Analysis Average
Professional Work · 2024-06-12
LiveBench leaderboard aggregate for livebench data analysis average. LiveBench is refreshed over time to reduce contamination and covers reasoning, coding, mathematics, language, data analysis, instruction following, and agentic coding.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 81.6 |
| GPT-5.4 | 79.3 |
| Claude Fable 5 | 78.7 |
| Gemini 3.1 Pro Preview | 78.5 |
| Opus 4.8 | 78.3 |
| Opus 4.7 | 78.3 |
| GPT-5.2-Codex | 78.2 |
| GPT-5.2 | 78.2 |
| Sonnet 4.6 | 78.0 |
| MiniMax M3 | 76.2 |
| DeepSeek-V4-Pro | 74.5 |
| Opus 4.5 | 74.4 |
| GLM-5.2 | 73.7 |
| Qwen3.7-Max | 71.8 |
| Sonnet 5 | 71.7 |
| GPT-5.4 Mini | 70.8 |
| Grok Build 0.1 | 70.8 |
| Qwen3.6 27B | 70.4 |
| Qwen3.6 Plus (2026-04-02) | 69.9 |
| Opus 4.6 | 69.9 |
| DeepSeek-V4-Flash | 68.0 |
| GPT-5.4 Nano | 67.6 |
| Kimi K2.6 | 65.1 |
| Gemini 3.5 Flash | 64.9 |
| Kimi K2.7 Code | 62.7 |
| Grok 4.3 | 55.8 |