LiveCodeBench (Artificial Analysis)
Code · 2024-03-12
Artificial Analysis LiveCodeBench leaderboard/evaluation surface. Keep separate from official LiveCodeBench windows, model-release comparison tables, and Vals AI public-dataset runs because the batch shows different scores for the same config under those sources.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3 Pro Preview | 91.7 |
| Gemini 3 Flash Preview | 90.8 |
| DeepSeek-V3.2-Speciale | 89.6 |
| GLM-4.7 | 89.4 |
| GPT-5.2 | 89.4 |
| gpt-oss-120b | 87.8 |
| Opus 4.5 | 87.1 |
| GPT-5.1 | 86.8 |
| MiMo-V2-Flash | 86.8 |
| DeepSeek-V3.2 | 86.2 |
| O4 Mini | 85.9 |
| Kimi K2 Thinking | 85.3 |
| GPT-5.1-Codex | 84.9 |
| GPT-5 | 84.6 |
| GPT-5-Codex | 84.0 |
| GPT-5 Mini | 83.8 |
| GPT-5.1-Codex mini | 83.6 |
| Grok 4 Fast | 83.2 |
| MiniMax M2 | 82.6 |
| Grok 4.1 Fast | 82.2 |
| Grok 4 | 81.9 |
| ERNIE 5.0 | 81.2 |
| MiniMax M2.1 | 81.0 |
| o3 | 80.8 |
| Apriel-1.6-15B-Thinker | 80.7 |
| Gemini 2.5 Pro | 80.1 |
| DeepSeek-V3.1-Terminus | 79.8 |
| DeepSeek-V3.2-Exp | 78.9 |
| GPT-5 Nano | 78.9 |
| Qwen3 235B A22B Thinking 2507 | 78.8 |
| DeepSeek-V3.1 | 78.4 |
| Qwen3-Next-80B-A3B-Thinking | 78.4 |
| gpt-oss-20b | 77.7 |
| INTELLECT-3 | 77.7 |
| DeepSeek-R1-0528 | 77.0 |
| Gemini 2.5 Pro Preview 05-06 | 77.0 |
| K-EXAONE 236B-A23B | 76.8 |
| Qwen3-Max (2025-09-23) | 76.7 |
| Doubao-Seed-Code | 76.6 |
| Seed OSS 36B Instruct | 76.5 |