Atlas

Benchmarks

← All benchmarks

LiveCodeBench (Artificial Analysis)

Code · 2024-03-12

Artificial Analysis LiveCodeBench leaderboard/evaluation surface. Keep separate from official LiveCodeBench windows, model-release comparison tables, and Vals AI public-dataset runs because the batch shows different scores for the same config under those sources.

Top models (higher is better)

ModelScore
Gemini 3 Pro Preview91.7
Gemini 3 Flash Preview90.8
DeepSeek-V3.2-Speciale89.6
GLM-4.789.4
GPT-5.289.4
gpt-oss-120b87.8
Opus 4.587.1
GPT-5.186.8
MiMo-V2-Flash86.8
DeepSeek-V3.286.2
O4 Mini85.9
Kimi K2 Thinking85.3
GPT-5.1-Codex84.9
GPT-584.6
GPT-5-Codex84.0
GPT-5 Mini83.8
GPT-5.1-Codex mini83.6
Grok 4 Fast83.2
MiniMax M282.6
Grok 4.1 Fast82.2
Grok 481.9
ERNIE 5.081.2
MiniMax M2.181.0
o380.8
Apriel-1.6-15B-Thinker80.7
Gemini 2.5 Pro80.1
DeepSeek-V3.1-Terminus79.8
DeepSeek-V3.2-Exp78.9
GPT-5 Nano78.9
Qwen3 235B A22B Thinking 250778.8
DeepSeek-V3.178.4
Qwen3-Next-80B-A3B-Thinking78.4
gpt-oss-20b77.7
INTELLECT-377.7
DeepSeek-R1-052877.0
Gemini 2.5 Pro Preview 05-0677.0
K-EXAONE 236B-A23B76.8
Qwen3-Max (2025-09-23)76.7
Doubao-Seed-Code76.6
Seed OSS 36B Instruct76.5
Loading Atlas data…