Atlas

Benchmarks

← All benchmarks

LiveCodeBench (Vals Public) - Hard

Code · 2024-03-12

The hard difficulty slice of Vals AI's LiveCodeBench public-dataset evaluation surface.

Top models (higher is better)

ModelScore
Claude Fable 579.7
Claude Opus 577.4
Gemini 3.1 Pro Preview76.6
GPT-5.3-Codex75.1
Opus 4.874.6
Gemini 3.6 Flash74.6
GPT-5.2-Codex74.6
Kimi K374.3
Grok 4.573.4
Qwen3.7-Max73.4
Gemini 3.5 Flash73.1
GPT-5.6 Terra72.9
DeepSeek-V4-Pro72.6
GPT-5 Mini71.7
Kimi K2.671.4
GPT-5.171.1
Gemini 3 Pro Preview70.9
GPT-5.270.1
Gemini 3 Flash Preview69.7
GPT-5.569.7
Muse Spark 1.169.1
Nemotron 3 Ultra 550B A55B69.1
GPT-5.1-Codex68.9
GPT-5-Codex68.6
Inkling68.3
GPT-5.467.7
GPT-5.6 Sol67.4
Qwen3.5 Plus (2026-02-15)67.1
Qwen3.6 Plus (2026-04-02)67.1
Opus 4.766.9
Sonnet 566.9
Opus 4.666.3
GPT-566.3
GPT-5.4 Nano66.0
Grok 4.2066.0
Grok 4.366.0
GPT-5.1-Codex-Max65.4
gpt-oss-120b64.6
Kimi K2.564.3
Opus 4.564.0
Loading Atlas data…