Atlas

Benchmarks

← All benchmarks

ProgramBench (problem counts) - Fully Resolved

Code · 2026-05-05

The Fully Resolved split of ProgramBench (problem counts). This child benchmark reports raw counts of tasks whose submitted implementation passes all hidden behavioral tests.

Top models (higher is better)

ModelScore
Claude Opus 56.0
Claude Fable 54.0
GPT-5.6 Sol3.0
Opus 4.82.0
Sonnet 4.61.0
GLM-5.21.0
GPT-5.41.0
GPT-5.51.0
GPT-5.6 Terra1.0
Haiku 4.50.0
Opus 4.70.0
Sonnet 50.0
DeepSeek-V4-Pro0.0
Gemini 3.1 Flash-Lite Preview0.0
Gemini 3.1 Pro Preview0.0
Gemini 3.5 Flash0.0
Gemini 3.5 Flash-Lite0.0
Gemini 3.6 Flash0.0
Gemini 3 Flash Preview0.0
GLM-5.10.0
GPT-5.4 Mini0.0
GPT-5.6 Luna0.0
Grok 4.30.0
Grok 4.50.0
Inkling0.0
Kimi K2.60.0
Kimi K2.7 Code0.0
Laguna M.10.0
Laguna XS.20.0
MiniMax M2.70.0
Muse Spark 1.10.0
Nemotron 3 Ultra 550B A55B0.0
Qwen3.6 Plus (2026-04-02)0.0
Loading Atlas data…