Atlas

Benchmarks

← All benchmarks

ProgramBench (problem counts) - Almost

Code · 2026-05-05

The Almost split of ProgramBench (problem counts). This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Claude Opus 583.0
Claude Fable 566.0
GPT-5.6 Sol46.0
Opus 4.831.0
Sonnet 527.0
Opus 4.719.0
GLM-5.219.0
GPT-5.519.0
GPT-5.6 Luna18.0
GPT-5.414.0
GPT-5.6 Terra13.0
Grok 4.59.0
GPT-5.4 Mini6.0
Gemini 3.6 Flash5.0
Muse Spark 1.15.0
Sonnet 4.64.0
Gemini 3.5 Flash3.0
Gemini 3.1 Pro Preview2.0
GLM-5.12.0
Qwen3.6 Plus (2026-04-02)2.0
Kimi K2.61.0
Kimi K2.7 Code1.0
Nemotron 3 Ultra 550B A55B1.0
Haiku 4.50.0
DeepSeek-V4-Pro0.0
Gemini 3.1 Flash-Lite Preview0.0
Gemini 3.5 Flash-Lite0.0
Gemini 3 Flash Preview0.0
Grok 4.30.0
Inkling0.0
Laguna M.10.0
Laguna XS.20.0
MiniMax M2.70.0
Loading Atlas data…