Atlas

Benchmarks

← All benchmarks

ProgramBench

Code · 2026-05-05

Reverse-engineering benchmark in which an agent is given only a compiled executable and its documentation and must re-implement the program without access to any of its source. Scored across 200 tasks, from small utilities to projects the size of FFmpeg and SQLite, against more than 248,000 hidden behavioral tests. There is no partial credit: a task resolves only if every test for it passes.

Top models (higher is better)

ModelScore
Claude Opus 593.0
Claude Mythos 593.0
Opus 4.890.0
Kimi K377.8
Sonnet 572.1
GPT-5.468.6
GPT-5.6 Luna68.3
GPT-5.6 Terra66.3
Grok 4.565.2
Opus 4.764.2
GLM-5.263.7
Gemini 3.6 Flash59.9
Sonnet 4.659.1
Gemini 3.5 Flash56.4
GPT-5.4 Mini54.4
GLM-5.150.9
Kimi K2.7 Code49.4
Kimi K2.648.0
DeepSeek-V4-Pro47.8
Gemini 3.1 Pro Preview39.5
Gemini 3.5 Flash-Lite33.9
Gemini 3 Flash Preview33.8
Qwen3.6 Plus (2026-04-02)28.6
Muse Spark 1.127.1
Haiku 4.525.4
Inkling24.4
Grok 4.322.1
Laguna M.119.8
Laguna XS.216.5
Gemini 3.1 Flash-Lite Preview14.3
MiniMax M2.713.5
Nemotron 3 Ultra 550B A55B10.0
Loading Atlas data…