ProgramBench
Code · 2026-05-05
Reverse-engineering benchmark in which an agent is given only a compiled executable and its documentation and must re-implement the program without access to any of its source. Scored across 200 tasks, from small utilities to projects the size of FFmpeg and SQLite, against more than 248,000 hidden behavioral tests. There is no partial credit: a task resolves only if every test for it passes.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 93.0 |
| Claude Mythos 5 | 93.0 |
| Opus 4.8 | 90.0 |
| Kimi K3 | 77.8 |
| Sonnet 5 | 72.1 |
| GPT-5.4 | 68.6 |
| GPT-5.6 Luna | 68.3 |
| GPT-5.6 Terra | 66.3 |
| Grok 4.5 | 65.2 |
| Opus 4.7 | 64.2 |
| GLM-5.2 | 63.7 |
| Gemini 3.6 Flash | 59.9 |
| Sonnet 4.6 | 59.1 |
| Gemini 3.5 Flash | 56.4 |
| GPT-5.4 Mini | 54.4 |
| GLM-5.1 | 50.9 |
| Kimi K2.7 Code | 49.4 |
| Kimi K2.6 | 48.0 |
| DeepSeek-V4-Pro | 47.8 |
| Gemini 3.1 Pro Preview | 39.5 |
| Gemini 3.5 Flash-Lite | 33.9 |
| Gemini 3 Flash Preview | 33.8 |
| Qwen3.6 Plus (2026-04-02) | 28.6 |
| Muse Spark 1.1 | 27.1 |
| Haiku 4.5 | 25.4 |
| Inkling | 24.4 |
| Grok 4.3 | 22.1 |
| Laguna M.1 | 19.8 |
| Laguna XS.2 | 16.5 |
| Gemini 3.1 Flash-Lite Preview | 14.3 |
| MiniMax M2.7 | 13.5 |
| Nemotron 3 Ultra 550B A55B | 10.0 |