Pencil Puzzle Bench - Agentic
Games · 2026-03-02
Pencil Puzzle Bench's Agentic setting gives models a tool loop for making moves, checking the board, and resetting after errors. The public leaderboard reports solve rates over model-specific puzzle subsets whose sizes vary, so each observation records the published aggregate rather than assuming a uniform item count.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 97.6 |
| GPT-5.5 | 83.3 |
| GPT-5.4 | 70.2 |
| GPT-5.2 | 56.0 |
| Opus 4.7 | 50.0 |
| Gemini 3.5 Flash | 41.9 |
| Qwen3.7-Max | 40.0 |
| Opus 4.6 | 36.7 |
| Gemini 3.1 Pro Preview | 33.3 |
| Sonnet 4.6 | 26.7 |
| GLM-5.2 | 26.7 |
| GPT-5.2 Pro | 26.7 |
| Kimi K2.6 | 20.0 |
| Gemini 3 Pro Preview | 16.7 |
| Kimi K2.7 Code | 16.7 |
| Qwen3.7-Plus | 16.7 |
| Qwen3.6 Plus (2026-04-02) | 10.0 |
| MiniMax M3 | 7.1 |
| Opus 4.5 | 6.7 |
| Gemini 3 Flash Preview | 6.7 |
| GPT-5.1 | 6.7 |
| Grok 4.20 | 6.7 |
| Sonnet 4.5 | 3.3 |
| GPT-5 | 3.3 |
| Grok 4.1 Fast | 3.3 |
| Grok 4.3 | 3.3 |
| Kimi K2.5 | 3.3 |
| MiniMax M2.5 | 3.3 |
| o3 | 3.3 |
| DeepSeek-V3.2 | 0.0 |
| DeepSeek-V4-Pro | 0.0 |
| Devstral 2 | 0.0 |
| Gemini 2.5 Pro | 0.0 |
| Gemma 4 31B IT | 0.0 |
| GLM-4.7 | 0.0 |
| GLM-5 | 0.0 |
| GPT-3.5 Turbo 0125 | 0.0 |
| GPT-4.1 | 0.0 |
| GPT-4o | 0.0 |
| Grok Code Fast 1 | 0.0 |