Atlas

Benchmarks

← All benchmarks

Pencil Puzzle Bench - Agentic

Games · 2026-03-02

Pencil Puzzle Bench's Agentic setting gives models a tool loop for making moves, checking the board, and resetting after errors. The public leaderboard reports solve rates over model-specific puzzle subsets whose sizes vary, so each observation records the published aggregate rather than assuming a uniform item count.

Top models (higher is better)

ModelScore
Claude Fable 597.6
GPT-5.583.3
GPT-5.470.2
GPT-5.256.0
Opus 4.750.0
Gemini 3.5 Flash41.9
Qwen3.7-Max40.0
Opus 4.636.7
Gemini 3.1 Pro Preview33.3
Sonnet 4.626.7
GLM-5.226.7
GPT-5.2 Pro26.7
Kimi K2.620.0
Gemini 3 Pro Preview16.7
Kimi K2.7 Code16.7
Qwen3.7-Plus16.7
Qwen3.6 Plus (2026-04-02)10.0
MiniMax M37.1
Opus 4.56.7
Gemini 3 Flash Preview6.7
GPT-5.16.7
Grok 4.206.7
Sonnet 4.53.3
GPT-53.3
Grok 4.1 Fast3.3
Grok 4.33.3
Kimi K2.53.3
MiniMax M2.53.3
o33.3
DeepSeek-V3.20.0
DeepSeek-V4-Pro0.0
Devstral 20.0
Gemini 2.5 Pro0.0
Gemma 4 31B IT0.0
GLM-4.70.0
GLM-50.0
GPT-3.5 Turbo 01250.0
GPT-4.10.0
GPT-4o0.0
Grok Code Fast 10.0
Loading Atlas data…