Atlas

Benchmarks

← All benchmarks

ChessBench Coherence

Games · 2026-07-28

ChessBench (chessbench.ai) has language models play full chess games against reference opponents and scores the moves; it is unrelated to the chess-bench.com puzzle benchmark carried as `chessbench`. Coherence is the rate of rule-compliant play, defined by the leaderboard as move_coherence * max(0.01, game_coherence), where move_coherence is the fraction of attempted moves that were legal and game_coherence the fraction of fully-legal games. A product of two rates rather than a proportion of scored items, but bounded on [0, 1] by construction and reported here in percentage points.

Top models (higher is better)

ModelScore
GPT-5.1100.0
GPT-5.4100.0
GPT-5.5100.0
Sonnet 597.7
GPT-591.9
GPT-5 Nano88.6
Gemini 3.6 Flash86.3
o385.3
Gemini 3.1 Pro Preview85.2
Opus 4.683.9
Claude Fable 582.6
Opus 4.879.7
Sonnet 4.679.1
Gemini 3.5 Flash78.0
GPT-5.4 Nano77.8
GPT-5.276.9
Gemini 3 Flash Preview75.8
O4 Mini73.0
GPT-5 Mini65.7
Claude Opus 563.8
GPT-5.4 Mini63.5
Gemini 3.1 Flash-Lite63.2
Opus 458.8
Gemini 3.5 Flash-Lite58.3
Opus 4.557.5
o3-mini55.8
Gemini 3.1 Flash-Lite Preview39.4
Opus 4.738.7
GPT-4 061336.9
Haiku 4.536.0
GPT-4.1 Mini35.4
Opus 4.134.4
GPT-4o (2024-08-06)33.8
Sonnet 4.531.6
Sonnet 431.4
GPT-4o (2024-11-20)31.4
Gemini 2.0 Flash27.9
GPT-4.126.2
GPT-4 Turbo25.5
GPT-3.5 Turbo 012515.2
Loading Atlas data…