ChessBench Coherence
Games · 2026-07-28
ChessBench (chessbench.ai) has language models play full chess games against reference opponents and scores the moves; it is unrelated to the chess-bench.com puzzle benchmark carried as `chessbench`. Coherence is the rate of rule-compliant play, defined by the leaderboard as move_coherence * max(0.01, game_coherence), where move_coherence is the fraction of attempted moves that were legal and game_coherence the fraction of fully-legal games. A product of two rates rather than a proportion of scored items, but bounded on [0, 1] by construction and reported here in percentage points.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.1 | 100.0 |
| GPT-5.4 | 100.0 |
| GPT-5.5 | 100.0 |
| Sonnet 5 | 97.7 |
| GPT-5 | 91.9 |
| GPT-5 Nano | 88.6 |
| Gemini 3.6 Flash | 86.3 |
| o3 | 85.3 |
| Gemini 3.1 Pro Preview | 85.2 |
| Opus 4.6 | 83.9 |
| Claude Fable 5 | 82.6 |
| Opus 4.8 | 79.7 |
| Sonnet 4.6 | 79.1 |
| Gemini 3.5 Flash | 78.0 |
| GPT-5.4 Nano | 77.8 |
| GPT-5.2 | 76.9 |
| Gemini 3 Flash Preview | 75.8 |
| O4 Mini | 73.0 |
| GPT-5 Mini | 65.7 |
| Claude Opus 5 | 63.8 |
| GPT-5.4 Mini | 63.5 |
| Gemini 3.1 Flash-Lite | 63.2 |
| Opus 4 | 58.8 |
| Gemini 3.5 Flash-Lite | 58.3 |
| Opus 4.5 | 57.5 |
| o3-mini | 55.8 |
| Gemini 3.1 Flash-Lite Preview | 39.4 |
| Opus 4.7 | 38.7 |
| GPT-4 0613 | 36.9 |
| Haiku 4.5 | 36.0 |
| GPT-4.1 Mini | 35.4 |
| Opus 4.1 | 34.4 |
| GPT-4o (2024-08-06) | 33.8 |
| Sonnet 4.5 | 31.6 |
| Sonnet 4 | 31.4 |
| GPT-4o (2024-11-20) | 31.4 |
| Gemini 2.0 Flash | 27.9 |
| GPT-4.1 | 26.2 |
| GPT-4 Turbo | 25.5 |
| GPT-3.5 Turbo 0125 | 15.2 |