ChessBench Accuracy
Games · 2026-07-28
ChessBench (chessbench.ai) has language models play full chess games against reference opponents and scores the moves; it is unrelated to the chess-bench.com puzzle benchmark carried as `chessbench`. Accuracy scores move quality against Stockfish as a weighted mean of branching, converting, and resisting accuracy (weights 10, 9, 1). Each part is scaled so that the random-move baseline is 0 and perfect play is 1, so the metric has a known ceiling but no floor: play worse than random scores below zero, and observed values reach -1.2. min_score is therefore left undeclared, which keeps this out of the score-level capability battery.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3.1 Pro Preview | 79.2 |
| o3 | 78.1 |
| GPT-5.2 | 77.4 |
| Gemini 3.5 Flash | 75.4 |
| Claude Fable 5 | 75.1 |
| GPT-5.4 | 75.0 |
| GPT-5 | 74.4 |
| Gemini 3 Flash Preview | 73.4 |
| Sonnet 5 | 71.7 |
| Gemini 3.6 Flash | 71.7 |
| GPT-5.1 | 71.6 |
| Opus 4.8 | 70.8 |
| GPT-5.5 | 70.5 |
| GPT-5 Nano | 67.6 |
| GPT-5 Mini | 67.4 |
| GPT-5.4 Nano | 67.1 |
| Opus 4.7 | 65.8 |
| O4 Mini | 65.3 |
| Gemini 3.1 Flash-Lite Preview | 64.8 |
| Gemini 3.1 Flash-Lite | 64.6 |
| Claude Opus 5 | 63.5 |
| Gemini 3.5 Flash-Lite | 59.8 |
| Opus 4.5 | 55.6 |
| GPT-5.4 Mini | 54.2 |
| o3-mini | 52.1 |
| Opus 4.6 | 50.7 |
| Sonnet 4.6 | 50.2 |
| GPT-4.1 | 48.7 |
| Sonnet 4 | 46.0 |
| Sonnet 4.5 | 44.7 |
| Opus 4 | 44.2 |
| GPT-4 0613 | 43.2 |
| Opus 4.1 | 41.1 |
| Haiku 4.5 | 39.3 |
| Gemini 2.0 Flash | 38.6 |
| GPT-4 Turbo | 38.2 |
| GPT-4o (2024-08-06) | 30.1 |
| GPT-4.1 Nano | 22.5 |
| GPT-4.1 Mini | 15.0 |
| GPT-4o (2024-11-20) | 14.3 |