ChessBench Elo
Games · 2026-07-28
ChessBench (chessbench.ai) has language models play full chess games against reference opponents and scores the moves; it is unrelated to the chess-bench.com puzzle benchmark carried as `chessbench`. Monte Carlo Elo ratings anchored to fixed reference players spanning the strength range. An unbounded comparison-derived scale, so it carries no declared bounds and does not enter the score-level capability battery.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 1274 |
| Gemini 3.6 Flash | 1205 |
| GPT-5.4 | 1198 |
| Gemini 3.1 Pro Preview | 1178 |
| GPT-5.1 | 1175 |
| Gemini 3.5 Flash | 1084 |
| Sonnet 5 | 1066 |
| Claude Fable 5 | 1053 |
| o3 | 981 |
| GPT-5 | 959 |
| Opus 4.8 | 957 |
| GPT-5.2 | 943 |
| GPT-5.4 Nano | 893 |
| Claude Opus 5 | 819 |
| GPT-5 Nano | 811 |
| Gemini 3 Flash Preview | 794 |
| GPT-5 Mini | 767 |
| GPT-5.4 Mini | 753 |
| Gemini 3.1 Flash-Lite | 737 |
| Opus 4.6 | 736 |
| Gemini 3.5 Flash-Lite | 690 |
| o3-mini | 561 |
| Gemini 3.1 Flash-Lite Preview | 545 |
| Opus 4.7 | 530 |
| Sonnet 4.6 | 514 |
| O4 Mini | 506 |
| Opus 4.5 | 481 |
| Opus 4.1 | 459 |
| Opus 4 | 443 |
| Sonnet 4.5 | 389 |
| Sonnet 4 | 375 |
| Gemini 2.0 Flash | 348 |
| GPT-4 Turbo | 339 |
| GPT-4 0613 | 333 |
| GPT-4.1 | 314 |
| Haiku 4.5 | 313 |
| GPT-4o (2024-08-06) | 286 |
| GPT-4.1 Mini | 273 |
| GPT-4o (2024-11-20) | 261 |
| Gemini 2.0 Flash-Lite | 168 |