Atlas

Benchmarks

← All benchmarks

ChessBench Accuracy

Games · 2026-07-28

ChessBench (chessbench.ai) has language models play full chess games against reference opponents and scores the moves; it is unrelated to the chess-bench.com puzzle benchmark carried as `chessbench`. Accuracy scores move quality against Stockfish as a weighted mean of branching, converting, and resisting accuracy (weights 10, 9, 1). Each part is scaled so that the random-move baseline is 0 and perfect play is 1, so the metric has a known ceiling but no floor: play worse than random scores below zero, and observed values reach -1.2. min_score is therefore left undeclared, which keeps this out of the score-level capability battery.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview79.2
o378.1
GPT-5.277.4
Gemini 3.5 Flash75.4
Claude Fable 575.1
GPT-5.475.0
GPT-574.4
Gemini 3 Flash Preview73.4
Sonnet 571.7
Gemini 3.6 Flash71.7
GPT-5.171.6
Opus 4.870.8
GPT-5.570.5
GPT-5 Nano67.6
GPT-5 Mini67.4
GPT-5.4 Nano67.1
Opus 4.765.8
O4 Mini65.3
Gemini 3.1 Flash-Lite Preview64.8
Gemini 3.1 Flash-Lite64.6
Claude Opus 563.5
Gemini 3.5 Flash-Lite59.8
Opus 4.555.6
GPT-5.4 Mini54.2
o3-mini52.1
Opus 4.650.7
Sonnet 4.650.2
GPT-4.148.7
Sonnet 446.0
Sonnet 4.544.7
Opus 444.2
GPT-4 061343.2
Opus 4.141.1
Haiku 4.539.3
Gemini 2.0 Flash38.6
GPT-4 Turbo38.2
GPT-4o (2024-08-06)30.1
GPT-4.1 Nano22.5
GPT-4.1 Mini15.0
GPT-4o (2024-11-20)14.3
Loading Atlas data…