Atlas

Benchmarks

← All benchmarks

ChessBench Elo

Games · 2026-07-28

ChessBench (chessbench.ai) has language models play full chess games against reference opponents and scores the moves; it is unrelated to the chess-bench.com puzzle benchmark carried as `chessbench`. Monte Carlo Elo ratings anchored to fixed reference players spanning the strength range. An unbounded comparison-derived scale, so it carries no declared bounds and does not enter the score-level capability battery.

Top models (higher is better)

ModelScore
GPT-5.51274
Gemini 3.6 Flash1205
GPT-5.41198
Gemini 3.1 Pro Preview1178
GPT-5.11175
Gemini 3.5 Flash1084
Sonnet 51066
Claude Fable 51053
o3981
GPT-5959
Opus 4.8957
GPT-5.2943
GPT-5.4 Nano893
Claude Opus 5819
GPT-5 Nano811
Gemini 3 Flash Preview794
GPT-5 Mini767
GPT-5.4 Mini753
Gemini 3.1 Flash-Lite737
Opus 4.6736
Gemini 3.5 Flash-Lite690
o3-mini561
Gemini 3.1 Flash-Lite Preview545
Opus 4.7530
Sonnet 4.6514
O4 Mini506
Opus 4.5481
Opus 4.1459
Opus 4443
Sonnet 4.5389
Sonnet 4375
Gemini 2.0 Flash348
GPT-4 Turbo339
GPT-4 0613333
GPT-4.1314
Haiku 4.5313
GPT-4o (2024-08-06)286
GPT-4.1 Mini273
GPT-4o (2024-11-20)261
Gemini 2.0 Flash-Lite168
Loading Atlas data…