Atlas

Benchmarks

← All benchmarks

ChessBench

Games · 2026-03-22

ChessBench measures whether a language model can solve 150 fixed Lichess tactics from FEN and an ASCII board. Models must return the exact expected lowercase UCI move line; the headline strict accuracy counts malformed, missing, and incorrect lines as failures and averages 30 puzzles from each of five tactical tracks.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash61.3
Grok 4.560.7
Grok 4.1 Fast58.7
Grok Build 0.158.7
Gemini 3.1 Pro Preview55.3
Grok 4.2054.7
Gemini 3 Flash Preview48.7
Grok 4.348.7
Claude Fable 546.0
Gemini 3.1 Flash Image Preview46.0
GPT-5.6 Sol42.7
GPT-5.6 Terra40.7
GPT-5.432.7
GPT-5.529.3
Qwen3.6 Plus (2026-04-02)28.0
Opus 4.824.0
GPT-5.6 Luna Pro24.0
Qwen3.6 Plus Preview22.0
Opus 4.718.7
Sonnet 518.7
Opus 4.616.7
GLM-5.213.3
GLM-5.112.7
GLM-512.0
Sonnet 4.610.7
Haiku 4.58.7
Gemini 2.5 Pro5.3
Loading Atlas data…