Atlas

Benchmarks

← All benchmarks

ChessBench — Mate in 2

Games · 2026-03-22

ChessBench Mate in 2 contains 30 fixed Lichess positions in which the model must return the exact expected three-ply UCI mating line. The track is stratified evenly across 800–1200, 1200–1600, and 1600–2000 puzzle-rating buckets.

Top models (higher is better)

ModelScore
Grok Build 0.186.7
Grok 4.2083.3
Gemini 3.5 Flash76.7
Grok 4.1 Fast73.3
Grok 4.573.3
GPT-5.6 Sol66.7
Grok 4.363.3
Gemini 3.1 Pro Preview56.7
Gemini 3 Flash Preview53.3
GPT-5.6 Terra53.3
Claude Fable 543.3
GPT-5.443.3
Gemini 3.1 Flash Image Preview40.0
GPT-5.520.0
Opus 4.816.7
Sonnet 516.7
GPT-5.6 Luna Pro16.7
Qwen3.6 Plus (2026-04-02)16.7
Qwen3.6 Plus Preview16.7
Opus 4.610.0
Opus 4.76.7
Sonnet 4.63.3
GLM-5.13.3
GLM-5.23.3
Haiku 4.50.0
Gemini 2.5 Pro0.0
GLM-50.0
Loading Atlas data…