Atlas

Benchmarks

← All benchmarks

WordleBench

Games · 2026-02-14

WordleBench evaluates multi-turn Wordle solving on the same 100 fixed historical solution words for every model. After each five-letter guess, the model receives green/yellow/black feedback and continues until it solves the puzzle or exhausts the game. The published success rate is the percentage solved in fewer than six guesses; this follows the official aggregation exactly, which excludes sixth-guess solves.

Top models (higher is better)

ModelScore
GPT-599.0
Sonnet 4.698.0
GPT-5.498.0
Claude Fable 597.0
GPT-5.597.0
GPT-5.6 Sol97.0
Sonnet 596.0
GPT-5.6 Terra Pro96.0
Qwen3.7-Plus96.0
Opus 4.695.0
GPT-5.195.0
Grok 4.2095.0
Opus 4.894.0
Gemini 3 Pro Preview94.0
GPT-5 Mini94.0
GPT-5.6 Luna94.0
GPT-5.6 Sol Pro94.0
Opus 4.593.0
Gemini 3.1 Flash-Lite Preview93.0
Qwen3.7-Max92.0
Grok 4.592.0
GPT-5.6 Luna Pro91.0
GPT-5.6 Terra91.0
Qwen3.6 Plus (2026-04-02)90.0
GPT-5.4 Mini89.0
Grok 4.1 Fast89.0
Opus 4.787.0
Sonnet 4.587.0
Gemma 4 31B IT87.0
GPT-5.287.0
Kimi K2.586.0
GPT-5.4 Nano86.0
Qwen3-Max-Preview (Thinking Mode)86.0
GPT-5 Nano85.0
Grok 4.385.0
MiMo-V2-Flash85.0
Mistral Medium 3.584.0
Qwen3-Next-80B-A3B-Thinking83.0
Kimi K2.681.0
MiMo-V2-Pro81.0
Loading Atlas data…