WordleBench
Games · 2026-02-14
WordleBench evaluates multi-turn Wordle solving on the same 100 fixed historical solution words for every model. After each five-letter guess, the model receives green/yellow/black feedback and continues until it solves the puzzle or exhausts the game. The published success rate is the percentage solved in fewer than six guesses; this follows the official aggregation exactly, which excludes sixth-guess solves.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5 | 99.0 |
| Sonnet 4.6 | 98.0 |
| GPT-5.4 | 98.0 |
| Claude Fable 5 | 97.0 |
| GPT-5.5 | 97.0 |
| GPT-5.6 Sol | 97.0 |
| Sonnet 5 | 96.0 |
| GPT-5.6 Terra Pro | 96.0 |
| Qwen3.7-Plus | 96.0 |
| Opus 4.6 | 95.0 |
| GPT-5.1 | 95.0 |
| Grok 4.20 | 95.0 |
| Opus 4.8 | 94.0 |
| Gemini 3 Pro Preview | 94.0 |
| GPT-5 Mini | 94.0 |
| GPT-5.6 Luna | 94.0 |
| GPT-5.6 Sol Pro | 94.0 |
| Opus 4.5 | 93.0 |
| Gemini 3.1 Flash-Lite Preview | 93.0 |
| Qwen3.7-Max | 92.0 |
| Grok 4.5 | 92.0 |
| GPT-5.6 Luna Pro | 91.0 |
| GPT-5.6 Terra | 91.0 |
| Qwen3.6 Plus (2026-04-02) | 90.0 |
| GPT-5.4 Mini | 89.0 |
| Grok 4.1 Fast | 89.0 |
| Opus 4.7 | 87.0 |
| Sonnet 4.5 | 87.0 |
| Gemma 4 31B IT | 87.0 |
| GPT-5.2 | 87.0 |
| Kimi K2.5 | 86.0 |
| GPT-5.4 Nano | 86.0 |
| Qwen3-Max-Preview (Thinking Mode) | 86.0 |
| GPT-5 Nano | 85.0 |
| Grok 4.3 | 85.0 |
| MiMo-V2-Flash | 85.0 |
| Mistral Medium 3.5 | 84.0 |
| Qwen3-Next-80B-A3B-Thinking | 83.0 |
| Kimi K2.6 | 81.0 |
| MiMo-V2-Pro | 81.0 |