WordleBench — Average Guesses on Successful Games
Games · 2026-02-14
WordleBench Average Guesses is the mean number of guesses only among games that the official aggregation counts as successful—that is, puzzles solved in fewer than six guesses. Because failures and sixth-guess solves are omitted, this conditional statistic can favor models that solve only easier words; Atlas retains it as a diagnostic rather than a capability factor.
Top models (lower is better)
| Model | Score |
|---|---|
| Trinity Large Thinking | 2.7 |
| DeepSeek-V4-Pro | 3.1 |
| Gemma 4 26B A4B IT | 3.4 |
| Gemini 3.5 Flash | 3.4 |
| Gemini 3 Flash Preview | 3.5 |
| GPT-5.6 Terra | 3.5 |
| Qwen3.6 35B A3B | 3.5 |
| GPT-5.6 Terra Pro | 3.5 |
| GPT-5.6 Sol Pro | 3.5 |
| Inkling | 3.5 |
| Qwen3.6 27B | 3.5 |
| GLM-5.1 | 3.5 |
| Claude Fable 5 | 3.5 |
| Opus 4.8 | 3.5 |
| DeepSeek-V4-Flash | 3.5 |
| GPT-5.6 Sol | 3.5 |
| Gemini 3 Pro Preview | 3.5 |
| Qwen3.7-Max | 3.5 |
| Opus 4.7 | 3.5 |
| GPT-5.6 Luna Pro | 3.5 |
| Qwen3.6 Plus (2026-04-02) | 3.6 |
| Gemma 4 31B IT | 3.6 |
| Qwen3.5 122B A10B | 3.6 |
| GPT-5.4 | 3.6 |
| GPT-5.5 | 3.6 |
| Kimi K3 | 3.6 |
| Opus 4.6 | 3.6 |
| GPT-5 | 3.6 |
| Kimi K2.5 | 3.6 |
| GLM-5 | 3.6 |
| GPT-5.1 | 3.6 |
| Gemini 2.5 Pro | 3.6 |
| Qwen3.5-27B | 3.6 |
| GPT-5 Mini | 3.6 |
| Opus 4.5 | 3.6 |
| Grok 4.5 | 3.6 |
| Sonnet 4.5 | 3.6 |
| Qwen3.5 397B A17B | 3.6 |
| Gemini 3.1 Pro Preview | 3.6 |
| MiMo-V2-Flash | 3.6 |