Atlas

Benchmarks

← All benchmarks

Eval Connections — Guess Accuracy

Games · 2025-07-31

Eval Connections Guess Accuracy is the percentage of submitted four-word guesses that match one of the target groups across the same 20-game canonical battery. The denominator varies with the number of guesses each model makes.

Top models (higher is better)

ModelScore
Claude Fable 5100.0
Sonnet 4.6100.0
Gemini 3.1 Pro Preview100.0
GPT-5.2 Pro100.0
GPT-5.6 Sol Pro100.0
OpenRouter Fusion100.0
Qwen3.7-Plus100.0
Sonnet 598.8
Gemini 3 Pro Preview98.8
Opus 4.797.6
GPT-5.6 Terra97.6
GPT-5.6 Terra Pro97.6
Grok 497.6
Grok 4.2097.6
o397.6
Opus 4.697.1
GLM-5.196.4
GPT-5.6 Sol96.4
DeepSeek-V4-Pro95.2
Gemini 3.5 Flash95.2
Gemini 3 Flash Preview95.2
GPT-5.3-Codex95.2
Grok 4 Fast94.1
GLM-5.293.0
Opus 4.592.9
GPT-5.3 Instant90.7
Kimi K2.589.9
Qwen3.6 Plus (2026-04-02)89.9
Grok 4.389.7
Qwen3 Max Thinking (2026-01-23)89.7
MiniMax M388.4
GLM-4.787.5
DeepSeek-V4-Flash86.5
GLM-586.5
GPT-5.6 Luna Pro85.7
Step 3.7 Flash85.6
Gemini 2.5 Pro84.3
Kimi K2.683.7
GPT-5.6 Luna83.1
Kimi K2.7 Code82.8
Loading Atlas data…