Atlas

Benchmarks

← All benchmarks

Eval Connections

Games · 2025-07-31

Eval Connections measures the percentage of 20 fixed word-grouping games a model solves. Each game presents 16 shuffled words forming four groups of four; the model receives multi-turn correctness feedback and must recover all four groups before making four mistakes.

Top models (higher is better)

ModelScore
Claude Fable 5100.0
Opus 4.7100.0
Sonnet 4.6100.0
Sonnet 5100.0
DeepSeek-V4-Pro100.0
Fugu Ultra v1.0100.0
Gemini 3.1 Pro Preview100.0
Gemini 3.5 Flash100.0
Gemini 3 Flash Preview100.0
Gemini 3 Pro Preview100.0
GLM-5.1100.0
GLM-5.2100.0
GPT-5.2 Pro100.0
GPT-5.3-Codex100.0
GPT-5.4100.0
GPT-5.6 Sol100.0
GPT-5.6 Sol Pro100.0
GPT-5.6 Terra100.0
GPT-5.6 Terra Pro100.0
Grok 4100.0
Grok 4.20100.0
Grok 4 Fast100.0
Kimi K2.5100.0
o3100.0
OpenRouter Fusion100.0
Qwen3.6 Plus (2026-04-02)100.0
Qwen3.7-Plus100.0
Opus 4.595.0
GLM-4.795.0
GLM-595.0
GPT-5.3 Instant95.0
GPT-5.6 Luna Pro95.0
Grok 4.395.0
Grok 4.595.0
Kimi K2.695.0
Kimi K2.7 Code95.0
Qwen3 Max Thinking (2026-01-23)95.0
Step 3.7 Flash95.0
DeepSeek-V4-Flash90.0
Gemini 2.5 Pro90.0
Loading Atlas data…