Atlas

Benchmarks

← All benchmarks

Eval Connections One-Shot with Traps

Games · 2026-07-22

Eval Connections one-shot board on 20 canonical puzzles with reviewed trap sets. One scored submission of all four groups: 0/1/2 points for that many correct groups, 3 for all four, plus 2 for a correct trap claim or correct N/A. Maximum 100 points. The current implementation permits two structural-format repairs without correctness feedback. Kept separate from classic multi-turn play; per-run inference settings are not published in the aggregate table.

Top models (higher is better)

ModelScore
GPT-6 Astra94.0
Grok 4.691.0
Muse Spark 1.190.0
GPT-6 Sol89.0
Gemini 3.8 Flash88.0
Claude Opus 5.586.0
Gemini 3.1 Pro Preview86.0
GPT-5.6 Terra86.0
Claude Fable 585.0
Grok 4.785.0
Gemini 3 Flash Preview84.0
Muse Spark 1.384.0
GPT-5.6 Sol Pro84.0
Kimi K384.0
Fugu Ultra v1.084.0
GPT-5.6 Sol83.0
Claude Fable 5.182.0
Grok 4.2082.0
Gemini 3.7 Flash80.0
DeepSeek V4 Pro 081380.0
GLM-5.3 Flash78.0
Muse Spark 1.277.0
GPT-5.6 Terra Pro77.0
DeepSeek-V4-Flash-073177.0
Grok 4.576.0
Qwen3.8 Flash75.0
Claude Opus 574.0
GPT-5.6 Luna74.0
GLM 5V Turbo74.0
DeepSeek V4.1 Flash73.0
GPT-5.6 Luna Pro72.0
Sonnet 571.0
Opus 4.570.0
Gemini 3.6 Flash69.0
MiMo-V2.6-Flash69.0
Gemini 3.5 Flash-Lite67.0
Gemini 3.5 Flash67.0
Inkling67.0
Qwen3.8 2.4T A95B66.0
Opus 4.765.0
Loading Atlas data…