Atlas

Benchmarks

← All benchmarks

LisanBench Difficulty-Weighted

Games · 2025-05-30

LisanBench Difficulty-Weighted Score sums valid word-chain moves across 50 starting words after weighting each move by inverse branching factor and word rarity, then averaging the three trials for each word. It rewards paths through sparse, uncommon parts of the word graph rather than raw chain length alone.

Top models (higher is better)

ModelScore
Opus 4.75123
Claude Opus 55068
Claude Fable 54562
Opus 4.63526
Kimi K33526
GPT-5.53316
Sonnet 4.62944
GPT-5.42738
Opus 4.82694
Opus 4.52204
Gemini 3.1 Pro Preview1929
Grok 41778
Sonnet 51737
o31523
DeepSeek-V3.2-Speciale1511
Grok 4.201465
GPT-5.21459
GPT-51457
GPT-5.6 Sol1403
GPT-5.6 Terra1197
Gemini 3 Pro Preview1131
Gemini 3.5 Flash1128
Sonnet 4.51091
DeepSeek-V4-Flash1063
DeepSeek-V4-Pro1060
DeepSeek-V3.2925
Step 3.5 Flash811
Grok 4 Fast806
GPT-5 Mini759
GPT-5.6 Luna648
Kimi K2.5642
Kimi K2 Thinking633
GPT-5 Nano627
Grok 4.1 Fast604
Sonnet 4603
GLM-5.2592
Gemini 3 Flash Preview592
GPT-5.4 Mini591
GPT-5.4 Nano543
o3-mini518
Loading Atlas data…