Atlas

Benchmarks

← All benchmarks

LisanBench Path Length

Games · 2025-05-30

LisanBench Path Length sums each model's valid edit-distance-one word-chain transitions across 50 starting words after averaging the three trials for each word. Chains must begin with the required word, use dictionary words, and never repeat a word.

Top models (higher is better)

ModelScore
Claude Opus 515428
Opus 4.714400
Opus 4.614206
Claude Fable 513680
Sonnet 4.612060
Kimi K310521
GPT-5.59845
Opus 4.89463
Gemini 3.1 Pro Preview6445
Sonnet 55724
Opus 4.55513
GPT-5.45335
DeepSeek-V3.2-Speciale5073
GPT-5.6 Sol4699
Gemini 3 Pro Preview4677
Gemini 3.5 Flash4123
Grok 44042
Grok 4.203856
GPT-5.23592
o33309
GPT-5.6 Terra3303
DeepSeek-V4-Pro3256
DeepSeek-V4-Flash2789
Grok 4 Fast2750
Sonnet 4.52665
Step 3.5 Flash2410
DeepSeek-V3.22348
GPT-52247
Sonnet 42236
GPT-5.6 Luna2105
Kimi K2 Thinking2002
Gemini 3 Flash Preview1996
Kimi K2.51852
GLM-5.21835
Grok 4.1 Fast1817
Doubao Seed 2.0 Pro1753
GPT-5 Mini1562
GPT-5.4 Nano1426
o3-mini1366
Qwen3.5 397B A17B1293
Loading Atlas data…