Atlas

Benchmarks

← All benchmarks

GraphWalks BFS - 256K

General QA · 2025-04-12

GraphWalks is OpenAI's long-context multi-hop reasoning benchmark: the prompt is filled with a directed graph of hexadecimal-hash node ids and the model must return every node at exactly a given breadth-first-search depth from a starting node. This row is the 256K-context BFS bin, the 100 BFS problems at roughly 256k tokens in the dataset's graphwalks_256k_to_1mil split, scored by F1 over the returned node set so each item earns partial credit. Higher is better.

Top models (higher is better)

ModelScore
Claude Mythos 591.1
GPT-5.6 Sol90.7
Opus 4.885.9
Claude Mythos Preview85.7
GPT-5.6 Luna81.3
Opus 4.776.9
GPT-5.6 Terra76.9
Sonnet 4.674.5
GPT-5.573.7
GPT-5.462.5
Opus 4.661.1
Loading Atlas data…