Atlas

Benchmarks

← All benchmarks

GraphWalks BFS - 1M

General QA · 2025-04-12

GraphWalks is OpenAI's long-context multi-hop reasoning benchmark: the prompt is filled with a directed graph of hexadecimal-hash node ids and the model must return every node at exactly a given breadth-first-search depth from a starting node. This row is the 1M-context BFS bin, the 100 BFS problems at roughly 1,024k tokens in the dataset's graphwalks_256k_to_1mil split, scored by F1 over the returned node set so each item earns partial credit. Higher is better.

Top models (higher is better)

ModelScore
Claude Mythos 579.4
GPT-5.6 Sol77.1
Claude Mythos Preview74.3
GPT-5.6 Terra71.2
Opus 4.868.1
GPT-5.6 Luna51.2
GPT-5.545.4
Opus 4.740.3
Opus 4.616.3
GPT-5.49.4
Loading Atlas data…