GraphWalks BFS - 1M
General QA · 2025-04-12
GraphWalks is OpenAI's long-context multi-hop reasoning benchmark: the prompt is filled with a directed graph of hexadecimal-hash node ids and the model must return every node at exactly a given breadth-first-search depth from a starting node. This row is the 1M-context BFS bin, the 100 BFS problems at roughly 1,024k tokens in the dataset's graphwalks_256k_to_1mil split, scored by F1 over the returned node set so each item earns partial credit. Higher is better.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos 5 | 79.4 |
| GPT-5.6 Sol | 77.1 |
| Claude Mythos Preview | 74.3 |
| GPT-5.6 Terra | 71.2 |
| Opus 4.8 | 68.1 |
| GPT-5.6 Luna | 51.2 |
| GPT-5.5 | 45.4 |
| Opus 4.7 | 40.3 |
| Opus 4.6 | 16.3 |
| GPT-5.4 | 9.4 |