GraphWalks BFS - 256K
General QA · 2025-04-12
GraphWalks is OpenAI's long-context multi-hop reasoning benchmark: the prompt is filled with a directed graph of hexadecimal-hash node ids and the model must return every node at exactly a given breadth-first-search depth from a starting node. This row is the 256K-context BFS bin, the 100 BFS problems at roughly 256k tokens in the dataset's graphwalks_256k_to_1mil split, scored by F1 over the returned node set so each item earns partial credit. Higher is better.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos 5 | 91.1 |
| GPT-5.6 Sol | 90.7 |
| Opus 4.8 | 85.9 |
| Claude Mythos Preview | 85.7 |
| GPT-5.6 Luna | 81.3 |
| Opus 4.7 | 76.9 |
| GPT-5.6 Terra | 76.9 |
| Sonnet 4.6 | 74.5 |
| GPT-5.5 | 73.7 |
| GPT-5.4 | 62.5 |
| Opus 4.6 | 61.1 |