SnorkelGraph
Math · 2026-06-08
SnorkelGraph is a graph-reasoning benchmark from Snorkel AI with 200 mathematical and spatial graph QA problems. Final answers are canonicalized and checked by a programmatic graph validator, with scores reported as accuracy@1.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.4 | 84.5 |
| Grok 4 Fast | 75.0 |
| O4 Mini | 75.0 |
| GPT-5 Mini | 72.5 |
| GPT-5 | 72.0 |
| o3 | 71.5 |
| o3-mini | 71.0 |
| Opus 4 | 64.5 |
| Grok 3 | 64.0 |
| GPT-4.1 | 63.0 |
| GPT-5 Nano | 62.5 |
| Qwen3-235B-A22B | 61.5 |
| Grok 4 | 61.0 |
| Sonnet 4 | 58.0 |
| Gemini 2.5 Pro | 58.0 |
| Gemini 2.5 Flash | 55.0 |
| Magistral Medium 1.0 | 53.5 |
| Claude 3.7 Sonnet | 50.0 |
| Nova Premier | 34.5 |
| Llama 4 Maverick | 34.0 |
| Mistral Large 1.0 | 30.0 |
| Llama 3.3 Nemotron Super 49B V1 | 29.0 |
| Nova Pro | 28.0 |
| Llama 4 Scout | 26.0 |
| Codestral 25.01 | 24.5 |
| Llama 3.3 70B Instruct | 23.5 |
| Llama 3.1 Nemotron 70B Instruct HF | 22.5 |
| Llama 3.1 405B Instruct | 20.5 |
| Nova Lite | 19.0 |
| Nova Micro | 17.5 |
| C4AI Command R+ | 15.0 |