SnorkelSpatial
General QA · 2025-10-24
SnorkelSpatial is a text-only spatial-reasoning benchmark from Snorkel AI. It evaluates LLMs on 330 procedurally generated and programmatically verified 2D grid-world problems requiring models to track board and particle movements, rotations, positions, orientations, and relative spatial relations. Scores are reported as accuracy@1.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.4 | 99.0 |
| Grok 4 Fast | 84.8 |
| o3 | 76.7 |
| GPT-5 | 73.9 |
| gpt-oss-120b | 52.7 |
| GPT-5 Mini | 45.5 |
| Opus 4.1 | 45.1 |
| Magistral Medium 1.2 | 44.2 |
| Opus 4 | 40.3 |
| o3-mini | 37.9 |
| Sonnet 4 | 33.3 |
| GPT-5 Nano | 26.7 |
| Claude 3.7 Sonnet | 21.5 |
| Gemini 2.5 Flash | 18.8 |
| Llama 4 Scout | 15.4 |
| Gemini 2.5 Pro | 15.2 |
| GPT-5 Chat (2025-08-07) | 14.8 |
| Mistral Large 1.0 | 14.8 |
| O4 Mini | 14.8 |
| GPT-4.1 | 14.6 |
| Llama 3.3 70B Instruct | 14.6 |
| Mistral Medium 3.1 | 14.6 |
| Nova Micro | 14.6 |
| C4AI Command R+ | 14.2 |
| Nova Premier | 14.2 |
| Qwen3-235B-A22B | 13.9 |
| Codestral 25.01 | 13.6 |
| Nova Lite | 13.3 |
| Grok 3 | 12.7 |
| Magistral Medium 1.0 | 12.4 |
| Llama 4 Maverick | 12.1 |
| Nova Pro | 12.1 |
| Command R | 11.8 |