Atlas

Benchmarks

← All benchmarks

SnorkelSpatial

General QA · 2025-10-24

SnorkelSpatial is a text-only spatial-reasoning benchmark from Snorkel AI. It evaluates LLMs on 330 procedurally generated and programmatically verified 2D grid-world problems requiring models to track board and particle movements, rotations, positions, orientations, and relative spatial relations. Scores are reported as accuracy@1.

Top models (higher is better)

ModelScore
GPT-5.499.0
Grok 4 Fast84.8
o376.7
GPT-573.9
gpt-oss-120b52.7
GPT-5 Mini45.5
Opus 4.145.1
Magistral Medium 1.244.2
Opus 440.3
o3-mini37.9
Sonnet 433.3
GPT-5 Nano26.7
Claude 3.7 Sonnet21.5
Gemini 2.5 Flash18.8
Llama 4 Scout15.4
Gemini 2.5 Pro15.2
GPT-5 Chat (2025-08-07)14.8
Mistral Large 1.014.8
O4 Mini14.8
GPT-4.114.6
Llama 3.3 70B Instruct14.6
Mistral Medium 3.114.6
Nova Micro14.6
C4AI Command R+14.2
Nova Premier14.2
Qwen3-235B-A22B13.9
Codestral 25.0113.6
Nova Lite13.3
Grok 312.7
Magistral Medium 1.012.4
Llama 4 Maverick12.1
Nova Pro12.1
Command R11.8
Loading Atlas data…