Atlas

Benchmarks

← All benchmarks

Blueprint-Bench 2

Multimodal · 2026-05-04

Blueprint-Bench 2 evaluates spatial reasoning by asking persistent agents to reconstruct floor plans for 50 apartments from approximately 20 photographs each. Published scores normalize connectivity similarity so the random baseline is zero and perfect reconstruction is one.

Top models (higher is better)

ModelScore
Claude Fable 50.4
GPT-5.50.4
GPT-5.6 Sol0.3
Gemini 3.5 Flash0.3
Gemini 3.6 Flash0.3
GPT-5.6 Terra0.3
Claude Opus 50.3
Kimi K30.3
Grok 4.50.3
GPT-5.40.3
Gemini 3.1 Pro Preview0.3
Opus 4.70.2
GPT-5.6 Luna0.2
Opus 4.80.1
Sonnet 4.60.1
Kimi K2.60.0
Haiku 4.50.0
Gemini 3 Flash Preview0.0
Grok 4.200.0
Grok 4.30.0
Loading Atlas data…