Atlas

Benchmarks

← All benchmarks

Furniture Assembly Benchmark

Multimodal · 2026-09-23

Epoch's 60 photographed IKEA assemblies across three products. Given the photo and assembly manual, an agent identifies whether a mistake exists and its step and description, using zoom and Python with an 80-step budget. Reports accuracy, with a model judge checking error descriptions.

Top models (higher is better)

ModelScore
Claude Opus 5.583.3
GPT-6 Astra80.0
Claude Fable 5.170.0
Claude Opus 560.8
GPT-6 Sol58.3
GPT-5.6 Sol56.7
GPT-5.6 Terra54.2
GPT-5.544.2
Opus 4.842.5
GPT-5.6 Luna42.5
Grok 4.640.0
GPT-5.238.3
GPT-5.437.5
Claude Fable 535.8
Kimi K334.2
Opus 4.733.3
Gemini 3.8 Flash31.7
Opus 4.528.3
Opus 4.628.3
GPT-6 Luna28.3
Gemini 3.1 Pro Preview26.7
Gemini 3.7 Flash26.7
Gemini 3.6 Flash23.3
Grok 4.522.5
Kimi K2.621.7
Grok 4.720.8
Qwen3.8 Max (0902)20.0
Loading Atlas data…