Atlas

Benchmarks

← All benchmarks

Humanity's Sixth Sense

Multimodal · 2026-10-07

Scale AI and Elorian's intuitive visual reasoning benchmark: a frozen release of 522 human-authored tasks (288 image and 234 silent-video tasks), across four domains and eleven subdomains. Pass@1 averages three attempts per task; Claude Opus 5 judges against reference answers and 723 rubric criteria. Published confidence intervals use a task-cluster bootstrap. Released October 7, 2026.

Top models (higher is better)

ModelScore
GPT-6 Astra53.6
GPT-6.1 Sol46.6
Claude Opus 5.544.6
Gemini 3.8 Flash41.6
Claude Fable 5.140.8
Gemini 3.7 Flash39.8
Muse Spark 1.337.4
Claude Fable 534.5
Gemini 3.5 Flash32.6
Gemini 3.6 Flash31.9
GPT-6 Sol31.2
Qwen3.8 Flash30.9
Claude Opus 530.5
GPT-5.6 Sol30.0
GLM-5.3 Flash26.8
Kimi K325.5
MiMo-V2.6-Flash25.4
MiMo-V2.6-Pro25.2
Sonnet 524.8
Muse Spark 1.224.5
Opus 4.823.5
GPT-5.6 Terra22.7
GPT-5.6 Luna21.6
GPT-6 Luna21.0
Loading Atlas data…