Atlas

Benchmarks

← All benchmarks

PerceptionBench

Multimodal · 2026-07-16

PerceptionBench evaluates atomic visual perception independently of reasoning and external knowledge using 3,000 verified, open-ended image questions. Its overall accuracy aggregates ten failure-derived capabilities, with short reference answers graded by GPT-oss-120B under a unified no-tools protocol.

Top models (higher is better)

ModelScore
GPT-5.6 Sol59.7
Kimi K358.5
Claude Fable 557.2
Gemini 3.1 Pro Preview56.2
GPT-5.555.8
Doubao Seed 2.1 Pro55.0
Gemini 3.5 Flash52.0
Qwen3.7-Plus51.1
Qwen3.5 397B A17B47.5
Opus 4.847.2
Kimi K2.642.6
Grok 4.541.0
Gemma 4 31B IT40.7
GLM 5V Turbo39.6
MiniMax M333.1
GLM-4.6V32.5
Loading Atlas data…