PerceptionBench
Multimodal · 2026-07-16
PerceptionBench evaluates atomic visual perception independently of reasoning and external knowledge using 3,000 verified, open-ended image questions. Its overall accuracy aggregates ten failure-derived capabilities, with short reference answers graded by GPT-oss-120B under a unified no-tools protocol.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol | 59.7 |
| Kimi K3 | 58.5 |
| Claude Fable 5 | 57.2 |
| Gemini 3.1 Pro Preview | 56.2 |
| GPT-5.5 | 55.8 |
| Doubao Seed 2.1 Pro | 55.0 |
| Gemini 3.5 Flash | 52.0 |
| Qwen3.7-Plus | 51.1 |
| Qwen3.5 397B A17B | 47.5 |
| Opus 4.8 | 47.2 |
| Kimi K2.6 | 42.6 |
| Grok 4.5 | 41.0 |
| Gemma 4 31B IT | 40.7 |
| GLM 5V Turbo | 39.6 |
| MiniMax M3 | 33.1 |
| GLM-4.6V | 32.5 |