Atlas

Benchmarks

← All benchmarks

WorldVQA Not Attempted

Multimodal · 2026-01-28

Percentage of WorldVQA questions in the first eight categories that the model did not attempt. Lower is better when accuracy and answer quality are held fixed.

Top models (lower is better)

ModelScore
Qwen3-VL-235B-A22B-Instruct0.0
Qwen3 VL 32B Instruct0.0
GLM-4.6V0.0
Gemini 2.5 Pro0.1
Grok 4.1 Fast0.1
GLM 4.6V Flash0.1
Grok 4 Fast0.2
Gemini 3 Pro Preview0.6
Seed1.5-VL1.6
Kimi K2.52.1
Opus 4.53.4
GPT-5.25.4
Sonnet 4.58.0
GPT-4o9.1
GPT-5.116.3
Loading Atlas data…