Atlas

Benchmarks

← All benchmarks

WorldVQA Correct Given Attempted

Multimodal · 2026-01-28

WorldVQA Correct Given Attempted measures correctness among answers the model chose to provide across the first eight categories, isolating precision when the model commits to an answer.

Top models (higher is better)

ModelScore
Gemini 3 Pro Preview47.7
Kimi K2.547.3
Opus 4.538.1
Gemini 2.5 Pro36.9
Seed1.5-VL35.5
GPT-5.229.5
GPT-5.129.3
GPT-4o24.4
Qwen3-VL-235B-A22B-Instruct23.5
Sonnet 4.521.8
Grok 4.1 Fast21.1
Grok 4 Fast19.0
GLM-4.6V19.0
Qwen3 VL 32B Instruct17.7
GLM 4.6V Flash14.8
Loading Atlas data…