WorldVQA Correct Given Attempted
Multimodal · 2026-01-28
WorldVQA Correct Given Attempted measures correctness among answers the model chose to provide across the first eight categories, isolating precision when the model commits to an answer.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3 Pro Preview | 47.7 |
| Kimi K2.5 | 47.3 |
| Opus 4.5 | 38.1 |
| Gemini 2.5 Pro | 36.9 |
| Seed1.5-VL | 35.5 |
| GPT-5.2 | 29.5 |
| GPT-5.1 | 29.3 |
| GPT-4o | 24.4 |
| Qwen3-VL-235B-A22B-Instruct | 23.5 |
| Sonnet 4.5 | 21.8 |
| Grok 4.1 Fast | 21.1 |
| Grok 4 Fast | 19.0 |
| GLM-4.6V | 19.0 |
| Qwen3 VL 32B Instruct | 17.7 |
| GLM 4.6V Flash | 14.8 |