WorldVQA ForceAnswer
Multimodal · 2026-07-16
WorldVQA accuracy under the later ForceAnswer protocol, which requires the model to provide an answer. It is kept separate from the original abstention-aware WorldVQA metrics because forced answering changes the measured behavior.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 56.7 |
| Kimi K3 | 51.0 |
| GPT-5.6 Sol | 41.8 |
| Opus 4.8 | 39.1 |
| GPT-5.5 | 38.5 |