Atlas

Benchmarks

← All benchmarks

WorldVQA ForceAnswer

Multimodal · 2026-07-16

WorldVQA accuracy under the later ForceAnswer protocol, which requires the model to provide an answer. It is kept separate from the original abstention-aware WorldVQA metrics because forced answering changes the measured behavior.

Top models (higher is better)

ModelScore
Claude Fable 556.7
Kimi K351.0
GPT-5.6 Sol41.8
Opus 4.839.1
GPT-5.538.5
Loading Atlas data…