Atlas

Benchmarks

← All benchmarks

HallusionBench (OpenVLM)

Multimodal

HallusionBench probes visual illusion and hallucination with paired yes/no questions. The leaderboard's headline figure combines the question-level, figure-level, and all-question accuracies that it also reports separately, so it is a composite rather than a single item proportion. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro64.1
GPT-4.560.0
Qwen VL Max 080959.2
InternVL3-78B59.1
InternVL3-38B58.4
Gemini 2.0 Flash58.0
InternVL2.5-78B57.4
InternVL2-40B56.5
Gemini 1.5 Pro 00255.9
InternVL3-14B55.9
Gemini 1.5 Flash 00253.9
Grok 2 Vision 121251.7
MiniCPM-o 2.651.1
NVLM-D-72B49.7
InternVL3-8B49.0
Gemini 1.5 Flash48.5
Kimi-VL-A3B-Instruct48.4
MiniCPM-V 2.648.1
LLaVA-OneVision 72B47.9
InternVL-Chat-V1-547.4
Pixtral 12B47.0
Molmo 7B-D46.4
Gemini 1.0 Pro45.7
Gemini 1.5 Pro45.6
Aria44.8
Llama 3.2 90B Vision Instruct44.1
InternVL3-2B41.9
Qwen VL Max (rolling alias)41.2
Llama 3.2 11B Vision Instruct40.3
InternVL3-1B37.2
Qwen-VL-Chat36.8
LLaVA v1.5 7B27.6
Loading Atlas data…