HallusionBench (OpenVLM)
Multimodal
HallusionBench probes visual illusion and hallucination with paired yes/no questions. The leaderboard's headline figure combines the question-level, figure-level, and all-question accuracies that it also reports separately, so it is a composite rather than a single item proportion. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 2.5 Pro | 64.1 |
| GPT-4.5 | 60.0 |
| Qwen VL Max 0809 | 59.2 |
| InternVL3-78B | 59.1 |
| InternVL3-38B | 58.4 |
| Gemini 2.0 Flash | 58.0 |
| InternVL2.5-78B | 57.4 |
| InternVL2-40B | 56.5 |
| Gemini 1.5 Pro 002 | 55.9 |
| InternVL3-14B | 55.9 |
| Gemini 1.5 Flash 002 | 53.9 |
| Grok 2 Vision 1212 | 51.7 |
| MiniCPM-o 2.6 | 51.1 |
| NVLM-D-72B | 49.7 |
| InternVL3-8B | 49.0 |
| Gemini 1.5 Flash | 48.5 |
| Kimi-VL-A3B-Instruct | 48.4 |
| MiniCPM-V 2.6 | 48.1 |
| LLaVA-OneVision 72B | 47.9 |
| InternVL-Chat-V1-5 | 47.4 |
| Pixtral 12B | 47.0 |
| Molmo 7B-D | 46.4 |
| Gemini 1.0 Pro | 45.7 |
| Gemini 1.5 Pro | 45.6 |
| Aria | 44.8 |
| Llama 3.2 90B Vision Instruct | 44.1 |
| InternVL3-2B | 41.9 |
| Qwen VL Max (rolling alias) | 41.2 |
| Llama 3.2 11B Vision Instruct | 40.3 |
| InternVL3-1B | 37.2 |
| Qwen-VL-Chat | 36.8 |
| LLaVA v1.5 7B | 27.6 |