AI2D (OpenVLM)
Multimodal
AI2 Diagrams, grade-school science diagram comprehension scored as multiple-choice accuracy. Option counts vary across items, so no single chance floor is declared. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.
Top models (higher is better)
| Model | Score |
|---|---|
| InternVL3-78B | 89.8 |
| Gemini 2.5 Pro | 89.5 |
| InternVL2.5-78B | 89.1 |
| InternVL3-38B | 88.7 |
| Qwen VL Max 0809 | 88.1 |
| GPT-4.5 | 87.2 |
| InternVL2-40B | 86.8 |
| LLaVA-OneVision 72B | 86.2 |
| MiniCPM-o 2.6 | 86.1 |
| InternVL3-14B | 86.0 |
| InternVL3-8B | 85.1 |
| Kimi-VL-A3B-Instruct | 84.5 |
| Gemini 1.5 Flash 002 | 84.4 |
| Gemini 1.5 Pro 002 | 83.3 |
| Gemini 2.0 Flash | 83.1 |
| Aria | 82.7 |
| MiniCPM-V 2.6 | 82.1 |
| Grok 2 Vision 1212 | 81.2 |
| Molmo 7B-D | 81.0 |
| InternVL-Chat-V1-5 | 80.6 |
| NVLM-D-72B | 80.1 |
| Gemini 1.5 Pro | 79.1 |
| Pixtral 12B | 79.0 |
| InternVL3-2B | 78.6 |
| Gemini 1.5 Flash | 78.5 |
| Llama 3.2 11B Vision Instruct | 77.3 |
| Qwen VL Max (rolling alias) | 75.7 |
| Gemini 1.0 Pro | 72.9 |
| InternVL3-1B | 69.7 |
| Llama 3.2 90B Vision Instruct | 69.5 |
| Qwen-VL-Chat | 63.0 |
| LLaVA v1.5 7B | 55.5 |