MMStar (OpenVLM)
Multimodal
MMStar is a curated multiple-choice benchmark of vision-indispensable samples, filtered so that questions cannot be answered from text alone. Accuracy over the whole set. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 2.5 Pro | 73.6 |
| InternVL3-78B | 73.4 |
| InternVL3-38B | 72.6 |
| InternVL2.5-78B | 69.5 |
| Gemini 2.0 Flash | 69.4 |
| GPT-4.5 | 69.3 |
| Qwen VL Max 0809 | 69.2 |
| InternVL3-14B | 68.9 |
| InternVL3-8B | 68.7 |
| Gemini 1.5 Pro 002 | 67.1 |
| LLaVA-OneVision 72B | 65.8 |
| Grok 2 Vision 1212 | 65.2 |
| InternVL2-40B | 64.7 |
| Gemini 1.5 Flash 002 | 64.4 |
| NVLM-D-72B | 63.7 |
| MiniCPM-o 2.6 | 63.3 |
| Kimi-VL-A3B-Instruct | 62.0 |
| InternVL3-2B | 61.1 |
| Aria | 60.7 |
| Gemini 1.5 Pro | 59.1 |
| MiniCPM-V 2.6 | 57.5 |
| InternVL-Chat-V1-5 | 57.1 |
| Molmo 7B-D | 56.1 |
| Gemini 1.5 Flash | 55.8 |
| Llama 3.2 90B Vision Instruct | 55.3 |
| Pixtral 12B | 54.5 |
| InternVL3-1B | 52.3 |
| Llama 3.2 11B Vision Instruct | 49.8 |
| Qwen VL Max (rolling alias) | 49.5 |
| Gemini 1.0 Pro | 38.6 |
| Qwen-VL-Chat | 34.5 |
| LLaVA v1.5 7B | 33.1 |