MM-Vet (OpenVLM)
Multimodal
MM-Vet evaluates integrated multimodal capabilities with open-ended answers graded by an LLM judge on a partial-credit scale, so the score is a mean of per-item grades rather than a pass rate. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 2.5 Pro | 83.3 |
| InternVL3-8B | 82.8 |
| InternVL3-38B | 81.1 |
| InternVL3-78B | 80.7 |
| InternVL3-14B | 80.5 |
| GPT-4.5 | 75.3 |
| Gemini 1.5 Pro 002 | 74.6 |
| Gemini 2.0 Flash | 73.6 |
| Qwen VL Max 0809 | 72.3 |
| Gemini 1.5 Flash 002 | 69.7 |
| MiniCPM-o 2.6 | 67.2 |
| InternVL3-2B | 67.0 |
| Kimi-VL-A3B-Instruct | 66.1 |
| Grok 2 Vision 1212 | 65.2 |
| Llama 3.2 90B Vision Instruct | 64.1 |
| Gemini 1.5 Pro | 64.0 |
| Gemini 1.5 Flash | 63.2 |
| InternVL2-40B | 61.8 |
| Qwen VL Max (rolling alias) | 61.8 |
| LLaVA-OneVision 72B | 60.6 |
| MiniCPM-V 2.6 | 60.0 |
| NVLM-D-72B | 58.9 |
| InternVL3-1B | 58.7 |
| Gemini 1.0 Pro | 58.6 |
| Pixtral 12B | 58.5 |
| Llama 3.2 11B Vision Instruct | 57.6 |
| Aria | 56.9 |
| InternVL-Chat-V1-5 | 55.4 |
| Qwen-VL-Chat | 47.3 |
| Molmo 7B-D | 41.5 |
| LLaVA v1.5 7B | 32.9 |
| InternVL2.5-78B | 16.9 |