MMMU Validation (OpenVLM)
Multimodal
The validation split of MMMU, college-level multimodal questions across six disciplines mixing multiple-choice and open-ended answers. Because the answer formats differ across items, no single guessing floor is well defined. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 2.5 Pro | 74.7 |
| InternVL3-78B | 72.2 |
| GPT-4.5 | 72.1 |
| InternVL2.5-78B | 70.0 |
| Gemini 2.0 Flash | 69.9 |
| InternVL3-38B | 69.7 |
| Gemini 1.5 Pro 002 | 68.6 |
| Grok 2 Vision 1212 | 67.1 |
| InternVL3-14B | 64.8 |
| Qwen VL Max 0809 | 64.6 |
| InternVL3-8B | 62.2 |
| NVLM-D-72B | 60.8 |
| Gemini 1.5 Flash 002 | 60.7 |
| Gemini 1.5 Pro | 60.6 |
| Llama 3.2 90B Vision Instruct | 60.3 |
| Gemini 1.5 Flash | 58.2 |
| Kimi-VL-A3B-Instruct | 57.8 |
| LLaVA-OneVision 72B | 56.6 |
| InternVL2-40B | 55.2 |
| Aria | 54.0 |
| Qwen VL Max (rolling alias) | 52.0 |
| Pixtral 12B | 51.1 |
| MiniCPM-o 2.6 | 50.9 |
| MiniCPM-V 2.6 | 49.8 |
| Molmo 7B-D | 49.1 |
| Gemini 1.0 Pro | 49.0 |
| InternVL3-2B | 48.7 |
| Llama 3.2 11B Vision Instruct | 48.0 |
| InternVL-Chat-V1-5 | 46.8 |
| InternVL3-1B | 43.2 |
| Qwen-VL-Chat | 37.0 |
| LLaVA v1.5 7B | 35.7 |