Atlas

Benchmarks

← All benchmarks

MMMU Validation (OpenVLM)

Multimodal

The validation split of MMMU, college-level multimodal questions across six disciplines mixing multiple-choice and open-ended answers. Because the answer formats differ across items, no single guessing floor is well defined. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro74.7
InternVL3-78B72.2
GPT-4.572.1
InternVL2.5-78B70.0
Gemini 2.0 Flash69.9
InternVL3-38B69.7
Gemini 1.5 Pro 00268.6
Grok 2 Vision 121267.1
InternVL3-14B64.8
Qwen VL Max 080964.6
InternVL3-8B62.2
NVLM-D-72B60.8
Gemini 1.5 Flash 00260.7
Gemini 1.5 Pro60.6
Llama 3.2 90B Vision Instruct60.3
Gemini 1.5 Flash58.2
Kimi-VL-A3B-Instruct57.8
LLaVA-OneVision 72B56.6
InternVL2-40B55.2
Aria54.0
Qwen VL Max (rolling alias)52.0
Pixtral 12B51.1
MiniCPM-o 2.650.9
MiniCPM-V 2.649.8
Molmo 7B-D49.1
Gemini 1.0 Pro49.0
InternVL3-2B48.7
Llama 3.2 11B Vision Instruct48.0
InternVL-Chat-V1-546.8
InternVL3-1B43.2
Qwen-VL-Chat37.0
LLaVA v1.5 7B35.7
Loading Atlas data…