Atlas

Benchmarks

← All benchmarks

MMStar (OpenVLM)

Multimodal

MMStar is a curated multiple-choice benchmark of vision-indispensable samples, filtered so that questions cannot be answered from text alone. Accuracy over the whole set. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro73.6
InternVL3-78B73.4
InternVL3-38B72.6
InternVL2.5-78B69.5
Gemini 2.0 Flash69.4
GPT-4.569.3
Qwen VL Max 080969.2
InternVL3-14B68.9
InternVL3-8B68.7
Gemini 1.5 Pro 00267.1
LLaVA-OneVision 72B65.8
Grok 2 Vision 121265.2
InternVL2-40B64.7
Gemini 1.5 Flash 00264.4
NVLM-D-72B63.7
MiniCPM-o 2.663.3
Kimi-VL-A3B-Instruct62.0
InternVL3-2B61.1
Aria60.7
Gemini 1.5 Pro59.1
MiniCPM-V 2.657.5
InternVL-Chat-V1-557.1
Molmo 7B-D56.1
Gemini 1.5 Flash55.8
Llama 3.2 90B Vision Instruct55.3
Pixtral 12B54.5
InternVL3-1B52.3
Llama 3.2 11B Vision Instruct49.8
Qwen VL Max (rolling alias)49.5
Gemini 1.0 Pro38.6
Qwen-VL-Chat34.5
LLaVA v1.5 7B33.1
Loading Atlas data…