Atlas

Benchmarks

← All benchmarks

MM-Vet (OpenVLM)

Multimodal

MM-Vet evaluates integrated multimodal capabilities with open-ended answers graded by an LLM judge on a partial-credit scale, so the score is a mean of per-item grades rather than a pass rate. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro83.3
InternVL3-8B82.8
InternVL3-38B81.1
InternVL3-78B80.7
InternVL3-14B80.5
GPT-4.575.3
Gemini 1.5 Pro 00274.6
Gemini 2.0 Flash73.6
Qwen VL Max 080972.3
Gemini 1.5 Flash 00269.7
MiniCPM-o 2.667.2
InternVL3-2B67.0
Kimi-VL-A3B-Instruct66.1
Grok 2 Vision 121265.2
Llama 3.2 90B Vision Instruct64.1
Gemini 1.5 Pro64.0
Gemini 1.5 Flash63.2
InternVL2-40B61.8
Qwen VL Max (rolling alias)61.8
LLaVA-OneVision 72B60.6
MiniCPM-V 2.660.0
NVLM-D-72B58.9
InternVL3-1B58.7
Gemini 1.0 Pro58.6
Pixtral 12B58.5
Llama 3.2 11B Vision Instruct57.6
Aria56.9
InternVL-Chat-V1-555.4
Qwen-VL-Chat47.3
Molmo 7B-D41.5
LLaVA v1.5 7B32.9
InternVL2.5-78B16.9
Loading Atlas data…