Atlas

Benchmarks

← All benchmarks

MathVista (OpenVLM)

Multimodal

MathVista's testmini split, mathematical reasoning over figures, charts, and diagrams, mixing multiple-choice and free-form numeric answers. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro80.9
InternVL3-78B79.0
InternVL3-38B76.3
InternVL3-14B74.4
MiniCPM-o 2.673.3
InternVL2.5-78B70.6
GPT-4.570.5
InternVL3-8B70.5
Gemini 2.0 Flash70.4
LLaVA-OneVision 72B68.4
Qwen VL Max 080968.3
Gemini 1.5 Pro 00267.9
Grok 2 Vision 121266.6
Kimi-VL-A3B-Instruct66.0
InternVL2-40B64.3
NVLM-D-72B64.0
Gemini 1.5 Flash 00263.7
Aria62.4
MiniCPM-V 2.660.8
Gemini 1.5 Pro58.3
Llama 3.2 90B Vision Instruct58.2
InternVL3-2B57.6
Pixtral 12B56.3
InternVL-Chat-V1-556.1
Gemini 1.5 Flash51.3
Molmo 7B-D48.7
Llama 3.2 11B Vision Instruct47.7
InternVL3-1B46.9
Gemini 1.0 Pro46.5
Qwen VL Max (rolling alias)43.6
Qwen-VL-Chat35.3
LLaVA v1.5 7B25.5
Loading Atlas data…