MathVista (OpenVLM)
Multimodal
MathVista's testmini split, mathematical reasoning over figures, charts, and diagrams, mixing multiple-choice and free-form numeric answers. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 2.5 Pro | 80.9 |
| InternVL3-78B | 79.0 |
| InternVL3-38B | 76.3 |
| InternVL3-14B | 74.4 |
| MiniCPM-o 2.6 | 73.3 |
| InternVL2.5-78B | 70.6 |
| GPT-4.5 | 70.5 |
| InternVL3-8B | 70.5 |
| Gemini 2.0 Flash | 70.4 |
| LLaVA-OneVision 72B | 68.4 |
| Qwen VL Max 0809 | 68.3 |
| Gemini 1.5 Pro 002 | 67.9 |
| Grok 2 Vision 1212 | 66.6 |
| Kimi-VL-A3B-Instruct | 66.0 |
| InternVL2-40B | 64.3 |
| NVLM-D-72B | 64.0 |
| Gemini 1.5 Flash 002 | 63.7 |
| Aria | 62.4 |
| MiniCPM-V 2.6 | 60.8 |
| Gemini 1.5 Pro | 58.3 |
| Llama 3.2 90B Vision Instruct | 58.2 |
| InternVL3-2B | 57.6 |
| Pixtral 12B | 56.3 |
| InternVL-Chat-V1-5 | 56.1 |
| Gemini 1.5 Flash | 51.3 |
| Molmo 7B-D | 48.7 |
| Llama 3.2 11B Vision Instruct | 47.7 |
| InternVL3-1B | 46.9 |
| Gemini 1.0 Pro | 46.5 |
| Qwen VL Max (rolling alias) | 43.6 |
| Qwen-VL-Chat | 35.3 |
| LLaVA v1.5 7B | 25.5 |