Atlas

Benchmarks

← All benchmarks

AI2D (OpenVLM)

Multimodal

AI2 Diagrams, grade-school science diagram comprehension scored as multiple-choice accuracy. Option counts vary across items, so no single chance floor is declared. Scored by the OpenVLM Leaderboard, which runs every model through OpenCompass's VLMEvalKit under one fixed harness. That is a different measurement protocol from the lab-reported and Artificial Analysis runs of the same underlying benchmarks, so it forms its own factor group. No evaluation_item_count is declared: the leaderboard publishes scores to one decimal place, which is coarser than the item lattice, so the denominator cannot be confirmed from the numbers and a binomial noise model is not claimed.

Top models (higher is better)

ModelScore
InternVL3-78B89.8
Gemini 2.5 Pro89.5
InternVL2.5-78B89.1
InternVL3-38B88.7
Qwen VL Max 080988.1
GPT-4.587.2
InternVL2-40B86.8
LLaVA-OneVision 72B86.2
MiniCPM-o 2.686.1
InternVL3-14B86.0
InternVL3-8B85.1
Kimi-VL-A3B-Instruct84.5
Gemini 1.5 Flash 00284.4
Gemini 1.5 Pro 00283.3
Gemini 2.0 Flash83.1
Aria82.7
MiniCPM-V 2.682.1
Grok 2 Vision 121281.2
Molmo 7B-D81.0
InternVL-Chat-V1-580.6
NVLM-D-72B80.1
Gemini 1.5 Pro79.1
Pixtral 12B79.0
InternVL3-2B78.6
Gemini 1.5 Flash78.5
Llama 3.2 11B Vision Instruct77.3
Qwen VL Max (rolling alias)75.7
Gemini 1.0 Pro72.9
InternVL3-1B69.7
Llama 3.2 90B Vision Instruct69.5
Qwen-VL-Chat63.0
LLaVA v1.5 7B55.5
Loading Atlas data…