Atlas

Benchmarks

← All benchmarks

Video-MME

Multimodal · 2024-05-31

Video-MME evaluates multimodal language models on multiple-choice video question answering across diverse domains, durations, and modalities. Rows in this dataset use the overall no-subtitles metric unless a source note states otherwise.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro86.9
Gemini 2.5 Pro Preview 05-0684.8
Gemini 1.5 Pro75.0
Gemini 1.5 Pro 00175.0
Qwen2.5-VL-72B-Instruct73.5
InternVL2.5-78B72.1
GPT-4o (2024-08-06)71.9
GPT-4o (2024-11-20)71.9
Qwen2 VL 72B Instruct71.2
Gemini 1.5 Flash70.3
Aria67.6
LLaVA-OneVision 72B66.3
GPT-4o Mini64.8
MiniCPM-o 2.663.9
Qwen2 VL 7B Instruct63.9
InternVL2-40B61.2
MiniCPM-V 2.660.9
Claude 3.5 Sonnet (June 2024)60.0
Claude 3.5 Sonnet (Oct 2024)60.0
GPT-4 Turbo with Vision59.9
Qwen VL Max (rolling alias)51.3
InternVL-Chat-V1-550.7
Qwen-VL-Chat41.1
Loading Atlas data…