Video-MME
Multimodal · 2024-05-31
Video-MME evaluates multimodal language models on multiple-choice video question answering across diverse domains, durations, and modalities. Rows in this dataset use the overall no-subtitles metric unless a source note states otherwise.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 2.5 Pro | 86.9 |
| Gemini 2.5 Pro Preview 05-06 | 84.8 |