Atlas

Models

← All models

Qwen2 VL 72B Instruct

Alibaba · 2024-09-19 · 73.4B parameters

Qwen2-VL-72B-Instruct is Alibaba's 72-billion-parameter instruction-tuned vision-language model from the Qwen2-VL series, released in late 2024. It processes images at arbitrary resolutions via a dynamic visual token scheme, understands videos longer than 20 minutes, and supports multilingual text within images across European languages, Japanese, Korean, Arabic, and more. At release it achieved performance comparable to GPT-4o and Claude 3.5 Sonnet on multimodal benchmarks and supports autonomous device-control tasks.

Benchmark scores

BenchmarkScore
BALROG12.8
BBH (Open LLM Leaderboard v2)69.5
Capability125.8
CharXiv Reasoning43.0
CharXiv-D81.3
GPQA (Open LLM Leaderboard v2)38.8
Humanity's Last Exam (Preview)4.7
Humanity's Last Exam Text Only (Preview)4.7
IFEval (Open LLM Leaderboard v2)59.8
MATH Level 5 (Open LLM Leaderboard v2)34.4
MATH-Vision25.9
MathVista (mini)70.5
MMLU-Pro (Open LLM Leaderboard v2)57.2
MuSR (Open LLM Leaderboard v2)44.9
ScreenSpot-Pro1.0
Video-MME71.2
VISTA28.6
ZeroBench (Main questions, pass@1)0.0
ZeroBench (Main questions, pass@5)2.0
ZeroBench (Subquestions, pass@1)13.0
ZeroBench Main Questions pass^50.0
Loading Atlas data…