Atlas

Models

← All models

Qwen2 VL 7B Instruct

Alibaba · 2024-08-29 · 8.3B parameters

Qwen2-VL 7B Instruct is Alibaba's roughly 8.3-billion-parameter instruction-tuned vision-language model (about a 7.6B language model plus a 675M vision encoder) from the Qwen2-VL series, released in 2024. It combines a Qwen2 language backbone with a visual encoder supporting dynamic resolutions and long-context processing, enabling document analysis, OCR, chart understanding, temporal reasoning over video, and structured output generation.

Benchmark scores

BenchmarkScore
Arena Vision — Chinese (No Style Control)994
Arena Vision — Chinese (Style Controlled)1036
Arena Vision — English (No Style Control)1007
Arena Vision — English (Style Controlled)1041
BALROG3.7
BBH (Open LLM Leaderboard v2)54.6
Capability112.0
CharXiv Reasoning34.6
CharXiv-D58.0
GPQA (Open LLM Leaderboard v2)32.0
IFEval (Open LLM Leaderboard v2)46.0
MATH Level 5 (Open LLM Leaderboard v2)19.9
MATH-Vision16.3
MMLU-Pro (Open LLM Leaderboard v2)40.9
MuSR (Open LLM Leaderboard v2)43.8
Video-MME63.9
Vision Arena Overall1032
Vision Arena Overall No Style Control990
Loading Atlas data…