Qwen2 VL 72B Instruct
Alibaba · 2024-09-19 · 73.4B parameters
Qwen2-VL-72B-Instruct is Alibaba's 72-billion-parameter instruction-tuned vision-language model from the Qwen2-VL series, released in late 2024. It processes images at arbitrary resolutions via a dynamic visual token scheme, understands videos longer than 20 minutes, and supports multilingual text within images across European languages, Japanese, Korean, Arabic, and more. At release it achieved performance comparable to GPT-4o and Claude 3.5 Sonnet on multimodal benchmarks and supports autonomous device-control tasks.