Qwen2 VL 7B Instruct
Alibaba · 2024-08-29 · 8.3B parameters
Qwen2-VL 7B Instruct is Alibaba's roughly 8.3-billion-parameter instruction-tuned vision-language model (about a 7.6B language model plus a 675M vision encoder) from the Qwen2-VL series, released in 2024. It combines a Qwen2 language backbone with a visual encoder supporting dynamic resolutions and long-context processing, enabling document analysis, OCR, chart understanding, temporal reasoning over video, and structured output generation.