Qwen3-VL-8B-Instruct
Alibaba · 2025-10-15 · 8.8B parameters
Qwen3-VL-8B-Instruct is the instruction-tuned, dense 8B vision-language model in Alibaba's Qwen3-VL family. It offers image and video understanding, multilingual OCR, grounding, and visual-agent capabilities at a smaller scale. Its context is 262,144 tokens natively and expandable to 1 million tokens.