Qwen3 VL 30B A3B Instruct
Alibaba · 2025-10-04 · 31.1B parameters
Qwen3-VL-30B-A3B-Instruct is Alibaba's non-thinking instruction-tuned 30B-A3B mixture-of-experts vision-language model. It accepts text, images, and video and supports multilingual OCR, spatial grounding, and visual-agent interaction. Its native context is 262,144 tokens and can be extended to 1,000,000 tokens with YaRN.