Qwen2.5-VL-3B-Instruct
Alibaba · 2025-01-26 · 3.8B parameters
Qwen2.5-VL-3B-Instruct is Alibaba's compact instruction-tuned 3B-class checkpoint in the Qwen2.5-VL family. It accepts text, image, and video inputs and generates text, with dynamic-resolution vision, OCR and document understanding, visual grounding, long-video understanding, and visual-agent capabilities. Its published configuration sets 128,000 maximum positions, while the model-card guidance identifies 32,768 tokens as the original text context and prescribes YaRN for longer text inputs.
Benchmark scores
| Benchmark | Score |
|---|---|
| Capability | 112.5 |
| CharXiv Reasoning | 31.3 |
| CharXiv-D | 58.6 |
| MATH-Vision | 21.2 |
| ScreenSpot-Pro | 16.1 |