InternVL3-1B
OpenGVLab · 2025-04-11 · 938M parameters
InternVL3-1B is OpenGVLab's 938,193,024-parameter vision-language checkpoint, pairing an InternViT-300M vision encoder with a Qwen2.5-0.5B language backbone through a two-layer MLP projector. Released after native multimodal pre-training, supervised fine-tuning, and mixed preference optimization, it accepts text, image, and sampled-video inputs and generates text. Its repository declares Apache-2.0, while the surrounding InternVL project code is MIT.
Benchmark scores
| Benchmark | Score |
|---|---|
| AI2D (OpenVLM) | 69.7 |
| Capability | 103.9 |
| CharXiv Reasoning | 21.0 |
| CharXiv-D | 47.1 |
| HallusionBench (OpenVLM) | 37.2 |
| MATH-Vision | 18.8 |
| MathVista (OpenVLM) | 46.9 |
| MM-Vet (OpenVLM) | 58.7 |
| MMMU Validation (OpenVLM) | 43.2 |
| MMStar (OpenVLM) | 52.3 |