InternVL-Chat-V1-5
OpenGVLab · 2024-04-18 · 25.5B parameters
InternVL-Chat-V1-5 is OpenGVLab's 25,514,186,112-parameter open-weight multimodal chat checkpoint, combining InternViT-6B-448px-V1-5, a two-layer MLP projector, and InternLM2-Chat-20B. Released on April 18, 2024, it uses dynamic high-resolution processing with 1–12 tiles of 448×448 pixels during training and zero-shot scaling to as many as 40 tiles (4K) at inference, plus an English–Chinese instruction dataset emphasizing OCR and documents; its official examples also support text-only, multi-image, and sampled-video conversations. Its language-model config exposes 32,768 maximum positions with 3× dynamic NTK RoPE scaling, although the published training recipe and tokenizer cap sequences at 4,096 tokens.
Benchmark scores
| Benchmark | Score |
|---|---|
| AI2D (OpenVLM) | 80.6 |
| Capability | 111.8 |
| CharXiv Reasoning | 29.2 |
| CharXiv-D | 58.5 |
| HallusionBench (OpenVLM) | 47.4 |
| MathVista (OpenVLM) | 56.1 |
| MM-Vet (OpenVLM) | 55.4 |
| MMMU Validation (OpenVLM) | 46.8 |
| MMStar (OpenVLM) | 57.1 |
| Video-MME | 50.7 |