Atlas

Models

← All models

InternVL3-38B

OpenGVLab · 2025-04-11 · 38.4B parameters

InternVL3-38B is OpenGVLab's 38,388,164,992-parameter vision-language checkpoint, combining an InternViT-6B-448px-V2.5 encoder with a Qwen2.5-32B language backbone through a two-layer MLP projector. Released after native multimodal pre-training, supervised fine-tuning, and mixed preference optimization, it accepts text, images, and sampled video and generates text; its repository metadata declares Apache-2.0, while the surrounding InternVL project code is MIT.

Benchmark scores

BenchmarkScore
AI2D (OpenVLM)88.7
Capability132.5
CharXiv Reasoning46.4
CharXiv-D87.2
HallusionBench (OpenVLM)58.4
MATH-Vision34.5
MathVista (OpenVLM)76.3
MM-Vet (OpenVLM)81.1
MMMU Validation (OpenVLM)69.7
MMStar (OpenVLM)72.6
Loading Atlas data…