InternVL3.5-8B
OpenGVLab · 2025-08-26 · 8.5B parameters
InternVL3.5-8B is OpenGVLab's dense 8,528,318,464-parameter vision-language checkpoint, combining a 24-layer InternViT-300M vision encoder and two-layer MLP projector with a 36-layer Qwen3-8B language backbone. The suffix-free release completed multimodal continual pre-training, supervised fine-tuning, and Cascade RL (MPO followed by GSPO); it accepts text, images, and sampled video frames, generates text, and exposes prompt-selectable Thinking mode. Its checkpoint config exposes 40,960 maximum positions while OpenGVLab reports 32K-token CPT/SFT sequences; unlike separately named InternVL3.5-Flash variants, it does not contain the Visual Resolution Router.