Seed1.5-VL
ByteDance Seed · 2025-05-12
Seed1.5-VL is ByteDance Seed's vision-language foundation model released in May 2025, combining a vision encoder reported at 532M parameters, an MLP adapter, and a Mixture-of-Experts language model described as approximately 20B active parameters. Its VLM pre-training stages used stated budgets of 16B, 3T, and 240B tokens after separate vision-encoder and language-model pre-training, so exact cumulative training exposure is not disclosed; it supports text, image, and video inputs, a 131,072-token context, and switchable thinking and non-thinking inference. ByteDance served it through Volcano Engine as doubao-1-5-thinking-vision-pro-250428, while the weights were not publicly released.