Atlas

Models

← All models

Seed1.5-VL

ByteDance Seed · 2025-05-12

Seed1.5-VL is ByteDance Seed's vision-language foundation model released in May 2025, combining a vision encoder reported at 532M parameters, an MLP adapter, and a Mixture-of-Experts language model described as approximately 20B active parameters. Its VLM pre-training stages used stated budgets of 16B, 3T, and 240B tokens after separate vision-encoder and language-model pre-training, so exact cumulative training exposure is not disclosed; it supports text, image, and video inputs, a 131,072-token context, and switchable thinking and non-thinking inference. ByteDance served it through Volcano Engine as doubao-1-5-thinking-vision-pro-250428, while the weights were not publicly released.

Benchmark scores

BenchmarkScore
WorldVQA Accuracy34.9
WorldVQA Brands F-score32.3
WorldVQA Correct Given Attempted35.5
WorldVQA Culture F-score33.4
WorldVQA Entertainment F-score33.6
WorldVQA F-score35.2
WorldVQA Geography F-score36.1
WorldVQA Nature F-score41.4
WorldVQA Not Attempted1.6
WorldVQA Objects F-score32.8
WorldVQA Sports F-score43.7
WorldVQA Transportation F-score35.0
ZeroBench (Main questions, pass@1)2.0
ZeroBench (Subquestions, pass@1)30.8
Loading Atlas data…