Step3-VL-10B
StepFun · 2026-01-14 · 10.2B parameters
Step3-VL-10B is StepFun's open-weight 10B-class vision-language model, pairing a 1.8B PE-lang visual encoder with an 8B Qwen3 language decoder; the complete released checkpoint contains 10,171,750,144 parameters. Released in January 2026, it targets multimodal reasoning across STEM, document understanding, and GUI interaction while using substantially less compute than much larger models.