MiMo-V2.5-Pro
Xiaomi · 2026-04-27 · 1T parameters
MiMo-V2.5-Pro is Xiaomi's most capable open-weight MiMo model at release, a sparse Mixture-of-Experts text model with approximately 1.02T total parameters and 42B active per token, 60 sliding-window and 10 global-attention layers, and a 1,048,576-token context window. It was pretrained on a reported 27T tokens using FP8 mixed precision and incorporates three Multi-Token Prediction modules, then post-trained with supervised fine-tuning, agentic reinforcement learning, and multi-teacher on-policy distillation for demanding agentic, software-engineering, and long-horizon tasks.