MiMo-V2-Flash
Xiaomi · 2025-12-16 · 309.8B parameters
MiMo-V2-Flash is Xiaomi MiMo's open-weight sparse MoE text model; Xiaomi reports rounded labels of 309B total and 15B active parameters, while its released FP8 checkpoint contains 309,766,601,088 model parameters excluding quantization-scale tensors. Pretrained across 27T tokens and post-trained with multi-teacher on-policy distillation and agentic RL, it supports optional thinking and a 262,144-token context. Its 48-layer design interleaves 39 128-token sliding-window GQA layers with 9 global GQA layers, routes each token to 8 of 256 experts, applies RoPE to 64 of each 192-dimensional query/key head, and includes three MTP blocks for self-speculative decoding.