Ring-flash-2.0
InclusionAI · 2025-09-19 · 102.9B parameters
Ring-flash-2.0 is InclusionAI's open-weight, intrinsic-thinking sparse MoE model, derived from Ling-flash-2.0-base through lightweight Long-CoT SFT, reinforcement learning with verifiable rewards, and RLHF. The released checkpoint has 102,889,705,216 parameters; InclusionAI's release materials round this to 100B total and 6.1B activated (4.8B excluding embeddings), with eight of 256 routed experts plus one shared expert used per MoE layer. It natively supports 32,768 tokens and is documented for extension to 131,072 tokens with YaRN, while InclusionAI reports generation above 200 tokens/s on four H20 GPUs.