Ling Flash 2.0
InclusionAI · 2025-09-17 · 102.9B parameters
Ling-flash-2.0 is InclusionAI's open-weight, 100B-class sparse MoE language model in the Ling 2.0 family; the official checkpoint manifest gives an exact total size of 102,889,705,216, while InclusionAI describes its activated size only as 6.1B parameters (4.8B excluding embeddings). It uses 32 layers, sigmoid routing, QK-Norm, GQA, RMSNorm, SwiGLU, and Partial RoPE, and the official card reports 20T+ high-quality training tokens together with supervised fine-tuning and multi-stage reinforcement learning. The released configuration has an exact native 32,768-token limit, while InclusionAI documents an optional fourfold YaRN extension only as a rounded 128K context; the supplied chat template fixes detailed thinking off, making this a non-thinking language model rather than a switchable reasoning checkpoint.