Ling-1T
InclusionAI · 2025-10-09 · 999.7B parameters
Ling-1T is inclusionAI's first flagship non-thinking model in the Ling 2.0 series, an open-weight sparse MoE with 1 trillion total parameters and approximately 51 billion activated parameters. Its 80-layer architecture uses GQA, 256 routed experts with eight selected per token plus one shared expert, SwiGLU, RMSNorm, QK normalization, and Partial RoPE; the released configuration has a native 32K-token window that the official instructions extend to 128K with YaRN.