Qwen3-Next-80B-A3B-Thinking
Alibaba · 2025-09-10 · 81.3B parameters
Qwen3-Next-80B-A3B-Thinking is Alibaba's thinking-only post-trained variant of the sparse Qwen3-Next-80B-A3B causal language model. It combines Gated DeltaNet with gated grouped-query attention and high-sparsity MoE layers; its native 262,144-token context can be extended with YaRN to 1,010,000 tokens.