Nanbeige4-3B-Thinking-2511
Nanbeige LLM Lab · 2025-11-21 · 3.9B parameters
Nanbeige4-3B-Thinking-2511 is the reasoning-enhanced checkpoint in the Nanbeige4-3B family and an upgraded iteration of Nanbeige4-3B-Thinking-2510, improved through knowledge distillation and targeted reinforcement learning. Its base model was pretrained on a reported 23 trillion-token corpus, followed by post-training with supervised fine-tuning, distillation, and multi-stage reinforcement learning.