DeepSeek-R1
DeepSeek · 2025-01-20 · 684.5B parameters
DeepSeek-R1 is DeepSeek's open-weight MoE reasoning checkpoint released January 20, 2025 and trained on top of DeepSeek-V3-Base. Its downloadable checkpoint contains exactly 684,489,845,504 model-weight parameters including one multi-token-prediction module; DeepSeek separately reports rounded figures of 671B main-model parameters, 37B activated per token, and 128K context support. Its multi-stage post-training combines cold-start supervised fine-tuning, reasoning-oriented GRPO, rejection-sampling supervised fine-tuning on about 600K reasoning and 200K non-reasoning examples, and a final reinforcement-learning stage.