DeepSeek-R1-Distill-Qwen-32B
DeepSeek · 2025-01-20 · 32.8B parameters
DeepSeek-R1-Distill-Qwen-32B is a dense reasoning checkpoint fine-tuned from Qwen2.5-32B using approximately 800,000 samples curated with DeepSeek-R1; DeepSeek states that the distilled models used only supervised fine-tuning, without reinforcement learning. Released January 20, 2025, it is the largest Qwen-based R1 distillation, and its released configuration declares a 131,072-token maximum while tokenizer_config.json separately sets model_max_length to 16,384.