Atlas

Models

← All models

DeepSeek-R1-Distill-Qwen-1.5B

DeepSeek · 2025-01-20 · 1.8B parameters

DeepSeek-R1-Distill-Qwen-1.5B is a dense reasoning checkpoint produced by supervised fine-tuning of Qwen2.5-Math-1.5B on approximately 800,000 samples curated with DeepSeek-R1, without an RL stage. Released January 20, 2025, it is the smallest model in the original R1 distillation series; its architecture config declares 131,072 positions while tokenizer_config.json sets model_max_length to 16,384.

Benchmark scores

BenchmarkScore
AA-LCR0.3
Artificial Analysis Intelligence Index3.7
BBH (Open LLM Leaderboard v2)32.4
BigCodeBench Complete7.9
BigCodeBench Instruct7.0
BigCodeBench-Hard Instruct0.0
BRIDGE Medical (chain-of-thought)13.4
BRIDGE Medical (few-shot)14.9
BRIDGE Medical (zero-shot)14.3
Capability84.8
Global PIQA - Non-Parallel (Strict Exact Match)27.9
Global PIQA - Parallel (Strict Exact Match)7.7
GPQA (Open LLM Leaderboard v2)25.6
GPQA Diamond9.8
HMMT Feb 202511.7
Humanity's Last Exam (Text-Only)3.3
IFBench (Artificial Analysis)13.2
IFEval (Open LLM Leaderboard v2)34.6
LiveCodeBench (Artificial Analysis)7.0
MATH Level 5 (Open LLM Leaderboard v2)16.9
MATH-500 (Artificial Analysis source)68.7
MEDIC (clinical summarization)82.3
MEDIC (closed-ended)34.6
MEDIC (open-ended Elo)988
MMLU Pro26.9
MMLU-Pro (Open LLM Leaderboard v2)11.9
MuSR (Open LLM Leaderboard v2)36.3
SciCode (Artificial Analysis)6.6
Loading Atlas data…