Phi-4-mini-instruct
Microsoft Research · 2025-02-26 · 3.8B parameters
Phi-4-mini-instruct is Microsoft's open-weight, instruction-tuned 3.8B-class text model, publicly launched on February 26, 2025. It is a dense 32-layer decoder-only Transformer with grouped-query attention, tied input/output embeddings, LongRoPE support for 131,072 tokens, and an o200k_base BPE tokenizer; Microsoft reports a 5-trillion-token pretraining corpus emphasizing high-quality web, synthetic reasoning, mathematics, code, and multilingual data. The released checkpoint was post-trained with supervised fine-tuning and direct preference optimization for instruction following, safety, and function calling; the report's separately continued-trained reasoning experiment is not this artifact.