Atlas

Models

← All models

Phi-4-mini-instruct

Microsoft Research · 2025-02-26 · 3.8B parameters

Phi-4-mini-instruct is Microsoft's open-weight, instruction-tuned 3.8B-class text model, publicly launched on February 26, 2025. It is a dense 32-layer decoder-only Transformer with grouped-query attention, tied input/output embeddings, LongRoPE support for 131,072 tokens, and an o200k_base BPE tokenizer; Microsoft reports a 5-trillion-token pretraining corpus emphasizing high-quality web, synthetic reasoning, mathematics, code, and multilingual data. The released checkpoint was post-trained with supervised fine-tuning and direct preference optimization for instruction following, safety, and function calling; the report's separately continued-trained reasoning experiment is not this artifact.

Benchmark scores

BenchmarkScore
AA-LCR13.7
AA-Omniscience Index-61.0
AIME 20256.7
Artificial Analysis Agentic Index0.3
Artificial Analysis Coding Index3.8
Artificial Analysis Intelligence Index6.0
Artificial Analysis Omniscience Accuracy8.3
Artificial Analysis Omniscience Hallucination Rate75.5
Artificial Analysis Openness Index50.0
BBH (Open LLM Leaderboard v2)56.9
BRIDGE Medical (chain-of-thought)22.5
BRIDGE Medical (few-shot)29.9
BRIDGE Medical (zero-shot)24.5
Capability108.9
CritPt0.0
GDPval-AA v2-120
GDPval-AA v2 (normalized)0.0
Global PIQA - Non-Parallel (Strict Exact Match)58.5
Global PIQA - Parallel (Strict Exact Match)33.5
GPQA (Open LLM Leaderboard v2)31.0
GPQA Diamond33.1
Humanity's Last Exam (Text-Only)4.2
IFBench (Artificial Analysis)21.1
IFEval (Open LLM Leaderboard v2)73.8
LiveCodeBench (Artificial Analysis)12.6
MATH Level 5 (Open LLM Leaderboard v2)17.0
MATH-500 (Artificial Analysis source)69.6
MMLU Pro46.5
MMLU-Pro (Open LLM Leaderboard v2)39.3
MuSR (Open LLM Leaderboard v2)38.7
Loading Atlas data…