Atlas

Models

← All models

Phi-3.5 Mini Instruct

Microsoft Research · 2024-08-22 · 3.8B parameters

Phi-3.5 Mini Instruct is Microsoft's dense, instruction-tuned Phi-3.5 small language model, announced on August 22, 2024. The 3.8B-class checkpoint supports a 131,072-token context through LongRoPE and was trained on a reported 3.4 trillion-token mixture, followed by SFT, PPO, and DPO, with improved multilingual and long-context capability over the June 2024 Phi-3 Mini update. Microsoft publishes the text-in, text-out weights under the MIT license.

Benchmark scores

BenchmarkScore
BBH (Open LLM Leaderboard v2)55.2
BigCodeBench Complete38.5
BigCodeBench Instruct32.8
BigCodeBench-Hard Complete12.8
BigCodeBench-Hard Instruct16.9
BoolQ78.0
BRIDGE Medical (chain-of-thought)23.9
BRIDGE Medical (few-shot)31.3
BRIDGE Medical (zero-shot)25.4
Capability107.5
EQ-Bench v254.7
EQ-Bench v2 + MAGI-Hard Combined53.8
Global PIQA - Non-Parallel (Strict Exact Match)56.9
Global PIQA - Parallel (Strict Exact Match)32.0
GPQA (Open LLM Leaderboard v2)34.0
GSM8K86.2
IFEval (Open LLM Leaderboard v2)57.7
MAGI-Hard52.9
MATH Level 5 (Open LLM Leaderboard v2)19.6
MEDIC (clinical summarization)81.6
MEDIC (closed-ended)39.0
MMLU-Pro (Open LLM Leaderboard v2)39.6
MuSR (Open LLM Leaderboard v2)40.2
PIQA81.0
SnakeBench Average Apples0.3
SnakeBench Best Apples2.0
SnakeBench Rating10.0
SnakeBench Total Apples4.0
SnakeBench Win Rate33.3
ZeroEval33.7
Loading Atlas data…