Atlas

Models

← All models

Phi-3.5 Mini Instruct

Microsoft Research · 2024-08-22 · 3.8B parameters

Phi-3.5 Mini Instruct is Microsoft's dense, instruction-tuned Phi-3.5 small language model, announced on August 22, 2024. The 3.8B-class checkpoint supports a 131,072-token context through LongRoPE and was trained on a reported 3.4 trillion-token mixture, followed by SFT, PPO, and DPO, with improved multilingual and long-context capability over the June 2024 Phi-3 Mini update. Microsoft publishes the text-in, text-out weights under the MIT license.

Benchmark scores

BenchmarkScore
BBH (Open LLM Leaderboard v2)55.2
BigCodeBench Complete38.5
BigCodeBench Instruct32.8
BigCodeBench-Hard Complete12.8
BigCodeBench-Hard Instruct16.9
BRIDGE Medical (chain-of-thought)23.9
BRIDGE Medical (few-shot)31.3
BRIDGE Medical (zero-shot)25.4
Capability107.5
EQ-Bench v254.7
EQ-Bench v2 + MAGI-Hard Combined53.8
Global PIQA - Non-Parallel (Strict Exact Match)56.9
Global PIQA - Parallel (Strict Exact Match)32.0
GPQA (Open LLM Leaderboard v2)34.0
IFEval (Open LLM Leaderboard v2)57.7
MAGI-Hard52.9
MATH Level 5 (Open LLM Leaderboard v2)19.6
MEDIC (clinical summarization)81.6
MEDIC (closed-ended)39.0
MMLU-Pro (Open LLM Leaderboard v2)39.6
MuSR (Open LLM Leaderboard v2)40.2
SnakeBench Average Apples0.3
SnakeBench Best Apples2.0
SnakeBench Rating10.0
SnakeBench Total Apples4.0
SnakeBench Win Rate33.3
ZeroEval33.7
ZeroEval CRUX42.1
ZeroEval GSM8K82.0
ZeroEval MATH Level 518.7
Loading Atlas data…