Atlas

Models

← All models

Phi-3.5-MoE-instruct

Microsoft Research · 2024-08-22 · 41.9B parameters

Phi-3.5 MoE is Microsoft's mixture-of-experts model in the Phi-3.5 family, released in August 2024, with a 16×3.8B MoE architecture and approximately 6.6B active parameters per token when routing to 2 experts. It was trained using Microsoft's GRIN (GRadient-INformed) MoE technique to improve expert specialization, on a reported 4.9 trillion tokens with a 128K context window and support for over 20 languages. Despite activating only a fraction of its total parameters during inference, it outperforms larger dense models on several reasoning and multilingual benchmarks.

Benchmark scores

BenchmarkScore
BBH (Open LLM Leaderboard v2)64.1
BoolQ84.6
BRIDGE Medical (chain-of-thought)25.3
BRIDGE Medical (few-shot)36.6
BRIDGE Medical (zero-shot)29.5
Capability119.5
EQ-Bench v277.0
EQ-Bench v2 + MAGI-Hard Combined72.1
GPQA (Open LLM Leaderboard v2)35.6
GSM8K88.7
IFEval (Open LLM Leaderboard v2)69.2
MAGI-Hard67.3
MATH Level 5 (Open LLM Leaderboard v2)31.2
MMLU-Pro (Open LLM Leaderboard v2)46.6
MuSR (Open LLM Leaderboard v2)45.6
PIQA88.6
Loading Atlas data…