Phi-3.5-MoE-instruct
Microsoft Research · 2024-08-22 · 41.9B parameters
Phi-3.5 MoE is Microsoft's mixture-of-experts model in the Phi-3.5 family, released in August 2024, with a 16×3.8B MoE architecture and approximately 6.6B active parameters per token when routing to 2 experts. It was trained using Microsoft's GRIN (GRadient-INformed) MoE technique to improve expert specialization, on a reported 4.9 trillion tokens with a 128K context window and support for over 20 languages. Despite activating only a fraction of its total parameters during inference, it outperforms larger dense models on several reasoning and multilingual benchmarks.