Mixtral 8x22B
Mistral AI · 2024-04-17 · 140.6B parameters
Mixtral 8x22B v0.1 is Mistral AI's pretrained base sparse Mixture-of-Experts text model, released on April 17, 2024. Its configuration routes each token through two of eight local experts, supports 65,536-token sequences, and uses grouped-query attention; Mistral describes the model at rounded precision as 141B total parameters with 39B active. The weights are released under Apache 2.0.