Mixtral 8x7B
Mistral AI · 2023-12-11 · 46.7B parameters
Mixtral 8x7B v0.1 is Mistral AI's pretrained, open-weight sparse mixture-of-experts base model, released on December 11, 2023. Each decoder layer has eight feed-forward experts and routes each token to two of them; Mistral reports rounded counts of 46.7B total parameters and 12.9B active parameters per token. The weights use Apache 2.0, and the checkpoint configuration specifies a 32,768-token context.