Mixtral 8x22B Release Notes
Mistral AI has announced the release of Mixtral 8x22B, a sparse Mixture-of-Experts (SMoE) model designed to provide high performance and cost efficiency. The model utilizes 39B active parameters out of a total of 141B, allowing it to maintain a high performance-to-cost ratio while remaining faster than dense 70B models.
Architecture and Efficiency
Mixtral 8x22B is built as a sparse Mixture-of-Experts model, which means it only activates a fraction of its total parameters during inference. With 39B active parameters and 141B total parameters, the model is designed for cost efficiency and speed. Mistral AI states that this architecture makes the model faster than any dense 70B model and more capable than other open-weight models available under both permissive and restrictive licenses.
Key Capabilities and Technical Specifications
Mixtral 8x22B offers several core technical strengths:
- Multilingualism: The model is fluent in English, French, Italian, German, and Spanish.
- Context Window: It features a 64K token context window, enabling precise information recall from large documents.
- Function Calling: The model is natively capable of function calling. When combined with the constrained output mode on la Plateforme, this supports large-scale application development and tech stack modernization.
- Reasoning, Math, and Coding: The model is optimized for reasoning and demonstrates strong performance in mathematics and coding tasks compared to other open models.
Performance Benchmarks
Mixtral 8x22B demonstrates superior performance across several industry benchmarks:
Reasoning and Knowledge
The model is optimized for reasoning and shows leading performance on benchmarks including MMLU (Measuring massive multitask language in understanding), HellaSwag, Wino Grande, Arc Challenge, TriviaQA, and NaturalQS.
Multilingual Performance
Mixtral 8x22B outperforms LLaMA 2 70B on the HellaSwag, Arc Challenge, and MMLU benchmarks specifically in French, German, Spanish, and Italian.
Mathematics and Coding
Compared to other open models, Mixtral 8x22B performs best in coding and maths tasks. The instructed version of the model achieves a score of 90.8% on GSM8K maj@8 and a 44.6% score on Math maj@4.
Licensing and Availability
Mixtral 8x22B is released under the Apache 2.0 license, the most permissive open-source license, allowing for unrestricted use. The availability of the base model also makes it a suitable foundation for fine-tuning use cases.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch