Mistral Large 2 release notes / what's new

Mistral AI has released Mistral Large 2, a 123 billion parameter model optimized for high-throughput single-node inference and cost efficiency. This model is designed to compete with leading frontier models like GPT-4o, Claude 3 Opus, and Llama 3 405B, particularly in coding, reasoning, and multilingual tasks.

Model Architecture and Deployment

Mistral Large 2 is a 123B parameter model featuring a 128k context window. It is specifically engineered for single-node inference, allowing it to maintain high throughput for long-context applications.

Licensing and Access

  • Research and Non-Commercial Use: Available under the Mistral Research License (MRL-0.1).
  • Commercial Use: Requires a Mistral Commercial License for self-deployment.
  • API Access: Available via la Plateforme (as mistral-large-2407) and le Chat.
  • Cloud Providers: Available via Google Cloud Vertex AI (Managed API), Azure AI Studio, Amazon Bedrock, and IBM watsonx.ai.

Technical Performance and Benchmarks

Mistral Large 2 demonstrates significant improvements over its predecessor and competitive performance against other leading open models.

General Knowledge and Reasoning

On the MMLU benchmark, the pretrained version of Mistral Large 2 achieves an accuracy of 84.0%, establishing a new point on the performance/cost Pareto front for open models.

Coding and Mathematics

Following the development of Codestral 22B and Codestral Mamba, Mistral Large 2 was trained on a high proportion of code. It performs on par with GPT-4o, Claude 3 Opus, and Llama 3 405B in code generation and reasoning tasks.

To reduce hallucinations, the model was fine-tuned to be more cautious and discerning, training it to acknowledge when it lacks sufficient information to provide a confident answer. This focus on accuracy is reflected in improved performance on the GSM8K (8-shot) and MATH (0-shot, no CoT) benchmarks.

Instruction Following and Alignment

Mistral Large 2 shows drastic improvements in following precise instructions and managing long multi-turn conversations, as measured by MT-Bench, Wild Bench, and Arena Hard.

Unlike some models that increase scores by generating lengthy responses, Mistral Large 2 is optimized for conciseness. This design choice aims to reduce inference costs and facilitate quicker interactions in business applications.

Multilingual Capabilities and Tool Use

Mistral Large 2 is designed for global business use cases by training on a large proportion of multilingual data.

Language Support

The model excels in the following languages:

  • English, French, German, Spanish, Italian, Portuguese, Dutch, Russian, Chinese, Japanese, Korean, Arabic, and Hindi.

Tool Use and Function Calling

The model is equipped with enhanced retrieval skills and function calling capabilities. It is trained to execute both parallel and sequential function calls, making it suitable for complex business application orchestration.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch