Mistral 3 Release Notes: Mistral Large 3 and Ministral 3 Family
Mistral AI has announced Mistral 3, a new generation of models including the high-capacity Mistral Large 3 and the edge-focused Ministral 3 series. All models in this release are available under the Apache 2.0 license to support distributed intelligence and developer customization.
Mistral Large 3: Frontier Open-Weight Performance
Mistral Large 3 is a sparse mixture-of-experts (MoE) model featuring 675B total parameters and 41B active parameters. Trained from scratch on 3,000 NVIDIA H200 GPUs, it is designed to achieve parity with the leading instruction-tuned open-weight models for general prompts.
Key capabilities and benchmarks include:
- LMArena Ranking: Mistral Large 3 ranks #2 in the OSS non-reasoning models category and #6 among all OSS models overall on the LMArena leaderboard.
- Multimodal and Multilingual: The model demonstrates image understanding and best-in-class performance for multilingual conversations in non-English and non-Chinese languages.
- Availability: Both base and instruction fine-tuned versions are released today, with a reasoning version scheduled for future release.
Ministral 3: Edge-Optimized Intelligence
The Ministral 3 series provides state-of-the-art performance-to-cost ratios for local and edge deployments. The family consists of three model sizes: 3B, 8B, and 14B parameters.
For each size, Mistral AI has released three variants:
- Base: The fundamental pre-trained model.
- Instruct: Optimized for following directions; these models match or exceed comparable models while often generating an order of magnitude fewer tokens.
- Reasoning: Designed for high-accuracy tasks, with the 14B variant achieving 85% accuracy on AIME ’25.
All Ministral 3 models include native multimodal (image understanding) and multilingual capabilities.
Hardware Optimization and Ecosystem Integration
Mistral AI collaborated with NVIDIA, vLLM, and Red Hat to ensure Mistral 3 is accessible and efficient across various hardware environments.
Data Center and Enterprise Deployment
- NVFP4 Format: A checkpoint in NVFP4 format, created with
llm-compressor, allows Mistral Large 3 to run efficiently on Blackwell NVL72 systems or a single 8×A100 or 8×H100 node using vLLM. - NVIDIA Integration: All Mistral 3 models were trained on NVIDIA Hopper GPUs using HBM3e memory. NVIDIA provided inference support for TensorRT-LLM and SGLang to enable low-precision execution.
- Advanced Kernels: For the MoE architecture of Large 3, NVIDIA integrated Blackwell attention and MoE kernels and collaborated on speculative decoding to support long-context, high-throughput workloads on GB200 NVL72.
Edge and Local Deployment
- Device Support: Ministral models are optimized for deployment on DGX Spark, RTX PCs, laptops, and Jetson devices.
Availability and Customization
Mistral 3 is available via Mistral AI Studio, Amazon Bedrock, Azure Foundry, Hugging Face, Modal, IBM WatsonX, OpenRouter, Fireworks, Unsloth AI, and Together AI. It will soon be available on NVIDIA NIM and AWS SageMaker.
For organizations requiring specialized AI, Mistral AI offers custom model training services to fine-tune or fully adapt models for domain-specific tasks or proprietary datasets.
Sources
- OriginalIntroducing Mistral 3
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch