Mistral Small 3 Release Notes

Mistral AI has introduced Mistral Small 3, a 24B-parameter model designed for low-latency generative AI tasks. Released under the Apache 2.0 license, the model is positioned as a high-efficiency alternative to larger open-weight models like Llama 3.3 70B and Qwen 32B, as well as proprietary models such as GPT-4o-mini.

High-Efficiency Performance and Latency

Mistral Small 3 is engineered to saturate performance at a size suitable for local deployment. By utilizing fewer layers than competing models, it substantially reduces the time required per forward pass, achieving a latency of 150 tokens per second.

Key performance metrics include:

  • MMLU Accuracy: Over 81%.
  • Speed: More than 3x faster than Llama 3.3 70B Instruct on the same hardware.
  • Efficiency: Mistral AI describes it as the most efficient model in its category.

Model Capabilities and Training

Mistral AI has released both pretrained and instruction-tuned checkpoints. Unlike some contemporary reasoning models, Mistral Small 3 was not trained using Reinforcement Learning (RL) or synthetic data. Mistral AI notes that the model is earlier in the production pipeline than models like DeepSeek R1, making it a strong base model for developers to build accrued reasoning capacities.

Target Use Cases

Mistral Small 3 is optimized for the "80%" of generative AI tasks that require robust instruction following and language performance with minimal latency. Recommended use cases include:

  • Conversational Assistance: Fast-response virtual assistants requiring near real-time interaction.
  • Agentic Workflows: Low-latency function calling for automated processes.
  • Domain Specialization: Fine-tuning the model to create subject matter experts in fields such as medical diagnostics, legal advice, and technical support.
  • Local Inference: Private deployment on a single RTX 4090 or a MacBook with 32GB RAM when quantized.

Industry-specific applications currently being evaluated by customers include fraud detection in financial services, customer triaging in healthcare, and on-device command and control for the automotive, robotics, and manufacturing sectors.

Availability and Ecosystem

Mistral Small 3 is available via mistral-small-latest or mistral-small-2501 on la Plateforme. It is also available through partners including Hugging Face, Ollama, Kaggle, Together AI, Fireworks AI, and IBM Watson X. Integration with NVIDIA NIM, Amazon SageMaker, Groq, Databricks, and Snowflake is expected soon.

Commitment to Open Source

Mistral AI is renewing its commitment to the Apache 2.0 license for general-purpose models, moving away from MRL-licensed models. This allows model weights to be downloaded, deployed locally, and modified freely.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch