Mistral AI releases Ministral 3B and 8B edge models

Mistral AI has announced the release of Ministral 3B and Ministral 8B, two new small language models (SLMs) optimized for edge computing and on-device use cases. These models are designed to provide high performance in knowledge, reasoning, and function-calling while maintaining the low latency and efficiency required for local deployment.

Technical Specifications and Architecture

Ministral 3B and 8B both support a context length of up to 128k tokens, though it is currently 32k on vLLM. To optimize performance, Ministral 8B utilizes a special interleaved sliding-window attention pattern, which enables faster and more memory-efficient inference.

Target Use Cases and Applications

Les Ministraux are built for scenarios requiring local, privacy-first inference and low latency. Key applications include:

  • On-device computing: Local analytics, internet-less smart assistants, and autonomous robotics.
  • Edge translation: On-device translation services.
  • Agentic workflows: Serving as efficient intermediaries for function-calling, input parsing, and task routing when used alongside larger models like Mistral Large.

Performance Benchmarks

Mistral AI reports that Ministral 3B and 8B consistently outperform their peers in the sub-10B category. According to the company's internal evaluation framework, the smallest model, Ministral 3B, already outperforms the original Mistral 7B on most benchmarks.

Availability, Pricing, and Licensing

Both models are available immediately via API. Pricing on la Plateforme is as follows:

Model API Endpoint Price per Million Tokens (Input/Output) License
Ministral 8B ministral-8b-latest $0.10 Mistral Commercial License / Mistral Research License
Ministral 3B ministral-3b-latest $0.04 Mistral Commercial License

For self-deployed use, commercial licenses are required. Mistral AI also offers assistance with lossless quantization to maximize performance for specific use cases. The model weights for Ministral 8B Instruct are available for research use, and both models will be available through cloud partners shortly.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch