Pixtral Large Release Notes

Mistral AI has introduced Pixtral Large, a 124B open-weights multimodal model designed for frontier-level image understanding. Built on top of Mistral Large 2, the model enables high-performance processing of documents, charts, and natural images without compromising the original text-only understanding capabilities of the base model.

Model Architecture and Specifications

Pixtral Large utilizes a multimodal decoder and a vision encoder to process visual and textual data. Its technical specifications include:

  • Parameters: A 123B multimodal decoder paired with a 1B parameter vision encoder (124B total).
  • Context Window: A 128K context window, which is capable of accommodating a minimum of 30 high-resolution images.
  • Availability: The model is available via the Mistral API as pixtral-large-latest, on le Chat, and as open weights on HuggingFace.
  • Licensing: Distributed under the Mistral Research License (MRL) for research and educational use, and a separate Mistral Commercial License for production and commercial experimentation.

Performance Benchmarks

Pixtral Large demonstrates state-of-the-art performance across several multimodal benchmarks, frequently outperforming both open-weights and proprietary models.

Mathematical and Document Reasoning

On the MathVista benchmark, which tests complex mathematical reasoning over visual data, Pixtral Large achieved a score of 69.4%, outperforming all other tested models. In evaluations of complex charts and documents using ChartQA and DocVQA, the model surpassed the performance of GPT-4o and Gemini-1.5 Pro.

Real-World and Human-Preference Evaluation

Pixtral Large outperformed Claude-3.5 Sonnet (new), Gemini-1.5 Pro, and GPT-4o (latest) on MM-MT-Bench, an open-source, judge-based evaluation designed to reflect real-world multimodal LLM use cases. Additionally, on the LMSys Vision Leaderboard, Pixtral Large is the highest-ranking open-weights model, surpassing its nearest competitor by nearly 50 ELO points and outperforming proprietary models such as the August 2024 version of GPT-4o.

Visual Understanding Capabilities

Pixtral Large is capable of complex reasoning across various visual formats, including:

  • Multilingual OCR and Reasoning: The model can extract data from images (such as receipts in different languages) and perform mathematical calculations based on that data.
  • Chart Analysis: The model can identify specific trends and anomalies in technical data, such as identifying the exact step count where training loss instability begins in a loss curve graph.
  • Visual Information Extraction: The model can accurately identify and list entities (such as company names) from website screenshots.

Mistral Large 24.11 Update

Concurrent with the Pixtral Large release, Mistral AI updated its state-of-the-art text model to version 24.11. This update provides significant improvements over the 24.07 version, specifically in:

  • Long context understanding
  • Function calling accuracy
  • System prompt handling

Mistral Large 24.11 is positioned for enterprise RAG and agentic workflows, including task automation and semantic document understanding. It is available via API as pixtral-large-latest and for self-deployment on HuggingFace.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch