Pixtral Large Release Notes
Mistral AI has introduced Pixtral Large, a 124B open-weights multimodal model designed for frontier-level image understanding. Built on top of Mistral Large 2, the model enables high-performance processing of documents, charts, and natural images without compromising the original text-only understanding capabilities of the base model.
Model Architecture and Specifications
Pixtral Large utilizes a multimodal decoder and a vision encoder to process visual and textual data. Its technical specifications include:
- Parameters: A 123B multimodal decoder paired with a 1B parameter vision encoder (124B total).
- Context Window: A 128K context window, which is capable of accommodating a minimum of 30 high-resolution images.
- Availability: The model is available via the Mistral API as
pixtral-large-latest, on le Chat, and as open weights on HuggingFace. - Licensing: Distributed under the Mistral Research License (MRL) for research and educational use, and a separate Mistral Commercial License for production and commercial experimentation.
Performance Benchmarks
Pixtral Large demonstrates state-of-the-art performance across several multimodal benchmarks, frequently outperforming both open-weights and proprietary models.
Mathematical and Document Reasoning
On the MathVista benchmark, which tests complex mathematical reasoning over visual data, Pixtral Large achieved a score of 69.4%, outperforming all other tested models. In evaluations of complex charts and documents using ChartQA and DocVQA, the model surpassed the performance of GPT-4o and Gemini-1.5 Pro.
Real-World and Human-Preference Evaluation
Pixtral Large outperformed Claude-3.5 Sonnet (new), Gemini-1.5 Pro, and GPT-4o (latest) on MM-MT-Bench, an open-source, judge-based evaluation designed to reflect real-world multimodal LLM use cases. Additionally, on the LMSys Vision Leaderboard, Pixtral Large is the highest-ranking open-weights model, surpassing its nearest competitor by nearly 50 ELO points and outperforming proprietary models such as the August 2024 version of GPT-4o.
Visual Understanding Capabilities
Pixtral Large is capable of complex reasoning across various visual formats, including:
- Multilingual OCR and Reasoning: The model can extract data from images (such as receipts in different languages) and perform mathematical calculations based on that data.
- Chart Analysis: The model can identify specific trends and anomalies in technical data, such as identifying the exact step count where training loss instability begins in a loss curve graph.
- Visual Information Extraction: The model can accurately identify and list entities (such as company names) from website screenshots.
Mistral Large 24.11 Update
Concurrent with the Pixtral Large release, Mistral AI updated its state-of-the-art text model to version 24.11. This update provides significant improvements over the 24.07 version, specifically in:
- Long context understanding
- Function calling accuracy
- System prompt handling
Mistral Large 24.11 is positioned for enterprise RAG and agentic workflows, including task automation and semantic document understanding. It is available via API as pixtral-large-latest and for self-deployment on HuggingFace.
Sources
- OriginalPixtral Large
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch