Grok 4.5 Release Notes
xAI has released Grok 4.5, a model designed for high-performance coding, agentic tasks, and knowledge work. The model is positioned as xAI's strongest model to date, developed in collaboration with Cursor.
Engineering and Coding Performance
Grok 4.5 demonstrates strong capabilities in real-world engineering tasks, particularly in software engineering benchmarks. It outperforms several leading models in specific resolution rates, though its performance varies across different evaluation harnesses.
Benchmark Results
- SWE Marathon (pass@1): Grok 4.5 achieved a 29.0% resolution rate, exceeding Opus 4.8 (max) at 26.0% and Fable (max) at 24.0%.
- Terminal Bench 2.1: Grok 4.5 scored 83.3%, placing it closely behind Fable (max) at 84.3% and GPT 5.5 (xhigh) at 83.4%.
- SWE Bench Pro: Grok 4.5 achieved a 64.7% resolve rate, outperforming Opus 4.7 (max) at 64.3%, GLM 5.2 at 62.1%, and GPT 5.5 (xhigh) at 58.6%.
- DeepSWE 1.0: Grok 4.5 scored 62.0%, trailing Fable (max) at 66.1% and GPT 5.5 (xhigh) at 64.31%.
- DeepSWE 1.1: Grok 4.5 scored 53%, trailing Fable (max) at 70%, GPT 5.5 (xhigh) at 67%, and Opus 4.8 (max) at 59%.
Technical Training and Infrastructure
Grok 4.5 was trained using tens of thousands of NVIDIA GB300 GPUs. The training process emphasized data quality over raw volume through deduplication, quality scoring, and domain-focused selection to maintain a high-signal data mixture.
Reinforcement Learning and Agentic Capabilities
To enhance per-token intelligence, xAI scaled reinforcement learning (RL) across hundreds of thousands of tasks. This RL training focused on multi-step software engineering and technical work using automated and model-based grading. The training stack supports highly asynchronous operations, allowing agentic rollouts to run for several hours while learning continues across the GPU cluster.
Efficiency and Serving Speed
Grok 4.5 is served at a speed of 80 tokens per second (TPS). It is characterized by high token efficiency, requiring significantly fewer tokens to resolve tasks compared to competitors.
On the SWE Bench Pro task, Grok 4.5 resolves tasks with an average of 15,954 output tokens, which is approximately 4.2x fewer tokens than Opus 4.8 (max), which requires 67,020 tokens.
Office Productivity and Integration
Grok 4.5 is the default model for Grok Build and is integrated into Cursor. It extends its technical capabilities to office software, including:
- Excel: Building complex models with web research, multi-sheet formulas, and internal notes.
- PowerPoint: Creating complex diagrams using native shapes and designing slide content.
- Word: Writing clear prose and outlining business reviews.
Pricing and Availability
- Pricing: Input tokens are priced at $2 per million, and output tokens are priced at $6 per million.
- Availability: The model is available via Grok Build, Cursor (all plans), and the SpaceXAI console API.
Sources
- OriginalIntroducing Grok 4.5
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch