Grok 4.5 Release: Engineering-Focused Model with High Token Efficiency
Grok 4.5 delivers high-intelligence reasoning with superior token efficiency
SpaceXAI has released Grok 4.5, a model specifically engineered for coding, agentic tasks, and technical knowledge work. The model is designed to provide "frontier-class" intelligence while significantly reducing the cost and time required to solve complex engineering problems through improved token efficiency and high throughput.
Key Performance and Efficiency Metrics
Grok 4.5 is served at speeds of 80 tokens per second (TPS). A primary differentiator of this release is its token efficiency; in SWE Bench Pro tasks, Grok 4.5 resolves issues using an average of 15,954 output tokens, which is approximately 4.2x fewer tokens than Opus 4.8 (max), which averages 67,020 tokens.
Pricing Structure:
- Input Tokens: $2 per million
- Output Tokens: $6 per million
- Cache Pricing: $0.50 per million (as noted by community analysis)
Benchmark Performance
In specialized engineering benchmarks, Grok 4.5 shows strong competitiveness, particularly in resolution rates for complex software tasks:
- SWE Marathon (pass@1): Grok 4.5 leads with a 29.0% resolution rate, surpassing Opus 4.8 (26.0%) and Fable (24.0%).
- SWE Bench Pro: Grok 4.5 achieved a 64.7% resolve rate, placing it among the top-tier models alongside Fable (80.4%) and Opus 4.8 (69.2%).
- DeepSWE 1.0: Grok 4.5 scored 62.0%, trailing slightly behind Fable (66.1%) and GPT 5.5 (64.31%).
- Terminal Bench 2.1: Grok 4.5 scored 83.3%, nearly identical to GPT 5.5 (83.4%) and Fable (84.3%).
Training Methodology and Data Strategy
Grok 4.5 was trained on tens of thousands of NVIDIA GB300 GPUs using a combination of high-signal data curation and scaled reinforcement learning (RL).
Integration with Cursor Data
The model was trained alongside Cursor, utilizing trillions of tokens of data that capture real-world user interactions with codebases and software tools. This allows the model to learn not just from static code, but from the iterative process of how developers interact with their environments and recover from mistakes.
Scaled Reinforcement Learning
SpaceXAI implemented an asynchronous training stack that allows agentic rollouts to run for several hours while learning continues across the GPU cluster. The RL training focuses on per-token intelligence across hundreds of thousands of tasks centered on multi-step software engineering.
Practical Applications and Tooling
Beyond raw coding, Grok 4.5 is integrated into several productivity workflows:
- Grok Build: Now the default model for the Grok Build CLI, capable of creating complex Excel models with web research and multi-sheet formulas.
- Office Integration: The model supports native PowerPoint shape creation for complex diagrams and professional prose generation in Word.
- ** uma Native App Development:** Users have reported success in building native iOS apps (SwiftUI and Metal) where other frontier models defaulted to simpler HTML/CSS implementations.
Community Insights and Technical Critique
Technical discussions on Hacker News highlight both the strengths and potential caveats of the Grok 4.5 release.
Performance and Value
Many users view the model as a "GLM-5.2 killer" due to its combination of Opus-class performance and Haiku-level pricing. One user reported that Grok 4.5 successfully debugged and fixed a Kubernetes manifest issue in under 30 minutes—a task that had previously taken hours of unsuccessful attempts with other models.
Data Integrity and Benchmarking
There are concerns regarding the validity of some benchmark scores. One community member pointed out that an earlier snapshot of the Cursor codebase was accidentally included in the training data, which may have artificially inflated results on CursorBench.
Privacy and Ethics
Critics have raised concerns about the data collection practices of Cursor, noting that the model's capabilities are derived from training on user interactions unless users explicitly opt-out, even for paying customers.
"It's very clear that typical, run-of-the-mill coding has been completely commoditized... novel technical solutions... could at least be capitalized on by simply being able to claim credit for it. But having them 'leaked' unwittingly via the training corpus... removes even that option."
Availability
Grok 4.5 is currently available via:
- Grok Build
- Cursor (all plans)
- SpaceXAI Console (API access)
Note: The model is not yet available in the EU, with expected availability in mid-July.
Sources
- HNGrok 4.5
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch