Grok 4.6 release notes / what's new

xAI has released Grok 4.6, a model specifically optimized for long-running agents and complex interactive and visual work. The model is designed to sustain progress across multi-step tasks such as codebase analysis, deep research, and the development of polished applications.

Frontier Performance in Agentic Coding and Knowledge Work

Grok 4.6 achieves frontier-level intelligence on several agentic benchmarks, matching GPT-5.6 Sol on the Artificial Analysis (AA) Intelligence Index, which aggregates nine different benchmarks.

According to the provided evaluation data, Grok 4.6 shows significant improvements over Grok 4.5 across multiple metrics:

Benchmark Grok 4.6 High Grok 4.5 High GPT-5.6 Sol Max Fable 5 Max
AA Intelligence Index 61 56 61 62
GDPVal-AA v2 1753 1526 1728 1741
CursorBench v3.2 69.9% 66.7% 67.2% 70.5%
DeepSWE v1.1 65.9% 54% 73% 70%
FrontierCode v1.1 (Ext) 61.3% 56.6% 60.6% 63.6%
APEX-Agents 57.5% 47.1% 56.7% 59.2%
Terminal-Bench v3.0 26% 15.7% 34.6% 34.1%
APEX-SWE 56.4% 53.6% 58.8%
AA-Briefcase 1577 1313 1502 1574
Harvey LAB (Vals) 15.8% 12.9% 2.5% 11.3%

Technical Training Methodology

Grok 4.6 was developed using a multi-stage training process designed to strengthen its foundation for reasoning and technical execution.

Supplemental Training and Optimization

The model underwent a longer supplemental training run than its predecessor, Grok 4.5. This phase utilized an improved optimizer and training recipe, incorporating high-quality engineering data and curated model-generated data focused on advanced technical concepts and reasoning.

SFT and RL Refinement

xAI used Grok 4.5 to regenerate Supervised Fine-Tuning (SFT) trajectories across software engineering, STEM, knowledge work, and agent harnesses. These trajectories were filtered using model-based checks to remove problematic traces. Following SFT, the model was trained on a wide array of agentic Reinforcement Learning (RL) tasks, including:

  • General coding and knowledge work.
  • Domain-specific environments for web development and computer-aided design (CAD).
  • Kernel optimization.

Enhanced Agentic Capabilities and Visual Work

Grok 4.6 demonstrates an increased ability to handle long-trajectory projects, specifically in turning broad product ideas into functional first versions.

Iterative Development and Self-Verification

The model can research unfamiliar domains, structure applications, and implement core interactions. On longer tasks, Grok 4.6 exhibits emergent behaviors such as self-testing and verification, where the model checks its own work before proceeding to the next step.

Visual and Interactive Projects

Grok 4.6 provides stronger initial passes for visual projects compared to Grok 4.5, establishing the structure and visual language of an application in a single pass. This allows users to start with a substantial foundation and iterate rapidly.

Safety and Deployment

Safety safeguards in Grok 4.6 have been calibrated to match the model's expanded capabilities. The safety stack is designed to maintain security while maximizing utility for legitimate high-technical use cases, such as AI research, engineering design cycles, and vulnerability patching. Pre-deployment testing included xAI's widest-ever suite of capability and safeguard calibrations, supplemented by post-deployment and third-party testing.

Availability and Pricing

Grok 4.6 is available via the following platforms:

  • Integrations: Cursor and Grok Build (with 2x included usage for the first week).
  • API: Available via the SpaceXAI API, OpenRouter, Vercel, and Cloudflare.

Pricing Structure:

  • Standard: $2 per million input tokens / $6 per million output tokens.
  • Fast Variant: Twice the price of the standard variant.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch