xAI Grok 4.7 Release Notes

xAI has released Grok 4.7, a model specifically optimized for complex coding and professional knowledge work. The model improves upon its predecessor, Grok 4.6, by utilizing a larger base model and a training regimen focused on tasks that require several hours to complete.

Technical Improvements and Training

Grok 4.7 is built on a larger base model than Grok 4.6 and was trained using an extended reinforcement learning run. This training focused on a harder mix of tasks, specifically weighting problems that take many hours to complete, which enables the model to better verify its own work and manage longer contexts.

Additionally, xAI trained Grok 4.7 to natively understand the Grok Bot harness, improving its performance in general knowledge work and conversational tasks.

Performance Benchmarks

Grok 4.7 demonstrates significant improvements over Grok 4.6 across several professional and technical benchmarks. It is positioned as a highly competitive option in terms of price-performance, particularly on CursorBench 4.0 for long-running coding tasks.

Software and Electrical Engineering

  • CursorBench 4.0: Grok 4.7 achieved a score of 46.3%, compared to 40.4% for Grok 4.6.
  • DeepSWE v1.1: Grok 4.7 reached 71.0% (high effort), outperforming Grok 4.6 (65.2%).
  • EEBench: Grok 4.7 scored 64.0%, significantly higher than Grok 4.6 (53.0%) and GPT-5.6 Sol Max (39.4%).

Professional Knowledge Work

  • AA Briefcase v1.1: Grok 4.7 scored 1,657, improving upon Grok 4.6 (1,546).
  • Harvey Legal Agent Benchmark: Grok 4.7 scored 19.6%, a substantial increase over Grok 4.6 (15.8%) and far exceeding GPT-5.6 Sol Max (2.5%) and Fable 5.1 Max (6.7%).
  • HealthBench Professional: Grok 4.7 scored 56.7%, compared to 48.5% for Grok 4.6.
  • Terminal-Bench 4.0: Grok 4.7 achieved 38.0%, compared to 20.3% for Grok 4.6.

Safety and Cybersecurity

Grok 4.7 implements an entirely new safeguard stack, making it the strongest model xAI has tested for jailbreak resistance and refusals.

Biosafety and Cyber Defense

  • Biosafety: The model tops LatchBio’s biosafety benchmark with a score of 62.4%.
  • Cybersecurity: On HackerBench v0.3, Grok 4.7 allows only 3.3% of risky dual-use prompts to pass, while maintaining low refusal rates for legitimate security work.

To support defense research, xAI has provided invite-only access to the model's red-team capabilities to select cybersecurity partners.

Pricing and Availability

Grok 4.7 is available via the Grok API, Grok Build, Cursor, third-party coding harnesses, model routers, and cloud platforms.

Pricing Structure:

  • Standard: $2 per million input tokens and $6 per million output tokens.
  • Fast Variant: Twice the output speed at twice the price of the standard variant.

Sources