SpaceXAI Grok 4.6 Release: Intelligence Frontier and Cost Efficiency

Grok 4.6 joins the intelligence frontier with high cost-efficiency

SpaceXAI's Grok 4.6 has reached a score of 61 on the Artificial Analysis Intelligence Index, placing it in line with GPT-5.6 Sol (max) and trailing only Anthropic's Claude Opus 5 (63) and Claude Fable 5 (62). This release marks a significant jump in intelligence, gaining 5 points over Grok 4.5 and 23 points over Grok 4.3.

Agentic performance and real-world knowledge work

Grok 4.6 demonstrates its strongest capabilities in agentic tasks rather than static reasoning. It achieves a GDPval-AA v2 Elo of 1753, placing it behind only Claude Opus 5 and making it statistically indistinguishable from Claude Fable 5 and Qwen3.8 Max.

Key performance metrics include:

  • $\tau^3$-Banking: Scores 50.7%, ranking it among the top two models alongside Qwen3.8 Max (51.3%) for multi-turn customer service with tool use.
  • Terminal-Bench v2.1: Scores 88.4%, performing on par with leading models for terminal-based software tasks.
  • AA-Briefcase: In this private benchmark for long-horizon agentic knowledge work, Grok 4.6 earns an Elo of 1577 (Fable 5-tier).

Notably, Grok 4.6 is highly turn-efficient. On AA-Briefcase, it completes tasks in approximately 53 turns and uses ~0.5B input tokens on average, compared to ~103 turns and ~2.0B input tokens for Claude Opus 5 (max).

Pricing and cost-per-task advantage

Grok 4.6 maintains the same headline pricing as Grok 4.5, which is unusual for frontier-level intelligence gains. It is priced at $2 per 1M input tokens and $6 per 1M output tokens, which is 60%+ lower than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30).

According to Artificial Analysis, the cost per task for Grok 4.6 is $0.84, placing it on the Intelligence vs. Cost per Task Pareto frontier. While headline pricing is flat, cache hit pricing has increased from $0.3 per 1M tokens in Grok 4.5 to $0.5 per 1M tokens in Grok 4.6.

Technical specifications

  • Context Window: 500k tokens (unchanged from Grok 4.5).
  • Pricing: $2/$6 per 1M input/output tokens.
  • Cache Hits: $0.5 per 1M tokens.

User perspectives and community insights

Community discussion reveals a divide in user experience and perceptions of the model's utility in professional workflows:

  • Coding and Productivity: Some users report that Grok 4.6, particularly when used via the Grok build CLI or within Cursor, is faster and more interactive than Claude Code. One user noted that Grok's communication style is more concise, avoiding "walls of text" and allowing for more steering during sessions.
  • Infrastructure Advantage: Observers suggest that SpaceXAI's ownership of its own compute and data centers may provide a long-term advantage in reducing token costs and improving tool integration.
  • Tool Use: Users have observed an improvement in Grok 4.6's propensity to verify tasks visually (e.g., taking screenshots), bringing it closer to the capabilities of the Claude Opus 5 and Fable families.
  • Criticisms: Some users remain skeptical of the model's adoption for coding, while others expressed concerns regarding the environmental impact of Grok's training due to the use of portable gas generators.

"Grok 4.6 though is no longer 'smart and fast'. It's about as fast as Sol though. A big improvement I noticed in 4.6 was tool use for verification... Now Grok is probably right behind them, perhaps tied with Sol on propensity to verify visually."

"I've been using grok 4.5 with grok build soon after it came out and dropped claude... It doesn't give me a wall of text, tells me what I need to know and I'll make the actual decisions."

"In my experience in heavy coding sessions most pricing is just cache read and cache write like 80% of my token bill."

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch