Grok 4 Fast release notes / what's new

xAI has announced Grok 4 Fast, a cost-efficient reasoning model designed to provide frontier-level performance across enterprise and consumer domains with high token efficiency. The model is built on learnings from Grok 4 and aims to make high-quality reasoning accessible to a broader user base of developers and consumers.

High Intelligence Density and Cost Efficiency

Grok 4 Fast achieves frontier-level performance while significantly reducing the cost of intelligence. By utilizing large-scale reinforcement learning to maximize "intelligence density," the model achieves comparable performance to Grok 4 on benchmarks but uses 40% fewer thinking tokens on average.

This increase in token efficiency, combined with lower per-token pricing, results in a 98% reduction in the price required to achieve the same performance as Grok 4 on frontier benchmarks. According to an independent review by Artificial Analysis, Grok 4 Fast maintains a state-of-the-art (SOTA) price-to-intelligence ratio on the Artificial Analysis Intelligence Index.

Reasoning Benchmarks

Benchmark pass@1 Grok 4 Fast Grok 4 Grok 3 Mini (High) GPT-5 (High) GPT-5 Mini (High)
GPQA Diamond 85.7% 87.5% 79.0% 85.7% 82.3%
AIME 2025 (no tools) 92.0% 91.7% 83.0% 94.6% 91.1%
HMMT 2025 (no tools) 93.3% 90.0% 74.0% 93.3% 87.8%
HLE (no tools) 20.0% 25.4% 11.0% 24.8% 16.7%
LiveCodeBench (Jan-May) 80.0% 79.0% 70.0% 86.8% 77.4%

Native Tool Use and Agentic Search

Grok 4 Fast is trained end-to-end with tool-use reinforcement learning (RL), enabling it to effectively decide when to invoke tools such as web browsing or code execution. The model features agentic search capabilities that allow it to browse the web and X in real-time, ingest media (including images and videos on X), and synthesize findings.

Search and Research Benchmarks

Benchmark pass@1 Grok 4 Fast Grok 4 Grok 3 (No Reasoning)
BrowseComp 44.9% 43.0%
SimpleQA 95.0% 94.0% 82.0%
Reka Research Eval 66.0% 58.0% 37.0%
BrowseComp (zh) 51.2% 45.0% 10.8%
X Bench Deepsearch (zh) 74.0% 66.0% 27.0%
X Browse* 58.0% 53.2% 20.8%

*X Browse is an internal benchmark evaluating multihop search and browsing capabilities on X.

General Domain Performance and LMArena Rankings

In LMArena's Search Arena, grok-4-fast-search (code name: menlo) ranks #1 with an Elo of 1163, leading o3-search by a margin of 17. In the Text Arena, grok-4-fast (code name: tahoe) ranks #8, performing on par with grok-4-0709 and outperforming other models in its weight class, which typically rank 18th or below.

Unified Architecture for Reasoning Modes

Grok 4 Fast introduces a unified architecture where both reasoning (long chain-of-thought) and non-reasoning (quick responses) are handled by the same model weights and steered via system prompts. This unification reduces end-to-end latency and token costs compared to previous iterations that required separate models for different reasoning modes.

Availability and API Pricing

Grok 4 Fast is available to all users, including free users, via grok.com in Fast and Auto modes. For developers, it is available via the xAI API as two models: grok-4-fast-reasoning and grok-4-fast-non-reasoning, both featuring a 2M token context window.

xAI API Pricing

Token Type <128k tokens …128k tokens
Input tokens $0.20 / 1M $0.40 / 1M
Output tokens $0.50 / 1M $1.00 / 1M
Cached input tokens $0.05 / 1M

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch