Grok 3 Beta release notes / what's new
xAI has introduced Grok 3 Beta, a model family that integrates extensive pretraining knowledge with advanced reasoning capabilities. Trained on the Colossus supercluster using 10x the compute of previous state-of-the-art models, Grok 3 demonstrates significant improvements in mathematics, coding, world knowledge, and instruction-following tasks.
Advanced Reasoning via Test-Time Compute
Grok 3 introduces specialized reasoning models, Grok 3 (Think) and Grok 3 mini (Think), which utilize large-scale reinforcement learning (RL) to refine their chain-of-thought processes. These models can employ test-time compute, spending from a few seconds to several minutes to backtrack, correct errors, and explore multiple problem-solving strategies before delivering a final answer.
Performance benchmarks for the reasoning models include:
- AIME 2025: Grok 3 (Think) achieved 93.3% using the highest level of test-time compute (cons@64).
- GPQA (Graduate-level expert reasoning): Grok 3 (Think) attained 84.6%.
- LiveCodeBench: Grok 3 (Think) achieved 79.4% for code generation and problem-solving.
- AIME 2024: Grok 3 mini reached 95.8%.
- LiveCodeBench: Grok 3 mini reached 80.4% for STEM tasks.
Users can access these capabilities via a "Think" button, which provides full transparency by allowing users to inspect the model's internal reasoning process.
Pretraining and General Performance
When reasoning is disabled, Grok 3 provides high-quality, instant responses and achieves state-of-the-art results among non-reasoning models. It features a context window of 1 million tokens, which is eight times larger than previous xAI models, enhancing its ability to process extensive documents and maintain instruction-following accuracy.
Comparative performance on key benchmarks (Grok 3 Beta vs. competitors):
| Benchmark | Grok 3 Beta | Gemini 2.0 | DeepSeek-V3 | GPT 4o | Claude 3.5 Sonnet | | :--- | :--- | :--- | :--- | :--- | :--- | :--- | | AIME’24 | 52.2% | — | 39.2% | 9.3% | 16.0% | | GPQA | 75.4% | 64.7% | 59.1% | 53.6% | 65.0% | | LCB | 57.0% | 36.0% | 33.1% | 32.3% | 40.2% | | MMLU-pro | 79.9% | 79.1% | 75.9% | 72.6% | 78.0% | | LOFT (128k) | 83.3% | 75.6% | — | 78.0% | 69.9% | | MMMU | 73.2% | 72.7% | — | 69.1% | 70.4% | | EgoSchema | 74.5% | 71.9% | — | 72.2% | — |
An early version of the model, codenamed chocolate, previously topped the LMArena Chatbot Arena leaderboard with an Elo score of 1402.
Grok Agents and DeepSearch
To extend the model's capabilities into real-world interaction, xAI is introducing Grok Agents, which combine reasoning with tool use, including internet access and code interpreters.
The first agent, DeepSearch, is designed to synthesize information from the entire corpus of human knowledge, reason through conflicting facts, and generate comprehensive reports. It is intended for use cases ranging from real-time news synthesis to in-depth scientific research.
Availability and Roadmap
- User Access: Grok 3 is available to ° Premium and Premium+ users on X and Grok.com. Premium+ users have immediate access to
ThinkandDeepSearchwith higher usage limits. - API Access: Grok 3 and Grok 3 mini, including both standard and reasoning versions, will be released via the API platform in the coming weeks.
DeepSearchwill be available to Enterprise partners via API. - Future Development: Training is ongoing. Future updates to the Enterprise API will include tool use, code execution, and advanced agent capabilities, with a focus on scalable oversight and adversarial robustness.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch