xAI Grok 4.1 Release Notes
xAI has released Grok 4.1, which improves real-world usability through enhanced emotional intelligence, creative writing capabilities, and a reduction in factual hallucinations. The model is now available to all users on grok.com, — and the associated iOS and Android applications.
State-of-the-Art General Capability
Grok 4.1 establishes a new benchmark in blind human preference evaluations, specifically on the LMArena Text Leaderboard. The model exists in two primary configurations:
- Grok 4.1 Thinking (code name:
quasarflux): This reasoning model holds the #1 overall position with 1483 Elo, leading the highest non-xAI model by 31 points. - Grok 4.1 Non-Reasoning (code name:
tensor): This version uses no thinking tokens for immediate responses and ranks #2 overall with 1465 Elo. Notably, the non-thinking version of Grok 4.1 surpasses every other model’s full-reasoning configuration on the public leaderboard.
Compared to the previous production model, Grok 4.1 is preferred 64.78% of the time in blind pairwise evaluations conducted during a silent rollout from November 1 to November 14, 2025.
Technical Optimization and Methodology
xAI utilized the same large-scale reinforcement learning (RL) infrastructure that powered Grok 4 to optimize the model's style, personality, helpfulness, and alignment. To handle non-verifiable reward signals, xAI developed new methods using frontier agentic reasoning models as reward models to autonomously evaluate and iterate on responses at scale.
Emotional Intelligence and Creative Writing
Grok 4.1 was evaluated using EQ-Bench3 and Creative Writing v3 benchmarks to measure interpersonal ability and creative output:
- EQ-Bench3: This LLM-judged test evaluates active emotional intelligence, empathy, and interpersonal skills across 45 roleplay scenarios. The evaluations were conducted using the default sampling parameters and Claude Sonnet 3.7 as the prescribed judge.
- Creative Writing v3: This benchmark measures performance across 32 distinct writing prompts over three iterations, using both rubrics and model battle normalized Elo.
Reduction in Factual Hallucinations
xAI focused on reducing factual hallucinations for information-seeking prompts during the post-training of Grok 4.1. This is particularly aimed at fast, non-reasoning models that use search tools but may suffer from limited reasoning depth.
Evaluations were performed on a stratified sample of real-world information-seeking queries from production traffic and the FActScore public benchmark (500 biography questions). The hallucination rate is defined as the macro-average of the percentage of atomic claims with major or minor errors over model responses. These evaluations were conducted using the non-reasoning model equipped with web search tools.
Sources
- OriginalGrok 4.1
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch