Grok-2 Beta Release

xAI has announced the beta release of Grok-2 and Grok-2 mini, marking a significant advancement over Grok-1.5 in chat, coding, and reasoning capabilities. These models are currently available to — Premium and Premium+ users and will be released via an enterprise API later in August 2024.

Frontier Performance and Benchmarks

Grok-2 demonstrates state-of-the-art performance across multiple academic benchmarks, competing directly with other frontier models.

An early version of Grok-2, tested under the name "sus-column-r" on the LMSYS Chatbot Arena, outperformed both Claude 3.5 Sonnet and GPT-4 Turbo in terms of overall Elo score. Internal evaluations by AI Tutors focused on instruction following and factual accuracy, showing significant improvements in tool use and reasoning with retrieved content, specifically in identifying missing information and discarding irrelevant posts.

Technical Benchmark Results

Both Grok-2 and Grok-2 mini show substantial gains over Grok-1.5 across reasoning, math, science, and coding.

| Benchmark | Grok-1.5 | Grok-2 mini | Grok-2 | GPT-4 Turbo | Claude 3 Opus | Gemini Pro 1.5 | Llama 3 405B | GPT-4o | Claude 3.5 Sonnet | | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | | GPQA | 35.9% | 51.0% | 56.0% | 48.0% | 50.4% | 46.2% | 51.1% | 53.6% | 59.6% | | MMLU | 81.3% | 86.2% | 87.5% | 86.5% | 85.7% | 85.9% | 88.6% | 88.7% | 88.3% | | MMLU-Pro | 51.0% | 72.0% | 75.5% | 63.7% | 68.5% | 69.0% | 73.3% | 72.6% | 76.1% | | MATH | 50.6% | 73.0% | 76.1% | 72.6% | 60.1% | 67.7% | 73.8% | 76.6% | 71.1% | | HumanEval | 74.1% | 85.7% | 88.4% | 87.1% | 84.9% | 71.9% | 89.0% | 90.2% | 92.0% | | MMMU | 53.6% | 63.2% | 66.1% | 63.1% | 59.4% | 62.2% | 64.5% | 69.1% | 68.3% | | MathVista | 52.8% | 68.1% | 69.0% | 58.1% | 50.5% | 63.9% | — | 63.8% | 67.7% | | DocVQA | 85.6% | 93.2% | 93.6% | 87.2% | 89.3% | 93.1% | 92.2% | 92.8% | 95.2% |

Note: Grok-2 MMLU, MMLU-Pro, MMMU, and MathVista were evaluated using 0-shot CoT; MATH results are maj@1; HumanEval results are pass@1.

Vision and Multimodal Capabilities

Grok-2 excels in vision-based tasks, specifically in visual math reasoning and document-based question answering.

According to the benchmark data, Grok-2 achieved 69.0% on MathVista and 93.6% on DocVQA, delivering state-of-the-art performance in these areas. xAI also announced that a preview of multimodal understanding as a core part of the Grok experience on — and the API will be released soon.

Integration and Availability

Grok-2 and Grok-2 mini are integrated into the — platform with a redesigned interface and new feature sets.

— Premium and Premium+ users can access these models via the Grok tab in the — app. Grok-2 serves as the state-of-the-art assistant for text and vision understanding, while Grok-2 mini provides a balance of speed and quality. Additionally, xAI is experimenting with the FLUX.1 model from Black Forest Labs to expand Grok's capabilities on the platform.

Enterprise API Platform

Developers will gain access to Grok-2 and Grok-2 mini through a new enterprise API platform later in August 2024. The API infrastructure features:

  • Multi-region inference deployments for low-latency global access.
  • Enhanced security, including mandatory multi-factor authentication (Yubikey, Apple TouchID, or TOTP).
  • Management tools, including a management API for integrating team, user, and billing management into in-house services, alongside rich traffic statistics and billing analytics.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch