Gemini 4 Argon Frontier Model Announcement

TL;DR

Gemini 4 Argon is Google DeepMind’s newest frontier model, offering a 1 M‑token context window and state‑of‑the‑art performance in software engineering, enterprise knowledge work, and defensive cybersecurity; it is initially available to trusted cyber defenders through the Fairwind Program.


Introducing Gemini 4 Argon

Gemini 4 Argon is positioned as a “frontier” model built to sustain deep reasoning across long‑horizon, complex workflows. The announcement emphasizes three target domains:

  • Real‑world software engineering
  • Enterprise knowledge work such as legal and finance
  • Cybersecurity defense

The model is being released in a phased manner, starting with a trusted‑defender cohort while DeepMind engages with the U.S. government’s voluntary pre‑release access process and iterates on safety guardrails.

Pricing

  • Input tokens: $2 per million
  • Output tokens: $10 per million
  • Cached input tokens: 95 % discount on the input price

Impact on Google’s Internal Workflows

Argon is already integrated into internal Google processes, where thousands of employees report gains in specialized coding, research depth, and writing quality. Highlighted use cases include:

  • Quantum algorithmic optimization – reduced spacetime resources (qubits × gates) by 40 % compared with the published baseline, within minutes.
  • Memory efficiency – analysis of fleet‑wide telemetry freed >300 TiB of memory, with projected total savings of 500 TiB – 1 PiB.
  • Large‑scale code migrations – Argon agents are converting C/C++ codebases to Rust, scaling from tens of thousands of lines (e.g., re2, libgav1) to >800 K lines for the Fuchsia OS Zircon kernel. For libgav1, Argon replaced 32 K lines of SIMD code with safe Rust, achieving a 2.7× speedup over the prior Rust port while preserving video output.

Extending Context Length to 1 M Tokens

To handle longer, more intricate tasks, Argon’s maximum output length has been increased from 64 K to 1 million tokens, an industry‑leading figure. This expanded context enables the model to generate extensive reasoning chains in a single pass, improving performance on tasks that previously required multiple interactions.


Enterprise‑Level Coding and Knowledge‑Work Performance

Argon’s multimodal reasoning and long‑context abilities translate into top‑tier results on several benchmark suites:

  • DeepSWE v1.1 – 77.9 % accuracy, setting a new state‑of‑the‑art for long‑horizon software‑engineering tasks.
  • Vals Index – Argon leads the composite economic‑impact benchmark that weights finance, coding, legal, and tax work by U.S. GDP contribution.
  • Vals Finance Agent v2 – best performance on multi‑step financial research.
  • Harvey’s Legal Agent Benchmark – highest scores on legal research and drafting.
  • AutomationBench (Zapier) – #1 ranking with a score of 51.3 % for end‑to‑end business‑process automation.
  • LVBench – state‑of‑the‑art long‑video understanding with a score of 91.7 %.

These results demonstrate Argon’s versatility across text‑heavy, multimodal, and multi‑step enterprise workflows.


Defensive Cybersecurity Capabilities

Argon has been fine‑tuned for cybersecurity defense and is being released to trusted defenders without cyber‑specific guardrails to expose its full capabilities. Notable achievements include:

  • Vulnerability discovery – on Google’s internal benchmark, Argon identified a broad spectrum of exposures across codebases in 20 programming languages.
  • Wiz Scan‑for‑Good partnership – Argon uncovered a critical vulnerability in global healthcare software that exposed personal data, a risk missed by prior frontier models.
  • CWE‑bench v1 – tied for first place with a 68 % remediation score, improving on the 3.8 Flash Cyber model’s performance on CWE‑bench v0.
  • Wiz internal penetration testing – outperformed 3.8 Flash Cyber in attack‑surface mapping, vulnerability identification, and proof‑of‑concept generation.

Strengthening Frontier Safeguards

Before a broad release, DeepMind is enhancing four safeguard categories:

  1. Misuse prevention – Argon refuses requests that facilitate cyber, chemical, biological, radiological, or nuclear attacks while allowing legitimate dual‑use research. Robustness testing involved internal and external red‑team exercises and monitoring of internal activations (see arXiv:2601.11516).
  2. Prompt‑injection resilience – Argon shows leading robustness on the Gray Swan Indirect Prompt Injection benchmark, achieved through automated red‑teaming and adversarial training.
  3. Misalignment monitoring – chain‑of‑thought and action monitoring stops execution when the model deviates from user intent; alerts are routed to a dedicated incident‑response team.
  4. System hardening – sandboxed environments are isolated and sealed before high‑risk training or evaluation, following DeepMind’s agent‑control roadmap.

DeepMind urges the broader AI community to adopt reasoning‑transparency practices to keep model thoughts observable during alignment work.


Rollout Timeline

Gemini 4 Argon will first be available to the Fairwind Program’s trusted cyber‑defender cohort. Subsequent phases will extend access to developers, enterprises, and consumers, beginning with paid API customers and Google AI Ultra subscribers.


Implications

  • Productivity boost – The 1 M‑token context and advanced reasoning allow developers and knowledge workers to tackle tasks that previously required extensive human effort or multiple model calls.
  • Security impact – By automating vulnerability discovery and patch generation, Argon could accelerate remediation cycles across the software supply chain.
  • Safety considerations – The phased release and extensive safeguard development illustrate DeepMind’s commitment to responsible deployment of frontier capabilities.

References


This article is based on the official DeepMind blog post “Gemini 4 Argon: our next era of frontier intelligence” published on 30 September 2026.

Sources