Skip to main content
Reasoning at scale

Gemini 4 Argon: Google's Next Frontier Model

Google unveiled Gemini 4 Argon, a frontier model for deep reasoning across complex workflows. Built initially for cybersecurity defenders through its Fairwind Program, the model demonstrates breakthrough capabilities in software engineering, enterprise work, and vulnerability detection. Already powering internal Google operations, Argon is proving its mettle through real-world impact—signaling a significant leap forward in tackling mission-critical tasks.
A large white number "4" glows against a blue gradient background with "Gemini 4 Argon" text and Google's multicolored star logo.
A large white number “4” glows against a blue gradient background with “Gemini 4 Argon” text and Google's multicolored star logo.

Google announced Gemini 4 Argon, a new frontier model rolling out to trusted cybersecurity defenders through its Fairwind Program. Built to handle deep reasoning across complex, long-horizon workflows, Argon delivers frontier performance in software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.

The company is taking a phased approach to safely release capabilities at this level, actively engaging with the U.S. government's voluntary pre-release model access process while gradually expanding access. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off the input token price.

Pricing Rate
Input tokens $2 per million
Output tokens $10 per million
Cached input tokens 95% off input token price

Internal Impact at Google

Gemini 4 Argon is already powering internal workflows across thousands of Google employees. In quantum algorithmic optimization, the model helped researchers improve resource efficiency by 40% above published baselines. Argon agents autonomously identified and applied memory optimizations across Google's data centers, freeing over 300 TiB of memory with estimated total savings of 500 TiB to 1 PiB.

The model is also managing large-scale codebase migrations, with Argon agents migrating C/C++ codebases to Rust—from tens of thousands of lines in core libraries like re2 and libgav1 to 800K+ lines for the Fuchsia OS Zircon kernel. For libgav1, Google's open source video decoder, Argon agents replaced 32K lines of SIMD code through profile-guided experiments, producing safe Rust that achieves 2.7x faster performance than the existing Rust port while maintaining identical video output.

Extended Reasoning Capabilities

a benchmark chart showing Gemini 4 Argon capabilities
a benchmark chart showing Gemini 4 Argon capabilities — Google

Gemini 4 Argon supports an industry-leading 1 million token output limit, up from the previous 64K tokens. This expanded capacity allows the model to think more deeply and generate hundreds of thousands of tokens in a single response, enabling more thorough problem-solving.

Enterprise and Coding Performance

Argon achieves state-of-the-art performance on DeepSWE v1.1 (77.9%), which measures real-world long-horizon software engineering tasks. The model leads on the Vals Index, which measures economic impact across finance, coding, legal, and tax work weighted by U.S. GDP contribution. It also ranks first on Zapier's AutomationBench with a score of 51.3%, which measures end-to-end execution across core business functions.

On domain-specific benchmarks, Argon leads performance on Vals Finance Agent v2 (multi-step financial research) and Harvey's Legal Agent Benchmark (legal research and drafting). For visual understanding tasks, the model achieves state-of-the-art performance on LVBench with a score of 91.7%, excelling at professional chart analysis, long video understanding, and document-based action taking.

Cybersecurity Capabilities

security vulnerabilities chart
security vulnerabilities chart — Google

Google trained Gemini 4 Argon specifically for cybersecurity defense, enabling it to autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and internal Google teams, the company is releasing Argon without cyber guardrails to leverage its full frontier-level capabilities.

Wiz is already using Argon through its Scan for Good initiative, a program protecting critical public infrastructure by finding and remediating high-risk exposures. In an early demonstration, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide—a risk previous frontier models had missed.

On CWE-bench v1, which evaluates vulnerability remediation ability, Argon ties for first place with a score of 68%. The model demonstrates significant improvements over 3.8 Flash Cyber in vulnerability discovery across Google's internal comprehensive benchmark, which spans 20 programming languages. On Wiz's internal black-box penetration testing benchmark, Argon outperforms 3.8 Flash Cyber in discovering attack surfaces, identifying vulnerabilities, and producing proof-of-concept evidence.

Frontier Safeguards

Gray Swan evaluation
Gray Swan evaluation — Google

Before broad availability, Google is strengthening safeguards across four areas:

DeepSwe evaluation chart
DeepSwe evaluation chart — Google
  • Defending against misuse: The model is designed to refuse harmful requests while preserving legitimate dual-use scientific research, per Google's Frontier Safety Framework. Safeguards include improved techniques to monitor internal model activations for misuse, tested by internal and external red teams using manual and automated attack methods.
  • Defending against prompt injection attacks: Argon is the company's most resilient model against indirect prompt injections, where malicious instructions hijack model behavior. Through automated red teaming and adversarial training, Argon leads on the Gray Swan's Indirect Prompt Injection benchmark.
  • Monitoring for misalignment: Google is deploying mitigations that monitor Argon's chain-of-thought and actions, stopping execution when necessary to prevent the model from stepping out of bounds. A similar system monitored training runs and sent alerts to an incident response team, with careful precautions against feeding findings back into training to avoid shaping reasoning to evade monitoring.
  • Hardening systems: Google is hardening sandboxed environments by isolating and sealing them before high-risk training or evaluations. The company commits to sharing these agent security practices with partners to improve ecosystem-wide security.

Availability

Argon will roll out to paid API customers and Google AI Ultra subscribers, with broader availability to developers, enterprises, and consumers coming soon.

Felipe Santos

“Artificial intelligence can process the world in milliseconds, but only the human heart can give meaning to every second lived” – Mr. Santos