Google announced Gemini 4 Argon, a new frontier model rolling out to trusted cybersecurity defenders through its Fairwind Program. Built to handle deep reasoning across complex, long-horizon workflows, Argon delivers frontier performance in software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
The company is taking a phased approach to safely release capabilities at this level, actively engaging with the U.S. government's voluntary pre-release model access process while gradually expanding access. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off the input token price.
| Pricing | Rate |
|---|---|
| Input tokens | $2 per million |
| Output tokens | $10 per million |
| Cached input tokens | 95% off input token price |
Internal Impact at Google
Gemini 4 Argon is already powering internal workflows across thousands of Google employees. In quantum algorithmic optimization, the model helped researchers improve resource efficiency by 40% above published baselines. Argon agents autonomously identified and applied memory optimizations across Google's data centers, freeing over 300 TiB of memory with estimated total savings of 500 TiB to 1 PiB.
The model is also managing large-scale codebase migrations, with Argon agents migrating C/C++ codebases to Rust—from tens of thousands of lines in core libraries like re2 and libgav1 to 800K+ lines for the Fuchsia OS Zircon kernel. For libgav1, Google's open source video decoder, Argon agents replaced 32K lines of SIMD code through profile-guided experiments, producing safe Rust that achieves 2.7x faster performance than the existing Rust port while maintaining identical video output.
Extended Reasoning Capabilities

Gemini 4 Argon supports an industry-leading 1 million token output limit, up from the previous 64K tokens. This expanded capacity allows the model to think more deeply and generate hundreds of thousands of tokens in a single response, enabling more thorough problem-solving.
Enterprise and Coding Performance
Argon achieves state-of-the-art performance on DeepSWE v1.1 (77.9%), which measures real-world long-horizon software engineering tasks. The model leads on the Vals Index, which measures economic impact across finance, coding, legal, and tax work weighted by U.S. GDP contribution. It also ranks first on Zapier's AutomationBench with a score of 51.3%, which measures end-to-end execution across core business functions.
On domain-specific benchmarks, Argon leads performance on Vals Finance Agent v2 (multi-step financial research) and Harvey's Legal Agent Benchmark (legal research and drafting). For visual understanding tasks, the model achieves state-of-the-art performance on LVBench with a score of 91.7%, excelling at professional chart analysis, long video understanding, and document-based action taking.
Cybersecurity Capabilities

Google trained Gemini 4 Argon specifically for cybersecurity defense, enabling it to autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and internal Google teams, the company is releasing Argon without cyber guardrails to leverage its full frontier-level capabilities.
Wiz is already using Argon through its Scan for Good initiative, a program protecting critical public infrastructure by finding and remediating high-risk exposures. In an early demonstration, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide—a risk previous frontier models had missed.
On CWE-bench v1, which evaluates vulnerability remediation ability, Argon ties for first place with a score of 68%. The model demonstrates significant improvements over 3.8 Flash Cyber in vulnerability discovery across Google's internal comprehensive benchmark, which spans 20 programming languages. On Wiz's internal black-box penetration testing benchmark, Argon outperforms 3.8 Flash Cyber in discovering attack surfaces, identifying vulnerabilities, and producing proof-of-concept evidence.
Frontier Safeguards

Before broad availability, Google is strengthening safeguards across four areas:

- Defending against misuse: The model is designed to refuse harmful requests while preserving legitimate dual-use scientific research, per Google's Frontier Safety Framework. Safeguards include improved techniques to monitor internal model activations for misuse, tested by internal and external red teams using manual and automated attack methods.
- Defending against prompt injection attacks: Argon is the company's most resilient model against indirect prompt injections, where malicious instructions hijack model behavior. Through automated red teaming and adversarial training, Argon leads on the Gray Swan's Indirect Prompt Injection benchmark.
- Monitoring for misalignment: Google is deploying mitigations that monitor Argon's chain-of-thought and actions, stopping execution when necessary to prevent the model from stepping out of bounds. A similar system monitored training runs and sent alerts to an incident response team, with careful precautions against feeding findings back into training to avoid shaping reasoning to evade monitoring.
- Hardening systems: Google is hardening sandboxed environments by isolating and sealing them before high-risk training or evaluations. The company commits to sharing these agent security practices with partners to improve ecosystem-wide security.
Availability
Argon will roll out to paid API customers and Google AI Ultra subscribers, with broader availability to developers, enterprises, and consumers coming soon.
Most Read
