METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Google DeepMind unveils Gemini 4 Argon

Google released its new frontier model to cyber defenders first. It raised the output token limit to 1 million and set no date for general availability.

Google DeepMind unveils Gemini 4 Argon

Image: METAL

Summary

  • Google DeepMind unveiled its new frontier model, Gemini 4 Argon, on September 30 and is rolling it out first to trusted cyber defenders in its Fairwind Program.
  • Argon scored 77.9% on DeepSWE v1.1, 68.9% on the Vals Index and 51.3% on AutomationBench, but other models led on FrontierSWE v2 and Terminal-bench 4.0.
  • Google raised the output token limit from 64K to 1M and said it will release Argon broadly after strengthening four safeguards: misuse defense, prompt injection defense, misalignment monitoring and sandbox hardening.

Google DeepMind unveiled its new frontier model, Gemini 4 Argon, on September 30 (local time). The model targets long, complex workflows such as real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. Its first users are not the general public but trusted cyber defenders taking part in Google's Fairwind Program.

The rollout order is itself the first message of the announcement. "Safely releasing frontier capabilities at this level requires a phased approach," Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect at Google, said in the announcement. Google is taking part in the U.S. government's voluntary process for pre-release model access and plans to refine its guardrails with feedback from early testers before expanding to developers, enterprises and consumers. General availability will start with paid API customers and Google AI Ultra subscribers.

Argon is already running internal work at Google. According to Google, thousands of employees have highlighted the model's strengths in specialized coding tasks, deeper research and writing quality. Quantum computing researchers used Argon to reduce the spacetime resources (the product of qubits and gates) of subroutines that bottleneck important applications, and in one example it beat the published baseline by 40% in a matter of minutes. A team of Argon agents analyzed fleet-wide profiling telemetry across Google's data centers, autonomously identified and applied memory optimizations, and freed more than 300 TiB of memory once rolled out. The company estimates total savings of 500 TiB to 1 PiB.

The code migration work is large as well. Argon agents are moving C/C++ code across Google to Rust, scaling from tens of thousands of lines in core libraries such as re2 and libgav1 up to more than 800,000 lines in Fuchsia's Zircon kernel. In libgav1, Google's open source video decoding software, the agents replaced 32,000 lines of SIMD code in an existing Rust port. By running many rounds of profile-guided experiments and studying the compiler's output, they produced safe Rust code the compiler could vectorize automatically, and the result was a memory-safe decoder that runs 2.7 times faster than the Rust port with identical video output. Google added that such large-scale rewrites go through automated and manual auditing, emulation testing and review before reaching production.

There is also a design change meant to sustain long tasks. Google raised Argon's output token limit from the previous 64,000 (64K) to 1 million (1M). Giving the model room to think deeply and generate hundreds of thousands of tokens in a single trajectory lets it solve hard problems in one go, according to the company.

In the comparison table Google published, Argon scored 77.9% on DeepSWE v1.1, which measures long-horizon software engineering tasks. In the same table GPT-6 Astra scored 74.1%, Claude Opus 5.5 74.2% and Claude Fable 5.1 67.4%. Argon ranked first on the Vals Index, which weights finance, coding, legal and tax work by contribution to U.S. GDP, with 68.9%, and on Zapier's business automation benchmark AutomationBench it scored 51.3%, ahead of Claude Opus 5.5 (42.5%) and GPT-6 Astra (41.4%). On Harvey's Legal Agent Benchmark it scored 19.6%, well above the single-digit scores of the other three models, and on LVBench, which measures long video understanding, it scored 91.7%.

The same table also listed benchmarks where Argon did not take first place. On FrontierSWE v2, GPT-6 Astra led with 65.5% against Argon's 55.0%, and on Terminal-bench 4.0, Claude Opus 5.5 scored highest at 66.4%. Argon scored 57.4%. Other models also topped PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0.

Gemini 4 Argon과 GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5의 지식 노동·코딩·과학·장문맥·컴퓨터 사용·멀티모달·사이버보안 벤치마크 점수를 비교한 구글 공식 표

Cybersecurity is the field that set this model's release path. Google trained Argon to autonomously find, validate and patch critical software vulnerabilities, and it is providing the model without cyber guardrails to trusted defenders and internal teams. Security company Wiz is using Argon in its Scan for Good initiative, which protects critical public infrastructure for free. According to Google, Argon uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide. It was a risk that previous frontier models had missed. On CWE-bench v1, which measures the ability to remediate vulnerabilities, Argon tied for first at 68% with Grok 4.7 and GPT-6 Astra, and on Google's internal vulnerability benchmark it found a wide range of exposures across code spanning 20 programming languages.

The Fairwind Program is a limited-access channel that opens Google's cyber defense tools to governments and trusted partners. According to reports, the program launched on September 2, has more than 650 participating partners worldwide, and its first offering paired Gemini 3.8 Flash Cyber with the CodeMender harness. METAL previously reported that Google opened cyber defense AI to governments and enterprises through the Fairwind Program. Argon is the first frontier model to enter that channel, about a month later.

취약점 수정 능력을 재는 CWE-bench v1 순위표. Grok 4.7, Gemini 4 Argon, GPT-6 Astra가 68%로 공동 1위다

Before broad availability, Google said it is strengthening safeguards in four areas. The first is misuse defense. The model is designed to refuse requests for cyber and chemical, biological, radiological and nuclear (CBRN) attacks while preserving legitimate dual-use scientific research, and Google improved its techniques for monitoring the model's internal activations to spot misuse. The second is defense against indirect prompt injection, and Google claimed Argon showed the leading robustness on Gray Swan's Indirect Prompt Injection benchmark. The third is misalignment monitoring. Argon's chain of thought and actions are watched, and execution is stopped if it goes beyond the user's intentions. The fourth is system hardening, isolating and sealing sandboxed environments before high-risk training and evaluations.

Misalignment monitoring also ran during training. Google explained that a similar system monitored training runs and sent alerts to a dedicated incident response team, and that it took care not to feed the findings back into training. The aim was to avoid the risk of the model shaping its reasoning to evade monitoring. "We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities," Kavukcuoglu said. The argument is that a model's thoughts must remain readable to identify and diagnose misalignment.

The full Google announcement reviewed by METAL gives no date for general availability. It only says Argon will be made available to developers, enterprises and consumers "as soon as possible." What was decided before the top scores in the comparison table was who gets it first, and Google put defenders, not attackers, at the front of that line. The measure of the frontier model race is shifting from performance rankings to release order.

Comments