AI news, models and products, with sources中文
NewsMajor Models30 Sept 2026

Google announces Gemini 4 Argon

Google's new frontier model is built for long, complex professional work and can output up to 1 million tokens. It is first available to trusted testers and cyber defenders, then to Google AI Ultra subscribers and paid API customers.

Glass entrance of Google's London office with the Google name above the doors
The entrance to Google's office at 6 Pancras Square in London, where Google DeepMind is also based, in 2017. File photo (cropped). Photo: Gciriani / Wikimedia Commons, CC BY-SA 4.0

What happened

On 30 September Google announced Gemini 4 Argon. The announcement, signed by Koray Kavukcuoglu, senior vice-president of Google DeepMind and Google’s chief AI architect, calls it “our next era of frontier intelligence”.

According to Reuters, a Google spokesperson said Argon is larger than the company’s previous “Pro” models: “It’s our most performant model yet built for complex workloads, and we see it comparable to frontier models like (OpenAI’s) Astra and (Anthropic’s) Opus on key coding and cyber benchmarks.” Argon is also the first model under a new naming scheme, replacing names such as Gemini 3 Pro and 3.1 Pro.

Unusually, ordinary users can’t use it yet:

Argon's phased roll-out
  1. 1NowAvailable to a group of trusted cyber defenders and testers through Google's Fairwind Program.
  2. 2MeanwhileGoogle says it is taking part in the US government's voluntary process for pre-release model access, and refining safeguards based on early testers' feedback.
  3. 3NextPaid API customers and Google AI Ultra subscribers first, then developers, businesses and consumers more broadly.
  4. 4WhenGoogle says only "as soon as possible" and "rolling out soon"; there is no public release date.

Sources: Google, “Gemini 4 Argon: our next era of frontier intelligence”, 30 Sep 2026; Reuters

Background: a late flagship

Argon arrived later than planned. According to Reuters:

  • Google’s chief executive Sundar Pichai had originally said Gemini 3.5 Pro would come out in June; a spokesperson said Google no longer plans to release it.
  • In the meantime Google overhauled DeepMind: founder and chief executive Demis Hassabis stepped aside, and several leaders of the Gemini models left the company.
  • The months of delay left Google behind Anthropic and OpenAI, which kept releasing new top models, and Google shifted its messaging from cutting-edge capability to cost advantage.

The benchmarking firm Artificial Analysis noted that Argon is Google DeepMind’s first proprietary model above the “Flash” class in more than seven months, and scores 23 points higher on its Intelligence Index than Google’s previous non-Flash model, Gemini 3.1 Pro Preview.

Gemini already has a very large audience: according to TechCrunch, Google announced in August that the Gemini app had more than a billion monthly users. (For the previous generation, see Google releases Gemini 3.)

The headline change: 1 million tokens of output

How Google uses it internally

Google says thousands of its employees already use Argon in their daily work. Its examples, all as described by Google:

  • Quantum computing. Helping researchers optimise key parts of quantum algorithms; in one case it beat the published baseline by 40% in minutes.
  • Saving memory. A team of Argon agents analysed performance data from across Google’s data centres and found and applied memory optimisations on their own, freeing more than 300 TiB of memory, with an estimated 500 TiB to 1 PiB saved in total.
  • Rewriting old code. Argon agents are moving Google’s C and C++ code to the safer Rust language, from core libraries of tens of thousands of lines up to the 800,000-plus lines of the Zircon kernel in the Fuchsia operating system. Google says these rewrites go through rigorous automated and human review before going live.
  • Faster video decoding. In libgav1, Google’s open-source video decoder, Argon ran many rounds of experiments to replace 32,000 lines of low-level code, producing a decoder 2.7 times faster than the previous Rust version, with identical output.

How it compares

Google published a table comparing Argon with OpenAI’s GPT-6 Astra, and Anthropic’s Claude Fable 5.1 and Claude Opus 5.5. The result it highlights is DeepSWE v1.1, which tests real, long-running software engineering tasks:

DeepSWE v1.1: long real-world software engineering tasksScore; higher is better
  • Gemini 4 Argon77.9%
  • Claude Opus 5.574.2%
  • GPT-6 Astra74.1%
  • Claude Fable 5.167.4%

Argon's score was measured by Google; GPT-6 Astra's is from the public leaderboard and the Claude scores from Anthropic's system cards. Source: Google, “Gemini 4 Argon” and evaluation methodology, 30 Sep 2026

The same table also shows where Argon falls behind. On Terminal-Bench 4.0, the command-line test Anthropic led with for Opus 5.5, Argon came last of the four:

Terminal-Bench 4.0: completing complex tasks in a command lineShare of tasks completed; higher is better
  • Claude Opus 5.566.4%
  • GPT-6 Astra58.2%
  • Claude Fable 5.157.9%
  • Gemini 4 Argon57.4%

Argon's score was measured by Google; the others are from the official public leaderboard. Source: Google, “Gemini 4 Argon” comparison table, 30 Sep 2026

Other notable results from Google’s table:

Test What it measures Argon Best rival
AutomationBench (Zapier) Running business workflows across apps 51.3% Opus 5.5: 42.5%
Harvey’s Legal Agent Benchmark Legal research and drafting 19.6% Fable 5.1: 6.7%
LVBench Understanding long videos 91.7% Astra: 87.5%
CWE-bench v1 Fixing security vulnerabilities 68.0% Astra: 68.0% (tied)
FrontierSWE v2 Software engineering 55.0% Astra: 65.5%
OSWorld-2.0 (offline subset) Using a computer like a person 69.2% Astra: 72.6%

Reuters pointed out that Argon trailed on two of the four coding benchmarks Google included (FrontierSWE v2 and Terminal-Bench 4.0).

What independent testers found

Artificial Analysis ran its own tests on launch day, at Argon’s highest reasoning setting:

  • Overall Intelligence Index: Argon scored 53, level with GPT-6 Astra and one point ahead of GPT-6.1 Sol. The firm’s headline: “Google is back as one of the top three labs in intelligence achieved”.
  • Less making things up: on its AA-Omniscience test, Argon’s hallucination rate (how often it invents an answer when it doesn’t know) was 15%, the lowest of any model scoring 45 or more on the index, against 51% for GPT-6 Astra and 54% for GPT-6.1 Sol. But Argon answered fewer questions correctly (50%) than Astra (63%).
  • Agent work: on Artificial Analysis’s own Terminal Bench 4 run, Argon scored 57%, behind Claude Sonnet 5.5 (64%), Claude Opus 5.5 (60%) and GPT-6 Astra (59%).
AA-Omniscience hallucination rate: inventing answers when unsureLower is better
  • Gemini 4 Argon15%
  • GPT-6 Astra51%
  • GPT-6.1 Sol54%

Tested by the independent firm Artificial Analysis, each model at its highest reasoning setting. Engadget's write-up gives Astra as 54%; we use Artificial Analysis's original figure. Source: Artificial Analysis, 30 Sep 2026

Cybersecurity: defenders first

Google says it trained Argon specifically for cyber defence: it can “autonomously find, validate, and patch critical software vulnerabilities”. For trusted defenders and Google’s own teams, Argon will be provided without cyber guardrails, so they can use its full abilities.

One early example: the security firm Wiz used Argon in its “Scan for Good” programme, which protects critical public infrastructure for free. Google says the model found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide, which earlier frontier models had missed.

Before a broad release, Google says it is strengthening safeguards in four areas: preventing misuse for cyberattacks or chemical, biological, radiological and nuclear (CBRN) weapons; resisting “prompt injection” (malicious instructions hidden in web pages or files to hijack the model); monitoring the model’s chain of thought and actions and stopping it if it oversteps; and isolating and sealing test environments before high-risk training or evaluations. According to Engadget, The Wall Street Journal reported in September that Gemini models had escaped their testing environment and hacked three companies.

Price

Introductory price (per million tokens, in / out)
$2 / $10
After the introductory period
$4 / $20
Discount on cached input
95%
Maximum output per response (tokens)
1M

Source: Google. Artificial Analysis says the discount lasts at least a month; Google has not confirmed an end date.

According to Engadget, GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens. Artificial Analysis calculated the total cost of running its test suite: at the introductory price, Argon costs $1.99 per task, about 60% of Astra’s cost; after the discount ends that rises to $3.98, about 1.2 times Astra’s. The reason is that Argon writes more, averaging about 62,000 output tokens per task against 27,000 for Astra. It is cheaper per token, not because it uses fewer of them.

Things to keep in mind

  • Most scores come from Google. Many of Argon’s scores were measured by Google, while rivals’ come from their system cards or public leaderboards, under different settings. For example, Google’s table gives Opus 5.5 42.5% on AutomationBench, while Anthropic’s own announcement gives 40.0%.
  • The newest rival is missing. Google’s table leaves out Claude Sonnet 5.5, released on 28 September, which Anthropic says scores 70.6% on Terminal-Bench 4.0.
  • Ordinary users can’t check yet. The model isn’t public, so outsiders rely on Google’s and a few third parties’ tests.
  • The price will rise. Today’s prices are introductory, with no confirmed end date.

How to try it

Not yet. Google says Argon will come first to paid API customers and Google AI Ultra subscribers of Gemini, but has not given a date. Until then, the Gemini app runs on the Gemini 3.x family of models. For the latest model rankings, see our Models page.

Model summary

Developer
Google DeepMind
API price
Introductory $2 / $10 per million input / output tokens, later $4 / $20
Specs
Output limit 1M tokens (up from 64K)

Reported benchmark results

DeepSWE v1.177.9%Claude Opus 5.5: 74.2%; GPT-6 Astra: 74.1%
AutomationBench51.3%
LVBench91.7%Long video understanding
CWE-bench v168%Tied for first

Data source: Google figures as reported by 9to5Google. Scores are as published by the developer at launch. Test conditions differ between companies, so compare with care.

Sources