Google announced Gemini 4 Argon, its new flagship AI model, on September 30. For now it goes only to a group of trusted cyber defenders and to Google’s own teams, and Google gave no date for developers or consumers. Google says Argon leads its rivals on most of the tests it published. Artificial Analysis, which tests models independently, scores it level with OpenAI’s GPT-6 Astra.
New to this? Read it in simple words
- Google has a new top AI model called Gemini 4 Argon.
- For now, only a small group of cyber security experts can use it.
- Google says it is best on 12 of 18 tests, but these are its own numbers.
- An independent tester rates it about as smart as OpenAI’s best model.
- Benchmark
- A standard test used to compare AI models.
- Token
- A small piece of text that AI models read and write.
- Cyber defender
- A security expert who protects computer systems from attacks.
Who gets it first
Argon is rolling out to “a set of trusted cyber defenders” through Google’s Fairwind Program, and to Google’s own teams. These early users get it without its cyber guardrails, Google says.
Paid API customers and Google AI Ultra subscribers will be next, followed by developers, businesses and consumers. Google said this will happen “as soon as possible”, but gave no date. Reuters notes that there is no timeline for a public release.
Google says a phased approach is needed to release capabilities at this level safely. It also takes part in the US government’s voluntary process for testing models before release.
Artificial Analysis Intelligence Index, checked September 30. Google’s own tables rank Argon higher.
Google’s scorecard
Google compared Argon with GPT-6 Astra, Claude Opus 5.5 and Claude Fable 5.1 on 18 tests. By its count, Argon leads on 12 and ties for first on one.
The biggest leads are on office-style work and long documents. On the Vals Index of professional tasks, Google puts Argon at 68.9%, against 67.0% for Claude Opus 5.5 and 63.1% for GPT-6 Astra.
Coding is more mixed. Argon sets a new best of 77.9% on DeepSWE. But it comes last of the four on FrontierSWE and on Terminal-bench 4.0, where Claude Opus 5.5 scores 66.4% to Argon’s 57.4%.
Argon can write up to 1 million tokens in one answer, up from 64,000. Its introductory price is $2 per million input tokens and $10 per million output tokens, rising later to $4 and $20.
The outside check
Artificial Analysis, which runs its own tests, gives Argon 53 points on its intelligence index. That puts it level with GPT-6 Astra and Claude Fable 5.1, and five points behind Claude Opus 5.5.
Running its tests cost $1.99 per task at Argon’s introductory prices. That is less than GPT-6 Astra at $3.26 and Claude Opus 5.5 at $5.98.
Google’s cyber claims are harder to check. It compares Argon’s security results only with its own models, The New Stack notes, and one example comes from Wiz, a security company that Google owns.
Sources
Every fact in this story comes from the sources below. Open them to check our work.
- 1
- 2Research · September 30, 2026Google announces Gemini 4 flagship AI model after months of delays Reuters, via Investing.com
- 3Research · September 30, 2026Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic — but in limited release VentureBeat
- 4Research · September 30, 2026Gemini 4 Argon is here: It’s great, and you can’t have it yet The New Stack
- 5Research · Accessed September 30, 2026Gemini 4 Argon (high) - Intelligence, Performance & Price Analysis Artificial Analysis
We read Google’s announcement and compared Reuters, VentureBeat and The New Stack. The scores in the second section are Google’s own, as its tables give them. We checked Artificial Analysis’s independent results on September 30; they cover Argon’s “high” reasoning setting.