Gemini 4 Argon Explained: Benchmarks, Access, and What We Know
Google has announced a frontier model that almost nobody outside a vetted list of security organizations can use yet, and it arrived by way of a cancellation. Gemini 4 Argon was unveiled on September 30, 2026, three months after the company quietly dropped Gemini 3.5 Pro, a model CEO Sundar Pichai had earlier said was expected in June. Google’s own charts claim a lead over Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Astra on most of the benchmarks it chose to show, and the same charts concede defeat on several others.
That mix of strong claims and restricted access is why this launch is worth reading slowly. The numbers are vendor-reported, there is no technical report, and the model’s size, context window and training recipe are not disclosed. What is disclosed is unusually specific about price, about who gets access first, and about where Argon loses.
By the end of this article you will know which facts are confirmed, which numbers are Google’s claims awaiting replication, and how to decide whether Argon belongs on your evaluation list.
What this covers: the cancellation of Gemini 3.5 Pro and where Argon fits in the lineage, what Google has and has not disclosed about architecture, the sourced benchmark table with its wins and losses, the Fairwind access program, pricing arithmetic, safety design, failure modes, and a decision matrix against Opus 5.5 and Astra.
Context and Background
Google’s Gemini line has moved quickly through 2026. Our earlier coverage of Gemini 3.5 Pro described the generation that was supposed to be the next Pro-tier flagship. The cheaper end of the family, covered in our look at Gemini 3.8 Flash, kept shipping while the Pro tier stalled. According to reporting that cites Google’s announcement and the Ars Technica account, Gemini 3.5 Pro never materialized despite the earlier June expectation, and Google has shifted its attention to Gemini 3.8 Flash and now to Argon.
The competitive backdrop explains the urgency. Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Astra, the model family we examined in our piece on OpenAI Astra and opaque recurrence, set the frontier that Argon is measured against. Anthropic’s recent mid-tier release, analyzed in Claude Sonnet 5.5 pricing and benchmarks, shows how much of the 2026 contest is about cost per completed task rather than raw leaderboard position.
Argon also introduces a new naming scheme. The Gemini number is now followed by a codename, and the 9to5Google report notes this marks a departure from earlier version numbering. Treat “Gemini 4” as the generation and “Argon” as the flagship model name within it. Google has not said whether other Gemini 4 models will follow, so any claim about a Gemini 4 family beyond Argon is speculation.
The release is also the first big Google model announcement to sit squarely inside a government process. On September 29, 2026, Google and other AI companies signed a voluntary safety agreement with the US government covering pre-release evaluations and safeguards against unintended hacking or unauthorized access, as reported by Tech Wire Asia. Argon participates in the US voluntary pre-release access process, and no public timeline for that review has been given.
One more piece of context matters for readers who budget for inference: price. Argon’s introductory price is $2 per million input tokens and $10 per million output tokens, converting later to $4 and $20, which equals Opus 5.5’s list rates as reported. We will return to what that does to your cost model in a dedicated section below.
Gemini 4 Argon: What Is Confirmed and How the Rollout Works
Gemini 4 Argon is Google DeepMind’s new frontier model, announced September 30, 2026 by Koray Kavukcuoglu. It is available only to vetted cyber defenders in the Fairwind Program for now, with paid API customers and Google AI Ultra subscribers next. It raises maximum output from 64,000 to 1 million tokens. Parameter count and context window are not disclosed.
The most decision-relevant fact about Argon is not a benchmark. It is that you probably cannot call it today. Google’s announcement says Argon is rolling out first to trusted testers and cyber defenders through the Fairwind Program, then to paid API customers and Google AI Ultra subscribers, and finally to “developers, enterprises, and consumers” as soon as possible, according to the secondary coverage of the post. No dates were given for the second or third phase.

Figure 1: The Gemini 4 Argon rollout as announced. Gemini 3.5 Pro was cancelled, Fairwind comes first, and the US voluntary pre-release review runs in parallel to the public phases.
Figure 1 shows the sequencing. Two inputs feed the first phase: the Fairwind Program, which Google opened in early September 2026 (reports differ between September 2 and September 3), and the pre-release process with US government evaluators. Phase two and three are gated on both.
What Fairwind is and who is in it
Fairwind is a Google program for vetted defenders. SiliconANGLE reports more than 650 organizations had signed up at launch, including CrowdStrike and Palo Alto Networks. Reporting describes the members as vetted Google Cloud customers, government agencies, cybersecurity partners and Google’s own internal teams. Before Argon, the program’s models were Gemini 3.8 Flash Cyber and the CodeMender system, per Cyber Kendra’s account.
Access carries conditions. Tech Wire Asia reports that multi-factor authentication is required and that use is restricted to employees in cybersecurity, incident response or penetration-testing roles. Fairwind members also get a version of the model with cyber guardrails removed, according to the same coverage. That is a notable design choice: the model that defenders receive is deliberately more permissive than the model the public will eventually see.
What Google has disclosed about the model itself
Strip the marketing and the disclosed technical specification is short. The maximum output limit is 1 million tokens per response, up from 64,000 for earlier Gemini models. One secondary source says the model was tested up to a 1M-token context, and another claims a 2 million-token context window. Google’s own post, as summarized, describes a “1 million token output limit” and does not state a context window in the text we could review, so we treat the context window as not confirmed.
Several widely repeated details are absent from every source we found. Parameter count: not disclosed. Dense versus Mixture-of-Experts layout: not disclosed. Training data scale and compute: not disclosed. Tokenizer and vocabulary: not disclosed. Modalities: the LVBench result implies video understanding, but a full modality list was not in the sources we could verify. The pre-announcement framing that Argon is “larger than prior Pro models” comes from earlier reporting and is not confirmed by Google’s published material.
Why a 1M output limit is more important than it sounds
A jump from 64K to 1M output tokens is a different category of capability, not a bigger number. Sixty-four thousand tokens is roughly a long essay or a mid-sized source file. One million tokens of output is enough for a whole code module tree, a full translated corpus or a very long agent trajectory in a single generation.
Google’s internal examples lean on exactly that scale. The company says Argon scaled a C and C++ to Rust migration to more than 800,000 lines in the Fuchsia operating system’s Zircon kernel, with the work still undergoing audit. It also reports replacing 32,000 lines of SIMD code in the libgav1 video decoder, yielding a decoder 2.7 times faster than the prior Rust port. These are Google-reported results on Google-owned code, which is both the most realistic test environment and the least independently checkable.
The practical caveat is about generation, not capability. Long outputs are expensive at $10 or $20 per million output tokens, they take real wall-clock time, and error rates compound over a long unverified generation. A model that can emit one million tokens still needs a harness that checks its work in chunks.
The naming and generation question
Because Google skipped the planned 3.5 Pro, the Gemini 3.x line now ends at Gemini 3.8 Flash for the cheaper tier while Gemini 4 starts at the very top. That inverts the usual rollout, where a small Flash model trails the Pro model. Analysts will read this as Google concluding that a mid-generation Pro model no longer justified its own launch once a stronger flagship was close.
That is interpretation, and Google has not offered a stated reason beyond messaging. Reporting characterizes the new emphasis as a cost advantage, which the pricing section below tests against the actual list prices.
Deeper Analysis: Reading the Gemini 4 Benchmarks Honestly
Every number in this section comes from Google’s announcement as relayed by the outlets named, and none has been independently replicated. The honest framing for Gemini 4 benchmarks is “reported by the vendor, under conditions the vendor chose.” That is not an accusation, it is the normal state of any launch week, but it changes how much weight a single score should carry.
The scorecard: where Argon leads and where it does not
VentureBeat counts Argon as leading or tying on 13 of 18 disclosed benchmarks, more than Astra with three outright leads or Opus 5.5 with two. Other coverage says 12 of 18 or 13 of 19 outright. The counts differ because outlets treat ties and rows differently, so use the shape of the result rather than the exact ratio.
| Benchmark | What it measures | Argon | Comparison (reported) |
|---|---|---|---|
| DeepSWE v1.1 | Real-world software engineering tasks | 77.9% | Opus 5.5 74.2%, Astra 74.1% |
| AutomationBench | Business workflow automation (Zapier) | 51.3% | Opus 5.5 42.5% |
| LVBench | Long video understanding | 91.7% | Astra 87.5% |
| Harvey legal agent benchmark | Legal analysis tasks | 19.6% | Astra 5.4% |
| GraphWalks, 256K to 1M tokens | Long-context graph reasoning | 84.2% | Astra 71.8% |
| CWE-bench v1 | Vulnerability remediation | 68% | Tied with Astra at 68% |
| FrontierSWE v2 | Frontier software engineering | 55.0% | Astra 65.5% (Argon trails) |
| Terminal-Bench 4.0 | Terminal agent tasks | 57.4% | Opus 5.5 66.4% (Argon trails) |
| PostTrainBench | Automating post-training research | 45.3% | Opus 5.5 49.3% (Argon trails) |
| OSWorld-2.0 | Computer-use agents | 69.2% | Astra 72.6% (Argon trails) |
The pattern is consistent. Argon’s wins cluster in enterprise knowledge work such as legal, workflow automation and long-context retrieval, in video understanding, and in one headline coding benchmark. Its losses cluster in agentic execution: terminal use, computer use and frontier-level software engineering. The planning brief for this post framed it as lagging on two of four coding benchmarks, and the published rows support that reading, with FrontierSWE v2 and Terminal-Bench 4.0 as the visible gaps.

Figure 2: The reported Gemini 4 Argon scorecard grouped into leads, trails and ties. The losses sit in terminal and frontier software engineering tasks.
Figure 2 collapses the table into three bins. If your workload is a terminal agent that runs shell commands for hours, the Terminal-Bench gap of roughly nine points is the number to care about, not the DeepSWE lead of roughly three.
Why the same benchmark can mean different things
Three caveats apply to nearly every row above. They are the reason a vendor chart should open a conversation rather than close one.
First, harness effects. The analysis from AlphaCorp notes that CWE-bench ran each model inside a different agent framework. A vulnerability-remediation score reflects the model plus its scaffolding, tool permissions, retry policy and time budget, so a tie at 68% between two models is a tie between two systems.
Second, different suites across vendors. Anthropic and OpenAI have favored different benchmark sets in their own announcements, and Google’s chart is built from the rows where a comparison could be made. A row that is missing from the chart is information too: nothing in the sources lists Argon’s scores on benchmarks Google chose not to show.
Third, saturation and contamination. Software engineering benchmarks leak into training data over time, and a three-point lead on a single benchmark is inside the noise that harness choice alone can create. A margin of 77.9% against 74.2% is meaningful as a direction and weak as a ranking.
The security numbers
Argon’s headline positioning is cyber defense, so the security results deserve their own look. Google reports 85.8% Pass@1 on real-world vulnerability discovery and 70.9% on the Wiz penetration-test benchmark, and reporting says Argon outperformed Gemini 3.8 Flash Cyber on the Wiz test. Google also says the model identifies vulnerabilities across more than 20 programming languages on internal benchmarks and that a Wiz Scan for Good engagement found a critical flaw in healthcare software that earlier models missed.
On adversarial robustness, Google cites the Gray Swan Indirect Prompt Injection benchmark and says Argon leads. VentureBeat’s reading puts the attack success rate at 0.7% versus 8.5% for Astra, and AlphaCorp reports 0.7% as the lowest of 13 models tested. A sub-one-percent rate is excellent, but an attack success rate depends heavily on the attack set, and red-teaming results from Google are, per AlphaCorp, unpublished.
Reported internal results
Google highlights four internal wins. A quantum subroutine optimization improved a baseline by 40%. Fleetwide profiling freed more than 300 tebibytes of memory across data centers, with a projected 500 TiB to 1 PiB in total. The Rust migration reached more than 800,000 lines, and the libgav1 rewrite was 2.7 times faster. None of these has an external benchmark, and the Zircon migration is still under audit.
These examples matter because they show what Google believes the model is for: long-horizon engineering work on large, well-tested codebases. The 300 TiB number is the kind of result that does not appear on a leaderboard but changes a fleet cost model. It is also the kind of result a reader cannot reproduce.
Pricing: Does the Cost Advantage Hold?
Argon’s introductory list price is $2 per million input tokens and $10 per million output tokens. Cached input tokens carry a 95% discount, which works out to $0.10 per million at the introductory rate. After the introductory period, standard pricing is $4 input and $20 output, with no date announced for the change.

Figure 3: Reported list prices per million tokens. Argon’s introductory rate is half of Opus 5.5, its standard rate matches Opus 5.5, and Astra is reported at $10 input and $50 output.
Figure 3 shows the triangle. VentureBeat reports Argon’s introductory rate as one-fifth the cost of Astra at $10 and $50 and half of Opus 5.5 at $4 and $20. A different outlet, Forkast, wrote that OpenAI plans introductory $2 and $10 rates that also double later. The two descriptions of OpenAI pricing conflict, so check OpenAI’s pricing page before building a model on either one.
A worked cost example
Consider an agent task that consumes 400,000 input tokens and generates 60,000 output tokens, with half the input served from cache. These are illustrative figures chosen for arithmetic, not measurements.
At Argon’s introductory rate, uncached input is 200,000 tokens at $2 per million, which is $0.40. Cached input is 200,000 tokens at $0.10 per million, which is $0.02. Output is 60,000 tokens at $10 per million, which is $0.60. The task costs about $1.02.
At the standard rate of $4 and $20, the same task is $0.80 for uncached input, roughly $0.04 for cached input if the 95% discount applies to the higher base, and $1.20 for output, about $2.04. The introductory period halves the bill, and standard pricing is simply Opus 5.5’s list rate.
The lesson is that output tokens dominate agentic cost. At a 1M output ceiling, one runaway generation can cost $10 at the introductory rate or $20 at the standard rate. Set explicit maximum output budgets even though the ceiling is generous.
What “cost advantage” really means
Reporting says Google’s messaging shifted toward cost advantage. List price is only half of cost. The other half is tokens per task: if Argon needs more reasoning or more retries to finish a terminal task than Opus 5.5, its cheaper rate can disappear. Our Sonnet 5.5 piece makes the same point about frozen list prices, where the efficiency lever is fewer tokens, not cheaper tokens.
The right metric is dollars per accepted outcome, measured on your own tasks. Run a fixed set of 50 to 100 real jobs through each model, record total tokens, tool calls and pass rate, and divide. A model that is 2.5 times cheaper per token but needs 3 times the tokens is not cheaper.
The introductory-to-standard transition is the other trap. Any budget approved at $2 and $10 should carry a line item for the $4 and $20 case, because a pilot that looks cheap can double when the introductory window closes.
Safety Design and the Dual-Use Problem
Google describes a layered safety design for Argon, and it is worth understanding because it explains the access restrictions. The announcement lists four mechanisms: monitoring of internal model activations to detect misuse, hardened sealed sandboxes for high-risk testing, chain-of-thought monitoring with the ability to stop execution, and robustness against indirect prompt injection. The company says it follows its Frontier Safety Framework, which AlphaCorp lists as version 3.1 dated April 17, 2026.
Monitors that can halt the model
The most architecturally interesting claim is the separate monitoring layer. Reporting says separate monitors track the model’s chain of thought and its actions, and can halt execution if the model exceeds user intent. This is a classic supervisory pattern: a second system with independent authority watching the first, rather than a single model policing itself.
The pattern has a known weakness. Chain-of-thought monitoring only works while the visible reasoning reflects the actual computation. Google has not published how it validates that, and red-teaming results are unpublished. Treat the monitor as a useful control, not a proof.
Why cyber capability drives the rollout
Autonomous vulnerability discovery, validation and patching are dual-use. The same capability that finds a flaw in healthcare software for a defender can find it for an attacker. That is why access is tiered: vetted defenders with multi-factor authentication and role restrictions first, then paying customers, then everyone.
The odd detail is that Fairwind members reportedly receive a version without cyber guardrails. Defenders need a model that will actually write exploit proofs of concept, and a guardrailed model that refuses is useless for penetration testing. The public version will presumably keep the guardrails, which means public benchmark behavior on security tasks may differ from the Fairwind results Google reports.

Figure 4: A decision flow for evaluating Gemini 4 Argon. Access status decides whether you test now or wait, and workload type decides which comparison matters.
Figure 4 turns the benchmark pattern into a procedure. The key branch is workload type: terminal agents point you to the Opus 5.5 comparison, legal, finance and long-video work point to an Argon pilot, and security agents need a check against the Astra tie.
How Gemini 4 Argon Compares
The following matrix uses only reported figures and is meant as a starting hypothesis for your own evaluation, not a verdict. “Not shown” means the sources we reviewed did not report a comparable number.
| Dimension | Gemini 4 Argon | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|
| Availability | Fairwind only, wider rollout undated | Generally available per Anthropic | Per OpenAI; see our Astra coverage |
| List price per M tokens | $2/$10 intro, $4/$20 standard | $4/$20 | $10/$50 as reported by VentureBeat |
| DeepSWE v1.1 | 77.9% | 74.2% | 74.1% |
| Terminal-Bench 4.0 | 57.4% | 66.4% | Not shown |
| FrontierSWE v2 | 55.0% | Not shown | 65.5% |
| Long-context graph reasoning | 84.2% at 256K to 1M | Not shown | 71.8% |
| Max output | 1M tokens | Not disclosed here | Not disclosed here |
Gemini 4 vs Astra and Opus by use case
For legal and finance agents, Argon’s reported Harvey legal result of 19.6% against Astra’s 5.4% and the Vals Index claim make it the first model to pilot, with the caveat that an absolute 19.6% still means most tasks fail. For terminal-heavy automation such as CI debugging or infrastructure agents, Opus 5.5’s 66.4% lead on Terminal-Bench 4.0 is the stronger signal. For long-video analysis, Argon’s 91.7% on LVBench against Astra’s 87.5% is a clear reported edge. For security work, Argon and Astra tie at 68% on CWE-bench v1, so the deciding factors will be access, price and harness integration.
If you want broader context on Claude’s flagship line, see our write-up of Claude Opus 5.5 for the model Argon is most often compared with on coding and agents.
Trade-offs, Gotchas, and What Goes Wrong
The first gotcha is the one the launch cannot hide: Argon is not generally available. Anything you read about performance in production is secondhand, and every number lives inside Google’s presentation. Do not rewrite an architecture plan around a model you cannot call.
The second is the missing technical report. As of September 30, 2026 AlphaCorp notes that no report has been published. Without a model card, you cannot check context behavior, long-output degradation, refusal rates or the evaluation protocol. Wait for the report before relying on any specific figure for compliance or procurement.
Specific failure modes to plan for
Long-output drift. A 1M-token generation compounds small errors. Break large jobs into verifiable chunks, run tests between chunks, and cap output per call. The Zircon migration is still under audit, which is a reminder that scale of output and correctness of output are different things.
Agentic weakness. The reported trailing scores on terminal use, computer use and FrontierSWE v2 suggest Argon may be less reliable on open-ended execution than on bounded analysis. If your workflow is a long autonomous shell session, test that specifically.
Guardrail mismatch. If you prototype on a Fairwind build and deploy on the public build, behavior will differ on security tasks. Validate on the same variant you will ship.
Price cliff. The $2/$10 introductory rate is temporary and its end date is not announced.
Benchmark overfitting. A lead on 12 or 13 of 18 benchmarks selected by the vendor is an invitation to test, not a guarantee. Weigh the rows that match your work and ignore the rest.
Anti-patterns
Do not route all traffic to a new model on launch week. Do not treat the 0.7% prompt-injection rate as a license to remove your own input filtering. And do not assume the safety monitors substitute for least-privilege tool permissions, which remain your job.
Practical Recommendations
Treat Gemini 4 Argon as a promising candidate that has not yet been proven in public. The best plan is to prepare now so that the day access opens you are running evaluations rather than starting them.
If you are a Fairwind member, run Argon on your real vulnerability backlog and measure remediation success on your harness, not Google’s. If you are not, build your evaluation set in advance: 50 to 100 representative tasks with automated pass or fail checks, covering your workflow type. Record tokens, tool calls, latency and cost per accepted outcome for your incumbent model today so you have a baseline.
Budget at the standard $4/$20 rate, even if you pilot at $2/$10. Put an output cap on every call. Keep a model-routing layer so that switching is a configuration change rather than a rewrite.
- Confirm your access path: Fairwind eligibility, paid API waitlist or AI Ultra.
- Build a private evaluation set that matches your workload, not Google’s benchmarks.
- Baseline your current model on cost per accepted outcome.
- Budget at $4/$20 and cap output tokens per call.
- Test the same model variant you will deploy.
- Read the technical report when it appears, and check Google’s pricing page for the end of the introductory period.
- Keep least-privilege tool permissions and your own prompt-injection defenses in place.
Frequently Asked Questions
What is Gemini 4 Argon?
Gemini 4 Argon is Google DeepMind’s new frontier model, announced on September 30, 2026. Google positions it for software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. It raises maximum output from 64,000 to 1 million tokens. Google has not disclosed its parameter count, architecture details or training data, and no technical report had been published as of the announcement.
Can I use Gemini 4 Argon today?
Probably not. Access is limited to the Fairwind Program, which includes vetted cybersecurity organizations, Google Cloud customers, government agencies and Google’s internal teams. Paid API customers and Google AI Ultra subscribers are next, and developers, enterprises and consumers follow, but Google gave no dates. A US voluntary pre-release access process also applies, with no public timeline.
How much does Gemini 4 Argon cost?
Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input discounted 95%, or about $0.10 per million. Standard pricing afterward is $4 input and $20 output. Google has not announced when the introductory period ends. These match reporting from 9to5Google and Tech Wire Asia, and you should confirm on Google’s pricing page.
Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?
On Google’s reported charts, it leads on most disclosed benchmarks, including DeepSWE v1.1 at 77.9% versus 74.2% for Opus 5.5 and 74.1% for Astra. It trails Astra on FrontierSWE v2 and OSWorld-2.0, and trails Opus 5.5 on Terminal-Bench 4.0 and PostTrainBench. All figures are vendor-reported and unreplicated, so run your own tests.
Why did Google cancel Gemini 3.5 Pro?
Google has not given a detailed public reason. Reporting says Gemini 3.5 Pro, which Sundar Pichai had said was expected in June, never materialized, and that Google shifted focus to Gemini 3.8 Flash and then Argon. Any explanation beyond that, such as a view that Argon made the mid-generation Pro model redundant, is interpretation rather than confirmed fact.
What is the Fairwind Program?
Fairwind is Google’s program for giving vetted defenders early access to security-capable models. It opened in early September 2026 and had more than 650 participating organizations at the Argon launch, including CrowdStrike and Palo Alto Networks according to SiliconANGLE. Members use multi-factor authentication and role restrictions, and reportedly receive a version of Argon with cyber guardrails removed.
Further Reading
- Google Gemini 3.5 Pro explained: architecture and benchmarks, the cancelled generation Argon replaces.
- Gemini 3.8 Flash explained: architecture, pricing and benchmarks, the cheaper tier that kept shipping.
- OpenAI Astra and opaque recurrence explained, the model Argon is most often compared against.
- Claude Sonnet 5.5 pricing and benchmarks, for the cost-per-task framing.
- Google’s primary announcement: Gemini 4 Argon: our next era of frontier intelligence.
- Independent reporting: VentureBeat on the benchmark lead and limited release and SiliconANGLE on the Fairwind rollout.
By Riju — about
