Gemini 4 Argon: Google's most capable model ships to a vetted few, not the public
Google launched Gemini 4 Argon with a 1M-token output limit and top-tier scores. We separate the vendor benchmarks from independent tests and the skeptic case.

Google announced Gemini 4 Argon on September 30 in a launch post signed by Koray Kavukcuoglu, its senior vice president for Google DeepMind. The model is pitched at real-world software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. It is not, however, going to the public. Argon rolls out first to vetted cyber defenders through Google's Fairwind Program, then to paid API customers and Google AI Ultra subscribers "as soon as possible". Google named no date and said it is still iterating on guardrails.
The benchmarks and the skeptics both arrive before any general user gets a turn.
The number Google leads with
The headline spec is an output limit. Argon writes up to 1 million tokens in a single response, up from 64,000, inside a 1-million-token context window. Inputs may be text, image, video or speech; output is text. A new Gemini API feature, Long Decode Continuation, pauses and resumes a long response so a request does not time out mid-reasoning. Whether any real workload needs a million output tokens is a separate question; the limit is a capability, not a promise of coherence across it.
What Google's own scorecard claims
Every figure Google published is vendor-reported and unaudited, and TechCrunch covered the launch with the expected superlatives.
| Google-claimed result | Figure |
|---|---|
| DeepSWE v1.1, agentic software engineering | 77.9%, state of the art |
| Vals Index | first place |
| Vals Finance Agent v2 and Harvey's legal benchmark | leading |
| Zapier AutomationBench | 51.3%, first place |
| LVBench, long-video understanding | 91.7%, state of the art |
| CWE-bench v1 | 68%, tied for first |
Google also reports internal engineering wins that are harder to check: a quantum-computing subroutine whose spacetime resources, qubits times gates, were cut 40% below the published baseline "in a matter of minutes"; more than 300 TiB of memory freed across Google data centers, with an estimated 500 TiB to 1 PiB total saving; and agents porting C and C++ to Rust, including re2, libgav1, and up to 800,000 lines of the Fuchsia Zircon kernel. Wiz used Argon in its Scan for Good program, where it found a critical vulnerability in hospital-used healthcare software that earlier frontier models missed.
What independent testing found
None of the above has been independently replicated except by Artificial Analysis, which measured Argon separately on September 30. It scores 53 on the AA Intelligence Index at high reasoning effort, tying GPT-6 Astra at maximum and edging GPT-6.1 Sol at maximum by a point, a tie for the top rather than a lead. That is 23 points above Google's previous non-Flash model, Gemini 3.1 Pro Preview, and AA calls Argon Google DeepMind's first proprietary model above the Flash class in more than seven months.
| Artificial Analysis measure | Gemini 4 Argon | GPT-6 Astra (max) |
|---|---|---|
| Intelligence Index | 53 | 53 |
| Cost per Index task, intro pricing | $1.99 | $3.26 |
| Output tokens per task | 62k | 27k |
| Terminal-Bench 4 | 57% | 59% |
| AutomationBench-AA | 77.5% | not reported |
The cost edge is real but thinner than it looks. Argon's introductory pricing is 60% of Astra's per-task cost, and that comes from lower token prices, not fewer tokens: Argon averages 62,000 output tokens per task against 27,000 for Astra. At standard pricing the gap narrows to roughly 1.2 times Astra. On Terminal-Bench 4 it trails Claude Sonnet 5.5 at 64%, Claude Opus 5.5 at 60% and GPT-6 Astra at 59%. The near-tie with the GPT-6 models rests on a comparison we flagged when GPT-6.1 Sol shipped: index equality can hide how many tokens a model burns to get there.
Benchmarks against the actual job
Some Google employees with direct access told Bloomberg that Argon performs well on public benchmarks but less well on real tasks, front-end design in particular, according to a summary of the report. Some staff believe Anthropic's Fable and OpenAI's Astra are improving faster than Gemini; others say Argon has closed the gap. Google called it "inaccurate to say that Gemini 4 is underperforming in areas such as coding". An employee familiar with model development told the news service there is "large consensus" internally that Gemini 4 sits at the frontier.
Kavukcuoglu answered on the record with conviction rather than data: "I have the utmost trust in the team. In my mind, it's a certainty that we are always gonna be at the frontier." Alphabet shares pared a gain of more than 2% to 0.5% after the story ran. Training runs for a model of this class can cost as much as $400 million, and Google previously scrapped a planned Gemini 3.5 Pro release announced at its May I/O conference.
Full capability, thin guardrails
Google's safety framing deserves to be read closely. The company says Argon refuses harmful cyber and CBRN requests while preserving dual-use research, monitors internal activations to spot misuse, and runs "monitoring for misalignment" that watches Argon's chain of thought and actions and can stop execution. The blog also urges the rest of the industry to preserve reasoning transparency. The posture is not unique to Google; Anthropic's own cyber model raised the same dual-use questions a day earlier.
Read those commitments next to the deployment and the tension is plain. The same post that promises to catch misuse ships the model without cyber guardrails to trusted defenders and to Google's own teams. A model with full cyber capability, handed to a vetted few before the guardrails are finished, is a bet the public cannot audit.
Price, and the end date nobody knows
Pricing is introductory until further notice: $2 per million input tokens and $10 per million output, with cached input at 95% off; standard rates will be $4 and $20. Google has not said when the promotion ends, so the cost of building on Argon cannot be modelled beyond the near term, and no public release date has been shared. A companion piece on the model's measured hallucination rate lands the same day.
What is not here is as telling as what is: no technical paper, no weights, no public access date, and no third-party replication beyond Artificial Analysis. That silence is standard for a frontier lab, and it is why the vendor's benchmark tables, not the phrase "most powerful model yet", are the wrong place to stop reading.


