Gemini 4 Argon Tops Benchmarks, But Google Staff Aren’t Sold

A cybersecurity analyst inspecting code on monitors in a dark office, lit by blue screen glow

Google has finally shipped the model it spent months being asked about. On September 30, it unveiled Gemini 4 Argon, the new flagship anchoring its Gemini 4 generation — larger than anything in its old “Pro” line and, by Google’s own scorecard, ahead of the competition on key coding and cyber benchmarks. TechCrunch notes it’s the company’s most powerful model yet, built to autonomously find, validate, and patch critical software vulnerabilities.

There’s just one catch. You can’t have it.

The Model You Can’t Touch Yet

Argon isn’t launching to developers, enterprises, or consumers. Instead, Google is rolling it out first to a small circle of trusted cybersecurity partners through its Fairwind Program, while participating in the U.S. government’s voluntary pre-release model access process. No public timeline was given. According to CNBC TV18, the launch came one day after CEO Sundar Pichai joined other tech chiefs at the White House to sign a voluntary AI safety accord.

The rollout is gated by design. “Starting this rollout in this way gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible,” Gemini product lead Tulsee Doshi told CNBC. Early proof is real: Wiz is already using Argon, and Google says it found a critical flaw in hospital software that other advanced models had missed.

What the Numbers Say

On Google’s disclosed benchmark table, Argon leads outright on 12 of 18 tests and ties one — more top scores than GPT-6 Astra or Claude Opus 5.5. The biggest wins sit where enterprises spend money: 19.6% on Harvey’s Legal Agent Benchmark, 51.3% on Zapier’s AutomationBench, 91.7% on long-video understanding, and a tie with Astra at 68% on vulnerability remediation.

But it isn’t a clean sweep. As CNBC TV18’s report noted, Argon lagged on two of the four coding benchmarks Google itself included — Astra still leads FrontierSWE v2 by 10.5 points, Opus 5.5 takes Terminal-bench 4.0 by 9.

Then there’s the price. Argon opens at $2 per million input tokens and $10 per million output — one-fifth of GPT-6 Astra’s $10/$50 and half of Claude Opus 5.5’s $4/$20, rising to $4/$20 after the introductory window, with the output ceiling jumping from 64,000 to 1 million tokens.

What Google’s Own People Say

Here’s where the launch story gets awkward. The Japan Times reports that Google engineers with direct access to the model are not nearly as impressed as the leaderboard. The model “does less well when employees actually put it to work,” they said, particularly on coding — with one insider adding that it “isn’t particularly adept at front-end design.” Some staff believe Anthropic’s Fable and OpenAI’s Astra are improving at a faster rate than Gemini.

Google pushed back hard, calling it “inaccurate” to say Gemini 4 is underperforming in coding. DeepMind chief Koray Kavukcuoglu declared “it’s a certainty that we are always gonna be at the frontier.” Investors weren’t reassured: Alphabet shares faded from up more than 2% to up 0.5% after the story landed.

Why this matters

This is the industry’s “benchmaxxing” problem playing out in public for the first time at a major lab: when benchmarks are the marketing, there’s enormous pressure to optimize for the test — and Google’s own engineers are now the ones telling reporters that test scores and real work don’t line up.

That tension is the real story of Argon, and it’s bigger than one launch. Google canceled the planned Gemini 3.5 Pro, overhauled DeepMind’s leadership with Demis Hassabis stepping aside, and spent months behind OpenAI and Anthropic at the frontier. Argon’s job isn’t just to win benchmarks — it’s to convince the market Google can still ship. The pricing does half that job: at one-fifth of Astra’s cost during the intro window, Google is buying its way back into the price conversation. But pricing is a promise about the future; the engineers’ skepticism is a report on the present. Until developers can actually touch Argon, the only people who can settle the argument are the ones Google employs.

Meanwhile in AI Tech: OpenAI is playing the same game from the other side — the DevDay pricing reveal for GPT-6.1 Sol.

Meanwhile in Tech News: the White House safety accord behind Argon’s gated launch — Trump’s meeting with AI CEOs on guardrails versus speed.

Written by
Ryan covers artificial intelligence and enterprise tech — from foundation models and AI chips to the business of machine intelligence. He tracks model releases, funding rounds, and the policy moves shaping the AI industry.