Google just did the thing everyone’s been waiting on for the better part of a year. Yesterday — September 30 — Google DeepMind announced Gemini 4 Argon, its new flagship model, the first of the Gemini 4 generation. And then, in the same breath, it told most of us: you can’t have it. Not yet, anyway.
Argon is rolling out first to a hand-picked group of cybersecurity defenders through something called the Fairwind Program. Developers, enterprises, the rest of us — we wait. Google says paid API customers and Google AI Ultra subscribers are next in line, with no date attached. Which makes this a weird launch to cover: the most significant Google model release since the Gemini 3 series, and there’s no signup page.
But here’s the thing — the details Google did share tell you a lot about where the frontier is heading, what your competitors will be building on in six months, and how to position your startup now so you’re not scrambling later. I went through the official announcement, the benchmark charts, and the press coverage so you don’t have to. Pricing, benchmarks, the cyber-first rollout, and five moves worth making this week. Let’s get into it.



What Gemini 4 Argon actually is
Strip away the rollout drama and Argon is a straightforwardly ambitious model. In the announcement post — from Google’s chief AI architect Koray Kavukcuoglu — Google describes it as built to “sustain deep reasoning across complex, long-horizon workflows.” Translation: this isn’t a chatbot upgrade. It’s a model designed to chew through big, messy, multi-step jobs: real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
The positioning matters. Google is explicitly not selling “ask it anything.” It’s selling sustained work — agents that can refactor a codebase, grind through a legal document set, or hunt vulnerabilities across twenty programming languages without losing the plot halfway through. That “long-horizon” phrasing shows up everywhere in the announcement, and it’s the thread connecting every benchmark Google chose to brag about.
A bit of context, because it matters here: VentureBeat notes this is Google’s first flagship model since the Gemini 3 series back in November 2025. That’s nearly a year of watching OpenAI and Anthropic trade blows — GPT-6 Astra, Claude Opus 5.5, the whole price war — while Google shipped smaller models and reshuffled its leadership. (Demis Hassabis moved up to chairman and chief scientist at Alphabet, Jeff Dean left to start his own company, and Kavukcuoglu took the DeepMind helm.) Argon is Google saying: we’re back in the flagship race. Whether the benchmarks back that up — we’ll get there.
The 1M output token thing is a bigger deal than it sounds
Okay, the headline spec. Argon can generate up to 1 million tokens in a single response. Google calls it industry-leading, and — for once — the marketing isn’t stretching. Previous Gemini models topped out at 64K. Every serious rival — Claude Opus 5.5, Claude Fable 5.1, GPT-6 Astra — caps a single response at 128K. Argon doesn’t edge past them. It laps them, roughly eight to sixteen times over.
Why should a founder care about output tokens? Because output length is the ceiling on how much work a model can do before it has to stop, hand control back, and hope the next turn remembers the plan. A 128K ceiling means big jobs get chopped into pieces — and every handoff is a chance for the plan to drift. A 1M ceiling means an agent can rewrite a whole module, draft a full technical report, or migrate a codebase in one continuous trajectory. Google’s phrase for it: the headroom to “think deeply and generate hundreds of thousands of tokens in a single trajectory… to solve tough problems in one go.”
Real talk: the cost is real too. At the introductory $10 per million output tokens, a maxed-out 1M-token response costs ten bucks. After the intro period, twenty. Most tasks won’t come close to that — but if you’re building long-horizon agents, your unit economics just gained a new variable: trajectory length. Worth modeling before you build, not after.
One honest caveat: Google hasn’t disclosed the input context window. Output is only half the story — a model that can write a million tokens but can’t read your whole repo is still bottlenecked. Rivals sit around 1M input tokens, with Astra stretching to about 1.05M. Until Google publishes the number, treat the 1M-output headline as exactly what it is: half a spec sheet.
Pricing: half the price of Claude Opus 5.5 (for now)
Here’s the part you scrolled for. Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, with cached input at 95% off — that’s $0.10 per million. After the introductory period, pricing moves to $4 in / $20 out. Stack that against the field:
- Gemini 4 Argon: $2 / $10 intro (then $4 / $20) — cached input $0.10
- Claude Opus 5.5: $4 / $20 — cached input $0.20
- Claude Fable 5.1: $10 / $50 — cached input $0.25
- GPT-6 Astra: $10 / $50 — cached input $1.00
That’s per million input/output tokens, with competitor and post-intro pricing figures from independent launch coverage. So at intro pricing, Argon undercuts Opus 5.5 by half and Astra by 80%. Even after the intro period ends, it merely matches Opus 5.5 — never exceeds it. And cached input at ten cents is the cheapest of the bunch, which matters more than it looks: agent workloads re-read context constantly, so cache pricing is where the real bill lives. If you want the full story on how we got here, the AI price war we broke down last week has the timeline.
One more data point worth pocketing: Artificial Analysis has Argon matching GPT-6 Astra on its Intelligence Index at roughly 60% of the cost per task at the discounted pricing. Take vendor-adjacent numbers with salt — but the direction is consistent: Google is pricing to win developers back, not to maximize margin.
Benchmarks: where it leads, and where it doesn’t
Standard disclaimer, and I mean it: these are Google’s published comparisons. Every lab picks the evals that flatter it. That said, the pattern across Argon’s chart is unusually coherent — it wins exactly the long-horizon, sustained-reasoning tasks it was built for, and loses the ones it wasn’t. That kind of honesty-by-omission is actually informative.
Where Argon leads:
- DeepSWE v1.1 (long-horizon software engineering): 77.9% — a new state of the art, ahead of Opus 5.5 (74.2%) and GPT-6 Astra (74.1%)
- Vals Index (economic impact across finance, coding, legal, and tax, weighted by US GDP): 68.9%, ranked #1
- AutomationBench (Zapier’s end-to-end business execution eval): 51.3%, ranked #1
- LVBench (long video understanding): 91.7%, state of the art
- CWE-bench v1 (vulnerability remediation): 68%, tied for first
- Harvey Legal Agent Benchmark: 19.6% against 5.4% for GPT-6 Astra — the legal knowledge-work gap is enormous
Where it trails:
- FrontierSWE v2: 55.0% vs GPT-6 Astra’s 65.5%
- Terminal-Bench 4.0: 57.4% vs Claude Opus 5.5’s 66.4%
- OSWorld-2.0 (computer use): 69.2% vs GPT-6 Astra’s 72.6%
Read that pattern like a founder, not a fan. Argon is the long-game model: sustained reasoning, big refactors, document-heavy knowledge work. It is not — today — the best pick for raw terminal-style agentic coding sprints or computer-use tasks, where Astra and Opus 5.5 still hold the crown. The practical takeaway: when access opens, benchmark Argon on your actual workload. If your product is “agent that does the long boring job,” this might be your model. If it’s “agent that furiously iterates in a terminal,” the incumbents still lead.

Why cyber defenders get it first
This is the strangest and most interesting part of the launch. Google trained Argon to autonomously find, validate, and patch critical software vulnerabilities — and then decided the capability was potent enough that its first users would be vetted cyber defenders, plus Google’s own internal teams, who get the model without its cyber guardrails so they can use the full defensive toolkit.
The proof point Google led with is genuinely striking: Wiz, the cloud security company, is already using Argon through its Scan for Good initiative — a program that protects critical public infrastructure for free. In an early run, Argon uncovered a critical vulnerability in healthcare software used by hospitals worldwide, exposing sensitive personal information — something Google says previous frontier models had missed. That’s not a benchmark. That’s a hospital.
And Google isn’t pretending the dual-use problem away. Before broader release, it’s hardening four areas: misuse defenses for cyber and CBRN risks (including monitoring the model’s internal activations, under its Frontier Safety Framework), indirect prompt-injection resistance (Argon leads Gray Swan’s IPI benchmark), misalignment monitoring that watches the chain-of-thought and can stop execution mid-run, and sealed, isolated sandboxes for high-risk training. It’s also participating in the US government’s voluntary pre-release model access process.
“Safely releasing frontier capabilities at this level requires a phased approach.” — Google, on the Gemini 4 Argon rollout
Look — you can be cynical about safety theater, and some of it is. But the direction is real: the most capable models are increasingly launching behind velvet ropes, and “frontier” increasingly means “gated.” For founders, that’s a planning input, not just a news item. The models that matter most to your roadmap may not be publicly available on day one — or day one hundred.
What Google’s already doing with it inside the building
My favorite part of the announcement wasn’t a benchmark — it was the internal receipts. Thousands of Googlers are already using Argon, and Google shared four results that read like a preview of what “long-horizon agents” actually do to a company:
- Memory-optimization agents analyzed fleet-wide telemetry and applied fixes across Google’s data centers, freeing over 300 TiB of memory — with 500 TiB to 1 PiB projected.
- Agents took an existing Rust port of libgav1 (Google’s open-source video decoder), replaced 32,000 lines of SIMD code through rounds of profile-guided experiments, and produced a memory-safe decoder that runs 2.7x faster with identical output.
- Agents are migrating C/C++ codebases to Rust at serious scale — from tens of thousands of lines in core libraries up to 800,000+ lines for the Fuchsia Zircon kernel — under rigorous automated and manual auditing.
- Argon helped quantum researchers optimize the spacetime resources of bottleneck subroutines, beating a published baseline by 40% in a matter of minutes.
Notice the shape of all four: none of them is “wrote a haiku.” They’re all long, grinding, deeply unglamorous engineering jobs — the kind that eat quarters. If you’re a founder squinting at this and thinking “that’s my backlog” — yeah. That’s the product. The agent that does the migration nobody wants to do.
What this means for your startup: 5 moves to make this week
You can’t use Argon today. Fine. Here’s what you can do today:
- Don’t rebuild your roadmap around it — yet. No public access, no date, no confirmed input context window. Treat Argon as a strong signal about where pricing and capability are heading, not as a dependency. The founders who get burned are the ones who architect around a model they can’t touch.
- Do the token math now. Model your agent workloads at both $2/$10 (intro) and $4/$20 (standard), and price in cached input at $0.10 — because long-trajectory agents re-read context constantly, cache pricing is where your bill actually lives. Our breakdown of what it actually costs to start an AI business in 2026 has the framework; plug Argon’s numbers into it.
- Benchmark on your workload, not Google’s. Vendor charts are a starting filter, not a decision. When access opens, run your own evals — especially if your product lives in Argon’s lane (long-horizon coding, legal and finance knowledge work). And if you’re still assembling your stack, our list of AI tools every startup founder should use is the honest starting point.
- Watch the rollout order. Fairwind defenders now; paid API customers and Google AI Ultra subscribers next; everyone else eventually. If you’re already a Google Cloud or AI Ultra customer, you’re closer to the front of the line than you think — worth knowing which of your accounts would qualify.
- Treat 1M outputs as an architecture question. Long trajectories need long guardrails: checkpointing, human review gates, cost ceilings per run. And if the integration plumbing — eval harnesses, API routing, fallback models — feels like more than your team can carry alone, an AI integration studio like AISquadX can get it production-ready before access opens. The winners won’t be the teams that try Argon first; they’ll be the teams whose systems were ready when it arrived.
Argon vs the field, at a glance
Quick cheat sheet — prices per million input/output tokens, availability as of October 1, 2026:
| Model | Access today | Max output / response | Price (in / out) |
|---|---|---|---|
| Gemini 4 Argon | Fairwind defenders (phased) | 1M tokens | $2 / $10 intro, then $4 / $20 |
| Claude Opus 5.5 | API + clouds | 128K | $4 / $20 |
| Claude Fable 5.1 | API + clouds | 128K | $10 / $50 |
| GPT-6 Astra | OpenAI API | 128K | $10 / $50 |
The honest summary: there’s no single “best model” anymore — there’s a best model per job. If you haven’t picked your default stack yet, our ChatGPT vs Claude for Business breakdown walks through how to choose without overthinking it.
FAQ
When can my startup actually use Gemini 4 Argon?
Not yet — and Google hasn’t named a date. Right now it’s limited to trusted cyber defenders in the Fairwind Program. Google says access will expand to paid API customers and Google AI Ultra subscribers first, then developers, enterprises, and consumers. If you’re a Google Cloud customer, keep an eye on your console announcements; that’s likely where early access shows up first.
How much will Gemini 4 Argon cost?
Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off ($0.10). After the intro period it moves to $4/$20. For perspective: that’s half of Claude Opus 5.5’s price at launch, and an 80% discount against GPT-6 Astra. A maxed-out 1M-token response would cost $10 at intro pricing.
Is Gemini 4 Argon really better than GPT-6 Astra and Claude Opus 5.5?
On the tasks it was built for — long-horizon software engineering (DeepSWE v1.1: 77.9%), knowledge-work economics (Vals Index: 68.9%), long video understanding (91.7%) — yes, according to Google’s published comparisons. On FrontierSWE v2, Terminal-Bench 4.0, and computer-use tasks, Astra and Opus 5.5 still lead. The right answer for your startup is whichever wins on your workload, so run your own evals when access opens.
What is Google’s Fairwind Program?
It’s the Google security program through which Argon is getting its first rollout — a vetted group of cybersecurity defenders who receive the model first, including versions without the standard cyber guardrails so they can use its full vulnerability-finding toolkit. Think of it as a controlled field test with the people most qualified to stress-test a model this capable.
What should my startup do right now to prepare?
Four things: run the token math on your agent workloads at both intro and standard pricing; write evals based on your actual tasks so you can test Argon the day access opens; check which of your Google accounts (Cloud, AI Ultra) would put you early in the rollout queue; and architect for long trajectories now — checkpointing, review gates, and per-run cost ceilings. Preparation beats early access.
The Bottom Line
Gemini 4 Argon is the most interesting Google launch in a year — not because of any single benchmark, but because of what it signals. The frontier is moving toward long-horizon agents. Output length is the new battleground. Prices are still falling. And the most capable models increasingly launch behind a velvet rope.
You can’t use it today. But you can be ready — cost models done, evals written, architecture built for long trajectories. When that Fairwind gate lifts and the API opens, the startups that move first won’t be the ones that read the announcement. They’ll be the ones that prepared like it was already here.



