Gemini 4 Argon: Google's Comeback? Benchmarks and Pricing
⚡ Tech
Tech
AI
Google
Gemini

Gemini 4 Argon: Google's Comeback? Benchmarks and Pricing

Gemini 4 Argon has my attention. I compare benchmarks, introductory and later prices, and early reactions before testing Google's new model.

Uygar DuzgunUUygar Duzgun
Oct 1, 2026
7 min read

I have spent much less time with Google's models than with the other AI tools in my workflow. Gemini 4 Argon gives me a reason to change that. The business automation results look useful, and the announced introductory price sits alongside GPT-6.1 Sol rather than a premium-only model.

I want to test it when access opens. For now, this is a look at published evidence, pricing and early reactions, not a hands-on Argon review.

Source check: October 1, 2026. The cover is AI-generated editorial artwork. The charts below redraw cited numerical results; they are not screenshots of my own tests.

What is Gemini 4 Argon, and can I use it yet?

Google announced Argon on September 30, 2026, initially for trusted cyber defenders through Fairwind. Wider access will start with paid API customers and Google AI Ultra subscribers; the announcement gives no firm public release date. See Google's Gemini 4 Argon announcement.

I would wait for actual account access before buying a subscription specifically for Argon. An announced rollout is different from a model I can select today, and Google has not promised a Swedish availability date in that announcement.

The strongest reason for me to pay attention is the possibility of finishing longer tasks with less intervention. My useful test will involve existing code, awkward business rules and evidence that the final result works. A polished answer alone will not settle it.

Gemini 4 Argon pricing: the introductory rate has a catch

Google announces $2 input and $10 output per million tokens. Its pricing footnote sets the later rate at $4/$20, without naming the introductory period's end date. Cached input receives a 95% discount, giving an introductory $0.10 rate by calculation.

Model or pricing periodInput / 1M tokensOutput / 1M tokens
Gemini 4 Argon, introductory$2$10
Gemini 4 Argon, after introduction$4$20
GPT-6.1 Sol, Standard$2$10
Claude Opus 5.5, Standard$4$20
GPT-6 Astra, Standard$10$50

Sources: Google's announcement and expanded pricing footnote, GPT-6.1 Sol model pricing, GPT-6 Astra model pricing, and Claude Platform pricing.

Gemini 4 Argon introductory and later API prices compared with GPT-6.1 Sol, Claude Opus 5.5 and GPT-6 Astra
Gemini 4 Argon introductory and later API prices compared with GPT-6.1 Sol, Claude Opus 5.5 and GPT-6 Astra

Base USD rates checked October 1. Argon's prices are announced rates, not proof of general API access. The comparison excludes tools, cache writes, premium modes, regional charges and taxes.

For a simple budget example, one million uncached input tokens plus 100,000 output tokens across multiple short-context requests would cost $3 at Argon's introductory rates, $6 afterward, $3 with Sol, $6 with Opus and $15 with Astra. That is arithmetic using a fixed token mix, not a measured workload. The same task can consume different token counts across models.

OpenAI also increases full-request prices above 272,000 input tokens in one prompt; Anthropic lists standard pricing across Opus 5.5's context window. A base-price table therefore cannot tell you the bill for a huge repository session. My GPT-6.1 Sol vs Opus 5.5 comparison covers those tradeoffs.

I would budget an integration using Argon's later rate, then treat the promotion as temporary savings. That avoids building a business case around a discount whose expiry I cannot yet schedule.

Gemini 4 Argon benchmarks: strong automation, mixed coding

Business automation is the result I care about most

Zapier's AutomationBench 1.0.6 tests end-to-end work across 47 tools in six business functions. It checks the resulting environment with deterministic assertions rather than grading an agent's persuasive final message.

The selected configurations below retain their effort settings and fallback behavior.

Gemini 4 Argon, Claude Opus 5.5 and GPT-6 Astra AutomationBench scores and costs per attempted task
Gemini 4 Argon, Claude Opus 5.5 and GPT-6 Astra AutomationBench scores and costs per attempted task

Zapier reports Argon High at 51.29% and $1.70 per attempted task using later list prices; its introductory equivalent is $0.85. Opus 5.5 Max with default fallbacks scores 42.47% at $1.44, and Astra Max scores 41.40% at $1.73. Effort labels do not match compute budgets between providers.

Argon has the higher score in this comparison, but its ordinary task cost is not the lowest. Nor does a roughly 51% benchmark score justify unattended production access. I still want approval around consequential actions, logs I can inspect and a way to recover from partial completion.

That distinction matters for business automation. I care about the right record being updated and the right recipient receiving a message. An agent telling me it finished cannot substitute for those checks.

Coding results depend on the task

Gemini 4 Argon vs GPT-6 Astra and Claude Opus 5.5 on DeepSWE v1.1 and Terminal-Bench 4.0
Gemini 4 Argon vs GPT-6 Astra and Claude Opus 5.5 on DeepSWE v1.1 and Terminal-Bench 4.0

Selected scores from Google DeepMind's model evaluation report. This combines Google-computed Argon results with published rival results, not a uniform independent head-to-head run.

Argon reaches 77.9% on DeepSWE v1.1, against Astra's 74.1% and Opus's 74.2%. On Terminal-Bench 4.0, Opus leads this trio at 66.4%, followed by Astra at 58.2% and Argon at 57.4%. The report also puts Astra ahead on FrontierSWE v2: 65.5% against Argon's 55.0%.

Google generally reports its highest thinking settings and maximum or best available rival results. Harnesses and budgets matter. A lead on one repository benchmark does not guarantee a better fix in my codebase, especially when the difference is only a few percentage points.

I would test a bounded regression repair and a multi-file change separately. Both need behavioral checks and a review of the diff. A model could pass the test while adding an abstraction I do not want to maintain.

Independent knowledge-work evidence adds context

Vals Index 2.1, updated September 30, lists Argon at 68.90%, Opus 5.5 at 66.97%, Astra at 63.13% and Sol 6.1 at 61.15%. Vals combines finance, coding, legal and tax tasks with sector weights based on US GDP.

Its cost column complicates the ranking: $15.68 per test for Argon at $4/$20 token rates, versus $3.24 for Sol. Those are benchmark-specific costs, not quotes for my work. They show why I would keep a cheaper model in the evaluation even if another model tops the score column.

I would route tasks after measuring them. Easy recurring work and expensive failures should not automatically use the same model.

What are other people saying about Argon?

The early Gemini 4 Argon release discussion contains excitement, frustration about restricted access and reminders about the later price. These are launch reactions from individual users, not a representative survey or evidence that they have all tried Argon. Comments predicting weeks or months of waiting are speculation.

The more useful debate for me appears in the Rust community discussion of Google's codebase migrations. One commenter, nicoburns, argues that translation between languages can work better because the original design already exists. Others describe maintainability problems and the effort they spend correcting generated code.

That discussion concerns agent-assisted engineering broadly. It does not establish Argon's real-world reliability. I take the practical lesson seriously: compare behavior against the original, keep changes reviewable and measure the human work left afterward.

How I plan to test Gemini 4 Argon

I am curious enough to give Google a proper chance. I will start with tasks I can verify, rather than asking the model to take over everything at once.

A regression fix: same starting files, failing test and acceptance criteria across models.
A document-heavy task: an answer with traceable evidence and an explicit admission when information is missing.
A business workflow: a sandbox with expected state changes, deliberate ambiguities and checks for unwanted actions.

I want to record elapsed time, token cost, retries, correctness and review effort. The question is whether Argon reduces my total work, including checking and repairing the result.

Google's announcement also raises the output limit to one million tokens. That is a ceiling, not a target. I would still set budgets and stopping conditions; more available reasoning should earn its cost. My article on AI reasoning tokens and compute costs explains why I care about that tradeoff.

Gemini 4 Argon looks worth testing. I have not used Google enough to give it a fair place in my workflow yet. These results make me want to change that when I can access the model, while keeping Sol, Opus and Astra as real alternatives until my own tests tell me more.

FAQ

Is Gemini 4 Argon better than GPT-6 Astra or Claude Opus 5.5?+
There is no universal winner in the published evidence. Argon leads the selected AutomationBench and DeepSWE results, while Opus leads Terminal-Bench 4.0 and Astra leads FrontierSWE v2 in Google's evaluation table. Effort, harnesses, fallback behavior and task costs differ.
Is Argon's introductory API price permanent?+
No. Google announces $2 input and $10 output per million tokens, with $4/$20 afterward. Its announcement does not specify the introductory period's end date. Budget long-running integrations against the later rate and confirm billing terms when access opens.
Does the AutomationBench score mean my business can run unattended?+
No. Benchmark performance does not guarantee a correct business outcome. Evaluate in a sandbox, check final records and recipients, and keep consequential actions behind approval. Zapier reports costs per attempted task, including failures.
Are these benchmark images from your own tests?+
No. The charts redraw values from Google DeepMind's evaluation report, Zapier's leaderboard and official pricing pages. They identify source dates, configurations and pricing assumptions. The hero is an AI-generated editorial illustration.
Have you tested Gemini 4 Argon yourself?+
Not yet. I have spent less time with Google's models and want to test Argon when I can access it. My planned comparison will measure correctness, elapsed time, token cost, retries and the review work left afterward.
✻

Recommended for you

GPT-6.1 Sol vs Opus 5.5: Benchmarks, Price, Best Uses

GPT-6.1 Sol vs Opus 5.5: Benchmarks, Price, Best Uses

GPT-6.1 Sol promises more agent work for the budget. I compare its benchmarks, API price and early user reactions with Claude Opus 5.5.

7 min read
AI Reasoning Tokens: More Tokens, Better Results, More Power

AI Reasoning Tokens: More Tokens, Better Results, More Power

OpenAI and Anthropic make the tradeoff visible: better AI answers often need more reasoning tokens, more inference compute, and more electricity.

7 min read
Claude Fable 5.1 Review: Benchmarks, Pricing, vs GPT-5.6 Sol

Claude Fable 5.1 Review: Benchmarks, Pricing, vs GPT-5.6 Sol

Claude Fable 5.1 keeps Fable 5's price, cuts cache reads 75% and targets the laziness complaints. Benchmarks vs Opus 5 and GPT-5.6 Sol, plus an evening on my stack.

11 min read