AI Tools Police
Reader-supported: we may earn a commission from links, at no cost to you. Rankings are never sold. How we investigate →

Frontier model review · Google DeepMind · the announcement, the evaluation methodology, the Fairwind pages, 12 outside evaluators and 24 dated user reports

Gemini 4 Argon review: launched at GPT-6 Sol's price, priced later at Opus 5.5's, and almost nobody outside Google can run it yet

Google prices Gemini 4 Argon at $2 and $10 per million tokens for an introductory period it does not date, then $4 and $20, and offers it only to a set of Fairwind security partners. Four of the eight benchmark claims in its announcement match the owners' own leaderboards, each run there at high effort; on Vals' run of Harvey's legal benchmark Argon is fifth, and three have no entry on their owners' boards. There is no model card, no safety report and no API model ID. Google's documents, 12 outside evaluators and benchmark owners, and 24 dated user reports, read on 1 October 2026.

By Mucahit Kaya · Founder and EditorOct 1, 2026~10 min read

Our verdict

Treat Gemini 4 Argon as a model to watch and not yet one to plan on: outside evaluators confirm four of Google's eight launch claims at high effort, but no customer outside a set of Fairwind security teams can get access, Google gives no date for anyone else, and any budget written today should assume the $4/$20 price that follows an introductory period of unstated length.

Gemini 4 Argon specs

Developer
Google DeepMind. Announced by Koray Kavukcuoglu, SVP, Google DeepMind and Chief AI Architect, Google
Announced
30 September 2026, per Google's announcement
Access
A set of Fairwind Program partners, inside their security teams only, per DeepMind's Fairwind page. No date for paid API customers, Google AI Ultra, enterprises or consumers, as of 1 October 2026
Model ID
Not published as of 1 October 2026. No Google document prints one, and no Gemini API or Google Cloud page lists Argon
Input context window
Not published by Google as of 1 October 2026. Vals and Artificial Analysis state 1M tokens; Google's own GraphWalks row ran inputs of up to 1M, which is not a published limit
Maximum output
1M tokens, up from 64K, per Google's announcement. Vals ran Argon at 262k; Artificial Analysis reached 1M through a continuation feature that resumes a response across calls
Knowledge cutoff
Not published as of 1 October 2026
Input and output
Not listed by Google as of 1 October 2026. Its benchmarks include chart and video understanding. Artificial Analysis's model page says text and image input, its launch article text, image, video and speech input, both with text output
Model card and safety report
None published as of 1 October 2026
Zero data retention
Supported when accessed as a managed model on Gemini Enterprise, per the Fairwind FAQ; the Google Cloud page that answer links to does not name Argon
What Google says it is for
Long-horizon software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense, per the announcement

Gemini 4 Argon pricing

Input, introductory
$2.00 per 1M tokens, per the announcement. Same list price as GPT-6 Sol, GPT-6.1 Sol and Claude Sonnet 5.5
Output, introductory
$10.00 per 1M tokens, per the announcement. Same list price as GPT-6 Sol, GPT-6.1 Sol and Claude Sonnet 5.5
Cached input, introductory
95% off the input price, per the announcement, which is $0.10 per 1M (our arithmetic; Artificial Analysis states the same). Collinear's CWE-bench run priced it at $0.20
After the introductory period
$4.00 input and $20.00 output per 1M tokens, per the announcement's footnote. Same list price as Claude Opus 5.5. Length and end date of the introductory period: not stated
Cached input after the introductory period
Not stated. $0.20 per 1M if the 95% discount carries over (our arithmetic)
Cache writes, batch, long-context and tool pricing
Not published as of 1 October 2026
Pricing note
Verified on 1 October 2026 against Google's announcement only, the one Google document that prints an Argon price. The Gemini API pricing page and Google Cloud's Agent Platform pricing page, both read the same day, have no Argon row

Rate card read at the vendor’s own documentation on Oct 1, 2026.

Gemini 4 Argon, which Google DeepMind announced on 30 September 2026, is "rolling out to a set of trusted cyber defenders through our Fairwind Program", in Google's words. Developers, enterprises and consumers follow with no date. The introductory price is $2 per million input tokens and $10 per million output tokens, then $4 and $20 once an introductory period of no stated length ends.

Google says the wider release will start "with paid API customers and Google AI Ultra subscribers". The introductory price is the list price of GPT-6 Sol, and the later one, set in a footnote, is the list price of Claude Opus 5.5.

Of the eight benchmark results the announcement cites, four match the owner's own leaderboard, where each owner ran Argon at high effort: the Vals Index, Vals Finance Agent v2, Zapier's AutomationBench and CWE-bench v1. On a fifth, Harvey's Legal Agent Benchmark, Vals' run puts Argon fifth of 73, behind four Meta Muse Spark models and ahead of every Anthropic and OpenAI model listed. Three cannot be checked against their owners: DeepSWE v1.1, LVBench and Gray Swan's prompt-injection benchmark. Google has published no model card, no safety report and no API model ID.

No prompt was run through Gemini 4 Argon for this review; nobody on our side has access. Everything here comes from documents read on 1 October 2026: Google's announcement, DeepMind's Gemini and Fairwind pages and its five-page evaluation methodology; the pages of 12 outside evaluators and benchmark owners, five of which list Argon; and 24 dated reports from Hacker News, Reddit and GitHub.

Can you use Gemini 4 Argon today?

Only if your organisation is one of the Fairwind partners Google has given it to, and then only inside its security teams. On 1 October 2026 Argon was absent from every Gemini API, pricing and Google Cloud page we checked, so a paid API account, a Google AI Ultra plan or an enterprise Cloud contract does not reach it.

Partners can run Argon on its own or inside CodeMender, Google's code-security agent. The announcement names one partner, Wiz, which "is already using Argon for cybersecurity defense through its Scan for Good initiative".

Who is in the Fairwind Program?

More than 650 organisations, by DeepMind's count, but only some of them have Argon. The program page says "We currently work with over 650 partners globally." and "A set of Fairwind Program partners get exclusive access to Gemini 4 Argon", without saying how many.

Priority goes to governments and national cyber authorities, critical infrastructure operators and core technology platforms. The terms are narrow: "Organizations may only grant Gemini 4 Argon access to internal cybersecurity, incident response, or penetration testing teams, and must track employee access and use." Partners also accept phishing-resistant multi-factor authentication, a ban on resale and a background check, and "Malicious tasks such as creating malware are not permitted."

A security team asks through a form on DeepMind's program page: "To request access to the Fairwind Program, fill in our form." Google gives no processing time: "We will review and respond to eligible partners who meet our criteria as soon as we can." Organisations that do not qualify are told "you can still protect your code using CodeMender with publicly available models".

Does Google AI Ultra include Gemini 4 Argon?

Not yet. Google names Ultra subscribers and paid API customers together as the first group after the Fairwind cohort, with no date for either. Eduardo1502 on Reddit's r/GeminiAI, 30 September 2026 reported: "I have ultra an no access to Gemini 4 Argon".

One reader took the rollout as Ultra first: omale1, posting as a Pro subscriber on r/singularity on 1 October 2026, wrote that "Argon goes to Ultra first". Google's sentence puts paid API customers in the same group.

When will developers get Gemini 4 Argon?

Google gives no date. The Gemini API release notes, last updated on 23 September, have no Argon entry, and Google Cloud's model list shows only 3.8 Flash Cyber under its cyber heading. sleegme, in a GitHub issue opened on 30 September 2026, found the same: "No public model id or API docs yet (checked 2026-10-01)."

Artificial Analysis's model page says Argon "is available via API through 1 provider", while its launch article says the model "is currently being rolled out to selected users and is not publicly available".

Does Gemini 4 Argon have an API model ID or a model card?

Neither, as of 1 October 2026. No Google document prints an API identifier, the Gemini API models page (updated 1 October) lists nothing newer than the gemini-3.8 family, and DeepMind's model card index has no Argon entry.

The safety paperwork is missing too. DeepMind's frontier safety page says evaluation results appear "in our safety reports below, and summarized in our model cards", but lists reports only for Gemini 3.7 Flash and Gemini 3 Pro. There is no Frontier Safety Framework report, Critical Capability Level finding or technical report for Argon, and no chemical, biological or offensive-cyber result. The announcement says the model "is designed to refuse harmful requests while preserving legitimate, dual-use scientific research" and that its safeguards "underwent robustness testing by internal and external red teams", with no result or team name.

Is $2/$10 what Gemini 4 Argon will cost?

Only for now. Google's footnote reads: "After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply." No Google document says when that happens, and the announcement is the only Google document with an Argon price; the Gemini API and Google Cloud pricing pages have no Argon row.

The introductory rate is also the list price of Claude Sonnet 5.5. OpenAI replaced GPT-6 Sol with GPT-6.1 Sol on 29 September at the same $2/$10, per Artificial Analysis, Vals and arena.ai, so the comparison in this review's title holds for the current Sol.

What changes after the introductory period?

The per-token bill doubles, and two outside boards already rank Argon at the later price. Vals' $15.68 per Vals Index test is priced at $4/$20, and Zapier ranks at standard list price, $1.70 per AutomationBench task at High, printing the promotional $0.85 only in a footnote: "These Gemini models have promotional pricing; ranking and Cost / task use standard list pricing."

Artificial Analysis, with Argon at high, gives $1.99 per Intelligence Index task now and $3.98 afterwards. Against rivals at max on the same index, the later figure sits above GPT-6 Astra at $3.26, below Claude Opus 5.5 at $5.98 and far above GPT-6.1 Sol at $0.72.

Nobody outside Google knows the end date. Artificial Analysis says the discount runs "for at least one month" without a source, and in the same article that "Google has not yet confirmed the promotion end date". funforgiven on r/singularity, 1 October 2026 mentioned a promotion "running through December 31, 2026". No Google document gives that date for Argon; Google's Cloud pricing page gives it for the Gemini 3.8, 3.7 and 3.6 Flash introductory price. mike1858, adding Argon to the splitrail usage tracker on 30 September 2026, noted that Google gives no end date and declined to guess one.

What is the cached-input price?

$0.10 per million tokens at the introductory price, by our arithmetic from "cached input tokens priced at 95% off input token price"; Artificial Analysis states the same. Collinear's CWE-bench page used $0.20, a 90% discount, with Argon "priced from its token counts at $2 per million input, $0.20 per million cached input and $10 per million output tokens".

Google gives no cached rate after the introductory period ($0.20 if the 95% carries over, our arithmetic) and no cache-write price, although Artificial Analysis puts $0.98 of Argon's $1.99 per index task, at high, on cache writes.

Which of Google's Gemini 4 Argon claims check out?

Four of the eight the announcement cites, each on its owner's page at high effort: the Vals Index, Vals Finance Agent v2, AutomationBench and CWE-bench v1. On Harvey's Legal Agent Benchmark, Vals' run puts Argon fifth of 73, and DeepSWE v1.1, LVBench and Gray Swan's Indirect Prompt Injection benchmark have no owner entry to check.

BenchmarkGoogle's figureOwnerArgon on the owner's pageOwner's settingResult
DeepSWE v1.177.9%, Google's own runDatacurveNo entry among 28 modelsnoneCannot be checked at the owner
Vals Index68.9%Vals AIFirst of 41, 68.90%highMatches
Vals Finance Agent v265.4%Vals AIFirst, 65.40%highMatches
Harvey's Legal Agent Benchmark19.6%Harvey, run by Vals AIFifth of 73 on Vals, 19.58%; none on Harvey's pageshighScore matches; fifth overall, ahead of every Anthropic and OpenAI model
AutomationBench51.3%ZapierFirst of 121, 51.29%HighMatches
LVBench91.7%LVBench teamNo entry; newest row dated 29 May 2025noneCannot be checked at the owner
CWE-bench v168%Collinear AITied first of 11, 68%highMatches
Gray Swan IPI0.7% attack success at 15 attempts, from a chart imageGray Swan AINo public leaderboardnoneCannot be checked at the owner

Google's figures from its announcement and the table on DeepMind's Gemini page. Owners' pages read 1 October 2026 between 12:20 and 12:30 UTC.

The eight claims follow what Google says Argon is for. Its announcement says the model "delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense": DeepSWE covers the coding, the Vals Index, Finance Agent v2, Harvey's benchmark and AutomationBench the knowledge work, CWE-bench and Gray Swan the security side, and LVBench the video understanding Google also names.

Google calls Argon "the leading model on the Vals Index" and claims "similarly leading performance" on Finance Agent v2 and Harvey's benchmark. Vals' comparator scores moved between 28 and 30 September without a note, so its figures here are as read on 1 October.

Is 77.9% on DeepSWE v1.1 a new state of the art?

It cannot be checked at the owner: Datacurve's leaderboard has no Argon entry and tops out at 74%, shared by three models. Artificial Analysis's own run puts Argon at 78.8%. Google's "new state of the art on DeepSWE v1.1 (77.9%)" is its own run, "self computed, using a mini-swe agent harness", set against GPT-6 Astra's leaderboard figure and Claude Opus 5.5's system-card figure.

Artificial Analysis ran DeepSWE v1.1 inside its Coding Agent Index v1.5 with Argon at high in Google's Antigravity CLI, and its 78.8% is the best of its 31 configurations, ahead of GPT-6.1 Sol at xhigh (73.2%), Claude Sonnet 5.5 at max (72.0%) and Claude Opus 5.5 at max (68.4%). The harness differs from Datacurve's, so this supports the direction of Google's claim, not its exact number.

Does Gemini 4 Argon rank first on AutomationBench?

Yes. Zapier's leaderboard 1.0.6 lists "Gemini 4 Argon (High)" first of 121 rows at 51.29%, which matches Google's "Argon ranks #1 with a score of 51.3%", and Argon at Medium second at 50.08%. Next come Claude Sonnet 5.5 at Max (44.75%), Claude Opus 5.5 at Max (42.47%) and GPT-6 Astra at Max (41.4%). Zapier names the setting; Google does not.

Finance is one Zapier domain Argon does not lead: Sonnet 5.5 at Max scores 50.0% there, Argon at Medium 49.17%. Artificial Analysis's AutomationBench-AA, where Argon at high scores 77.5%, is a partial-score metric and does not compare with Zapier's figure.

Does Gemini 4 Argon lead Harvey's Legal Agent Benchmark?

It leads every Anthropic and OpenAI model, not the board. On Vals' held-out run, the only published run that includes Argon, it scores 19.58% at high and places fifth of 73, behind four Meta Muse Spark models led by Muse Spark 1.2 at 25.42%, and ahead of every Anthropic and OpenAI model listed. Harvey itself has published nothing on Argon.

On the same run GPT-6 Astra and GPT-6.1 Sol score 5.42%, Opus 5.5 3.75% and Sonnet 5.5 2.92%, each at max. Google's table shows a matching 19.6% but compares it only with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5, and its methodology says "Harvey's Legal Agent Benchmark results are sourced from Vals AI."

Does Gemini 4 Argon tie for first on CWE-bench v1?

Yes, at 68% programmatic pass@1, a three-way tie with Grok 4.7 and GPT-6 Astra on Collinear AI's leaderboard, where "Each model gets four rollouts on each of the 120 tasks, at high reasoning." Opus 5.5 is at 67% and GPT-6 Sol at 52%. Grok 4.7 is not in Google's main table.

Argon trails on the board's other columns: 75% at pass@4 against Grok 4.7's 81% and Opus 5.5's 79%, and 62% on the judge-panel pass@1 against Opus 5.5's 67%. Its $6.63 per rollout, priced at the introductory rate, is the highest of the top four; Opus 5.5's is $0.79.

Is Gemini 4 Argon the most resistant model on Gray Swan's prompt-injection benchmark?

That cannot be checked: Gray Swan publishes no public leaderboard for its Indirect Prompt Injection benchmark, and none of its pages read on 1 October mentions Argon. Google's claim of "leading in prompt injection robustness" rests on its own chart image, read off the image: 0.7% attack success for Argon at 15 attempts, against 1.0% for Claude Opus 5.5 and Claude Fable 5.1, 8.5% for GPT-6 Astra and 10.1% for GPT-6 Sol. Google's chart names no setting for any model. The margin over the two Claude models is 0.3 points (our arithmetic); values at one and ten attempts are drawn but not printed.

Can Google's internal results for Gemini 4 Argon be checked?

No. Google's internal examples come with no outside document: a quantum subroutine where Argon "beat the published baseline by 40% in a matter of minutes", memory work "freeing up over 300 TiB of memory once rolled out", and Rust migrations that Google says are still being audited. Its two vulnerability-discovery charts rest on internal benchmarks, Google's and Wiz's, and compare Argon only with Gemini 3.8 Flash Cyber: 85.8% against 71.0% and 70.9% against 58.2%, read from the images.

Those charts are most of what Google says about how Argon differs from Gemini 3.8 Flash Cyber, the earlier Fairwind model. The announcement says "Gemini 4 Argon demonstrates impressive leaps in vulnerability discovery over 3.8 Flash Cyber.", and Google's Gray Swan chart shows 6.0% attack success for 3.8 Flash Cyber at 15 attempts against Argon's 0.7%. The Fairwind launch post of 2 September 2026, which never mentions Argon, sells 3.8 Flash Cyber on cost, "at a fraction of the operating cost of traditional frontier models", and Google Cloud lists it while Argon is absent. No Google document gives both models' prices or compares them on anything else.

Which settings does Google not state for Gemini 4 Argon?

The effort level, for every figure. The announcement names no setting, and the methodology says only that results are run "with the highest thinking settings", without naming a level. Every outside owner that prints one ran Argon at high, which Artificial Analysis calls "the highest available", so Google's phrase probably means high; Google does not say so.

Row in Google's tableArgon's run, per Google's methodology
DeepSWE v1.1 and Terminal-Bench 4.0Google's own runs; rivals from leaderboards or system cards
Terminal-Bench Science 0.1Google's own run "with 6x verifier timeout"
OSWorld-2.0Best of three: "maxed over 3 runs with a single attempt per run"
LVBench1 frame per second for Gemini; 800 frames for GPT-6 Astra, 600 for Opus 5.5, 300 for Fable 5.1
Smaller benchmarks"To reduce variance, we average over multiple trials for smaller benchmarks." Which ones, and how many trials, is not stated

Gemini 4 Argon Model evaluation, pp.2-4, read 1 October 2026.

Two of those rows still favour a rival. With six times the verifier timeout, Argon scores 57.6% on Terminal-Bench Science against GPT-6 Astra's 68.1%, and its best of three on OSWorld-2.0, 69.2%, trails Astra's 72.6%. The PDF also dates its results both "as of September, 2026" and "as of October, 2026".

What does "without cyber guardrails" mean for Gemini 4 Argon?

Google does not define it. Its one sentence reads: "For trusted defenders and our own internal teams at Google, we'll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities." No Google document says which refusals or classifiers are off, or whether the Argon Fairwind partners use is the model paid API customers will later get.

The restraints Google describes are the contractual Fairwind terms above. Before broad release, it says it is still strengthening safeguards in four areas: misuse, including cyber and CBRN attacks; indirect prompt injection; misalignment, through mitigations that "monitor Argon's chain-of-thought and actions and stop execution when necessary"; and sealed sandboxes for high-risk training and evaluation. No metric goes with any of them, and no Hacker News or Reddit comment collected for this review discussed the guardrails sentence as of 1 October 2026.

No government body has said it evaluated Argon either. Google writes: "We are actively engaged in the U.S. government's voluntary process for pre-release model access while we gradually expand access." It names no agency and reports no outcome. NIST's Center for Advancing Innovation and Standards for Super Intelligence (CAISSI, formerly CAISI), whose remit includes voluntary agreements with developers, had published nothing on Argon by 1 October, nor had the UK AI Security Institute or METR; Google does not say CAISSI is the body it means.

Gemini 4 Argon vs GPT-6 Astra and Claude Opus 5.5

On Google's own 19-row table, Argon is top on 13 rows, tied on one and behind on five, each time behind GPT-6 Astra or Claude Opus 5.5. On Artificial Analysis's Intelligence Index v4.3.2, with Argon at high and both rivals at max, Argon scores 52.56, level with Astra's 52.67 and 5.1 points below Opus 5.5's 57.62 (our arithmetic).

Google's table leaves out GPT-6 Sol and Claude Sonnet 5.5, the two models whose list price Argon's introductory price matches. On arena.ai's Text board "gemini-4-argon-high" is first at 1525, marked preliminary; on the WebDev board it is eighth, with Opus 5.5 at max first.

Where does Google's own table put a rival first?

On five of 19 rows, none of which the announcement mentions. Google reports each rival at the maximum reasoning setting its provider published, or the best available result where none was published (methodology, p.2).

RowGemini 4 ArgonTop score in Google's table
FrontierSWE v255.0%GPT-6 Astra, 65.5%
OSWorld-2.0, offline subset, partial score69.2%GPT-6 Astra, 72.6%
Terminal-Bench Science 0.157.6%GPT-6 Astra, 68.1%
Terminal-bench 4.057.4%Claude Opus 5.5, 66.4%
PostTrainBench45.3%Claude Opus 5.5, 49.3%

DeepMind's Gemini model page, read 1 October 2026; the same values appear on page 5 of the methodology PDF. Rivals at the maximum reasoning setting each provider published, or the best available (methodology, p.2).

Outside runs agree on Terminal-Bench 4.0: Vals has Argon at high fifth at 57.58%, behind Opus 5.5 at max on 65.15%, and Artificial Analysis has Argon at high on 57.1% against Opus 5.5 at max on 59.6%.

What do Artificial Analysis's numbers add?

A cost and a hallucination figure: Argon at high costs $1.99 per index task against GPT-6 Astra's $3.26 at max, and its hallucination rate is 15.1% against Astra's 51.3%. Argon was run at high, the only Argon setting Artificial Analysis lists, and each rival at max.

Model and settingIntelligence Index v4.3.2Cost per index taskHallucination rate
Gemini 4 Argon, high52.56$1.99 ($3.98 at $4/$20)15.1%
GPT-6 Astra, max52.67$3.2651.3%
Claude Opus 5.5, max57.62$5.9858.6%
Claude Sonnet 5.5, max55.98$7.6247.0%
GPT-6.1 Sol, max51.83$0.7254.3%
GPT-6 Sol, max47.63$1.0460.1%

Artificial Analysis model pages, from each page's chart data, read 1 October 2026. Argon's $1.99 uses the introductory $2/$10; $3.98 is Artificial Analysis's figure at $4/$20. Hallucination rate is from AA-Omniscience.

Argon at high also has the lowest accuracy of the six on the same AA-Omniscience evaluation, 49.9%, which Artificial Analysis reads as Argon being "much more likely to acknowledge when it does not know an answer". Its lower cost than GPT-6 Astra comes from the token price: at high it averages 61.6k output tokens per task, 2.3 times Astra's 27.2k at max (our arithmetic). On the Coding Agent Index v1.5, Argon at high is third of 31 at 63.8, behind Sonnet 5.5 at max (68.4) and Opus 5.5 at max (66.0).

What does Gemini 4 Argon's 1M output limit change?

On Google's figures, one response can run to 1M tokens, "up from the previous 64K tokens". Neither outside run used 1M in one request: Vals ran Argon with "262k max output tokens", and Artificial Analysis reached 1M only through Long Decode Continuation, "a new Gemini API feature that pauses long responses and resumes them across follow-up calls", which no Google API page read on 1 October describes.

A full 1M-token response would cost $20 in output alone at the later price, $10 at the introductory one (our arithmetic). Google publishes no input context window; Vals and Artificial Analysis state 1M, and Google's GraphWalks row ran inputs "between 256k and 1M tokens", which records a run, not a published limit. murkt on Hacker News, 30 September 2026 wrote, without a source: "Input token limit is 1M for Gemini models for a long time." MisterBiggs, 1 October 2026 estimated that "At Gemini 3.8 Flash speeds Argon would be outputting for 1hr 10mins", the commenter's own arithmetic; no outside speed figure for Argon exists.

What do users say about Gemini 4 Argon?

None of the user reports found by 1 October 2026 describes what Argon did on a task. Across two Hacker News threads, Reddit and GitHub, four accounts reported looking for Argon and not finding it, and three claimed use without saying what it did.

Has anyone outside Google used Gemini 4 Argon?

Outside evaluators have: Vals, Zapier, Collinear, Artificial Analysis and arena.ai all ran Argon at launch, and none of their pages says how it got access. No user reports a task result. clduab11 on r/GeminiAI, 1 October 2026 asked: "Who is doing these benchmarks if it's only released for people who qualify for Fairwind?"

Four accounts looked for Argon and did not find it. Besides Eduardo1502's Ultra account and sleegme's API check, both cited above, I-Kernel, in a post on r/GeminiAI on 1 October 2026, wrote: "Even with corporation and credentials I only found 3.8 Flash???" PuffyVaccination_9, 1 October 2026, replied: "my company has legit enterprise access and we're still seeing 3.8 flash everywhere, no sign of argon at all". dom96 on Hacker News, 30 September 2026 named what the closed release blocks: "I'd love to run it on my benchmark but alas, Google not making it public prevents this."

Three accounts claim use, and none can be checked. asdfman123 on Hacker News, 30 September 2026: "Argon is doing my job for me while I'm writing this comment." kridsdale1, replying on 30 September 2026: "We've been super impressed by Argon for a few weeks." Neither names an employer. ThatOtherSwimmer on r/GeminiAI, 1 October 2026 described having "worked on model eval at Google and used Argon". None says what the model did or where it failed.

What do users make of the rollout and the price?

These are opinions, not reports of use. modeless on Hacker News, 30 September 2026: "When I said I was tired of Google launching waitlists I didn't think they would respond by simply not having a waitlist." readthethreshold on r/artificial, 30 September 2026 named the evidence problem: "with fairwind still gated most people cant even poke the same model those charts were run on". FateOfMuffins on r/singularity, 30 September 2026 compared it with a rival's limited release, "similarly restricted as Project Glasswing from Anthropic in April for Mythos", a comparison not checked here.

Two withheld judgement on the numbers. jjcm on Hacker News, 30 September 2026: "I'll wait for hands on before getting too hyped that Google is back." MaximumIntention on r/singularity, 1 October 2026: "It's more cherrypicking of benchmark results IMO."

On price, readers made the comparisons in this review's title. ehsankia on Hacker News, 30 September 2026: "It's exact same price as Sol 6.1 announced yesterday." GodelNumbering, 30 September 2026: "So, they are basically offering opus 5.5 pricing." nl, 1 October 2026, using Artificial Analysis's figures for Argon at high and Sol at max: "Sol 6.1 scores one point less than Gemini 4 on intelligence AND costs less than half ($0.72 vs $1.99) per task."

Wrong prices are circulating too. Emotional-Cut2952 on r/ClaudeAI, 30 September 2026 wrote that the price would rise to "$4/$25", and AlyoshaV on r/singularity, 1 October 2026 gave the introductory price as "$2/$20". Google's figures are $2 and $10, then $4 and $20.

What is still unknown about Gemini 4 Argon?

When anyone outside Fairwind can use it, and when the $2/$10 price ends. Every absence named above stood at 12:30 UTC on 1 October 2026.

The record so far: on 2 September 2026 Google launched Fairwind with Gemini 3.8 Flash Cyber; on 30 September it announced Argon, and Vals, Artificial Analysis and arena.ai posted results the same day; on 1 October Argon was still missing from the Gemini API docs and pricing page, both updated that day. Google's pages and every outside board cited here get a fresh read at least once a month, plus an extra read whenever access widens or a new Argon result appears, and any change that follows is dated where it lands on the page.

Gemini 4 Argon: independent evaluations

Five organisations outside Google had published results on Gemini 4 Argon by 12:30 UTC on 1 October 2026, and every one that names a setting ran it at high. Vals AI ranks it first of 41 on Vals Index v2.1 at 68.90%, $15.68 per test priced at the later $4/$20, first on Finance Agent v2 at 65.40%, and fifth of 73 on its run of Harvey's Legal Agent Benchmark at 19.58%, with Argon capped at 262k output tokens and its comparators at max. Zapier's AutomationBench 1.0.6 lists Argon (High) first of 121 at 51.29%, $1.70 per task at standard list price and $0.85 at the promotional price. Collinear AI's CWE-bench v1, at high reasoning in Google's Antigravity harness, has Argon in a three-way tie for first at 68% programmatic pass@1, $6.63 per rollout priced at $2/$10. Artificial Analysis lists only a high setting for Argon: 52.56 on Intelligence Index v4.3.2 at $1.99 per task on $2/$10 ($3.98 at $4/$20), against GPT-6 Astra at 52.67 and Claude Opus 5.5 at 57.62, both at max; its Coding Agent Index v1.5 run, Argon at high in the Antigravity CLI, gives 78.8% on DeepSWE v1.1. arena.ai lists gemini-4-argon-high first on its Text board at 1525, marked preliminary, and eighth on WebDev, showing the $2/$10 price. Datacurve (DeepSWE), the LVBench team, Gray Swan, Harvey, Epoch AI, ARC Prize, METR, the UK AI Security Institute and NIST's CAISSI had published nothing on Argon. None of the five says how it obtained access.

What we read

What we did not read

  • A Gemini 4 Argon model card, Frontier Safety Framework report or technical report. None exists to read: the model card index, the list of safety reports and every likely address for a card, a safety report or a technical report returned no Argon document on 1 October 2026.
  • An Argon page in the Gemini API docs, an Argon row on any Google pricing page, an API model ID or an Argon release note. None exists; the price on this page comes from the announcement alone.
  • Values that Google prints only inside chart images. The Gray Swan attack rates, the two vulnerability-discovery scores and the CWE-bench bars were read off the images or their alt text and are marked as such where they appear; the Gray Swan one-attempt and ten-attempt values are drawn but not printed, so none is given.
  • Most rows of Google's 19-row benchmark table against their own owners: FrontierSWE v2, PostTrainBench, Terminal-Bench Science 0.1, RiemannBench, LABBench 2, both GraphWalks rows, Agent's Last Exam, OSWorld-2.0 and Chartography, plus Vibe Code Bench and Terminal-Bench 4.0 beyond the values inside Vals' index data.
  • Any Wiz document on the Scan for Good finding or on the Wiz penetration-testing benchmark; both are described here as Google describes them.
  • The Fairwind application form itself, and any Gemini 3.8 Flash Cyber model card or price page; Argon is compared with 3.8 Flash Cyber only where Google's Argon and Fairwind documents do so.
  • Launch-day press coverage, which is not used as evidence, including a reported Bloomberg story on employee doubts and the quotations from Google staff in news articles.
  • X, including the post by a Google employee repeating the introductory price, and Anthropic's or OpenAI's pages on their own limited cyber releases, which one Reddit comment cited here mentions.
  • Reddit comments on reddit.com itself: Reddit refused scripted reads, so account names, times and links come from a public archive of Reddit's own fields, and an edit or deletion after archiving would not show.
  • G2, Trustpilot and Capterra, which carry no reviews of a model without a public release.
  • The model itself. We ran no prompt through Gemini 4 Argon for this page, have no access to it, and report no result of our own.
  • Anything published after 12:30 UTC on 1 October 2026.

Disclosure. Our writing workflow runs on Anthropic models, including Claude Opus 5.5, a direct competitor of the model reviewed here. So every judgement on this page rests on Google's own documents and on outside evaluators, each named with its setting.

We run no hands-on tests. This review is built from the lab’s own published documents, independent evaluations by other organisations, and dated user reports, each named above with the date we read it. How we investigate →