Models·8 min read
By BitsMindsSource: Google

Gemini 4 Argon: Benchmarks, Price and Who Can Use It

Google's Gemini 4 Argon tops 13 of the 19 tests in Google's own table, but so far only a group of cyber defenders can use it. Paid API customers and Google AI Ultra subscribers come next, at an introductory $2 per million input tokens and $10 per million output tokens. Artificial Analysis already puts it level with GPT-6 Astra at 53 points, for 61% of Astra's cost per task.

Gemini 4 Argon Editorial illustration: a thick glass element tile for argon, element 18, glowing the violet of an argon discharge, with Google’s Gemini sparkle inlaid at its centre, standing on a machined plinth against a Gemini blue-to-coral field. The tile is a metaphor for the model’s name, not a Google product. GOOGLE DEEPMIND Gemini 4 Argon BITSMINDS.COM 18 Ar ARGON 39.948
Share:

Gemini 4 Argon is Google’s new frontier model, announced on 30 September 2026 by Koray Kavukcuoglu, who runs Google DeepMind day to day. It is the first model of the Gemini 4 generation, and Google pitches it at long, multi-step work in software engineering, finance, law and cyber defence. You cannot use it yet. Google’s announcement says it is “rolling out to a set of trusted cyber defenders through our Fairwind Program”, with paid API customers and Google AI Ultra subscribers next and no date for either. Argon is the name that circulated in September leaks, which we treated as unverified in our Gemini 4 explainer. It turned out to be the product name.

Price, limits and a slow rollout

Argon will launch at an introductory $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper, which works out to $0.10 per million. After the introductory period the price doubles to $4 and $20, and Google has not said how long that period lasts. The launch sticker matches GPT-6.1 Sol, which OpenAI released the day before at the same $2 and $10.

The biggest technical change is the output limit. Argon can write up to 1M tokens in a single response, up from 64K before, which Google says gives it room to think through hard problems in one pass. Google’s post does not state a context window. Artificial Analysis lists it at 1M tokens.

The rollout is unusually cautious. Google says it is taking part in the U.S. government’s voluntary process for pre-release model access, and that it will keep collecting feedback from early testers “as we iterate on guardrails” before opening Argon to developers, businesses and consumers. It comes days after OpenAI scrapped GPT-6.1 Astra for failing its own safety tests.

Google’s benchmark table

Google compared Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 19 tests. Argon has the top score on 13, ties Astra on one and trails on five. The clearest wins are in knowledge work. On Harvey’s legal agent benchmark Argon scores 19.6%, against 6.7% for the next-best model. On AutomationBench, Zapier’s test of end-to-end business workflows, it scores 51.3% against 42.5% for Opus 5.5, and it leads the Vals Index, which weights finance, coding, legal and tax work by their share of U.S. GDP. On DeepSWE v1.1, the long-horizon software-engineering test, it reaches 77.9%, ahead of Opus 5.5 (74.2%) and Astra (74.1%).

Google’s table: where Gemini 4 Argon leads, and where it doesn’t% · higher is better · rivals at their highest reported setting · Google, 30 Sept 2026Gemini 4 ArgonGPT-6 AstraClaude Opus 5.502040608068.963.167.0Vals Index51.341.442.5AutomationBench77.974.174.2DeepSWE v1.155.065.562.3FrontierSWE v257.458.266.4Terminal-Bench4.045.344.349.3PostTrainBench57.668.163.3Terminal-BenchScience69.272.6n/rOSWorld-2.0
Eight of the 19 rows in Google’s table, chosen to show its largest wins beside every test it loses, against the two strongest rivals. Opus 5.5 has no OSWorld-2.0 offline score. Claude Fable 5.1 is in the full table below. Data: Google.

Where Argon trails, the gaps are wider than its coding leads. Astra is 10.5 points ahead on FrontierSWE v2 and on Terminal-Bench Science, Opus 5.5 is 9 points ahead on Terminal-Bench 4.0 and 4 points ahead on PostTrainBench, and Astra leads OSWorld-2.0, a computer-use test, by 3.4. In the table’s own agentic-coding group Argon wins two tests, DeepSWE and Vibe Code Bench, and loses two, so it is not a clean sweep in coding.

Google’s methodology document is worth reading beside the table. Rival scores are mostly their makers’ own reported figures or public leaderboards, at the highest reasoning setting available. Several of Argon’s were run by Google itself: DeepSWE with a mini-SWE-agent harness, Terminal-Bench 4.0, Terminal-Bench Science with a verifier timeout six times the standard one, and OSWorld-2.0, where Google reports the best of three runs. Two newer rivals are missing: Claude Sonnet 5.5, released on 28 September, and GPT-6.1 Sol, released the day before Argon. OpenAI’s own DeepSWE 1.1 figure for GPT-6.1 Sol is 75.2%, measured in OpenAI’s set-up rather than Google’s.

Here is the full table as Google published it. The best score in each row is in bold.

AreaBenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
Knowledge workVals Index68.9%63.1%65.8%67.0%
Knowledge workAutomationBench51.3%41.4%31.4%42.5%
Knowledge workVals Finance Agent v265.4%53.5%58.9%58.6%
Knowledge workHarvey’s Legal Agent Benchmark19.6%5.4%6.7%3.8%
Agentic codingDeepSWE v1.177.9%74.1%67.4%74.2%
Agentic codingFrontierSWE v255.0%65.5%56.3%62.3%
Agentic codingVibe Code Bench91.9%89.6%90.3%90.3%
Agentic codingTerminal-Bench 4.057.4%58.2%57.9%66.4%
ML engineeringPostTrainBench45.3%44.3%40.2%49.3%
Science and mathsTerminal-Bench Science 0.157.6%68.1%52.6%63.3%
Science and mathsLABBench 288.8%85.4%68.6%73.1%
Science and mathsRiemannBench76.0%72.0%65.6%69.6%
Long contextGraphWalks, up to 128K (F1)99.7%98.7%91.4%90.6%
Long contextGraphWalks, 256K to 1M (F1)84.2%71.8%65.0%66.8%
Computer useAgent’s Last Exam (pass rate)39.5%34.2%—38.2%
Computer useOSWorld-2.0 (offline, partial score)69.2%72.6%——
MultimodalChartography71.6%71.0%46.2%66.3%
MultimodalLVBench91.7%87.5%79.7%83.7%
CybersecurityCWE-bench v168.0%68.0%58.0%67.0%

The independent number

Artificial Analysis, which runs its own ten-test Intelligence Index, has already measured Argon at its high setting, the only one it lists. Argon scores 53 on version 4.3.2 of the index. That is level with GPT-6 Astra and Claude Fable 5.1 at their maximum settings, three points behind Claude Sonnet 5.5 and five behind Claude Opus 5.5, which still leads at 58. For Google it is a large step: Gemini 3.8 Flash, its best-scoring model until now, sits at 41.

The independent check: level with Astra, for lessArtificial Analysis Intelligence Index v4.3.2 · best listed setting per model · higher is better · read 30 Sept 2026COST PER TASK0102030405060Claude Opus 5.5 (max)58$5.98Claude Sonnet 5.5 (max)56$7.62Claude Fable 5.1 (max)53$7.63GPT-6 Astra (max)53$3.26Gemini 4 Argon (high)53$1.99GPT-6.1 Sol (max)52$0.72Gemini 3.8 Flash (high)41$1.24
Cost per task is what one index task cost to run, at the prices Artificial Analysis used (Argon at the introductory $2/$10). Data: Artificial Analysis, read 30 September 2026.

Cost is where Argon stands out. Running the index cost $1.99 per task, against $3.26 for Astra and $7.63 for Fable 5.1, so it matches their scores for 61% and 26% of their cost. That is despite being, in Artificial Analysis’s words, “somewhat verbose”: it wrote 110 million output tokens to finish the index. Two caveats apply. The figure uses the introductory price, and at $4 and $20 the same run would cost about twice as much, or roughly $4 a task (our arithmetic, not Artificial Analysis’s). And two rivals are better value even now. GPT-6.1 Sol scores 52 for $0.72 a task, and Claude Opus 5.5 at its high setting scores 54 for $1.82, a point more for slightly less money. On our AI model leaderboard Argon enters in third place: it ties Astra and Fable 5.1 on score, and the board settles ties on cost per task.

Cyber defenders first, without guardrails

Cybersecurity is the one area where Google is releasing Argon early, and deliberately without its usual restrictions. Google says the model can find, validate and patch serious vulnerabilities on its own, and trusted defenders get a version with the cyber guardrails removed. Wiz is already using it in Scan for Good, its free programme for protecting public infrastructure, where the model found a critical flaw exposing personal data in healthcare software used by hospitals worldwide. Google says earlier frontier models had missed it. On CWE-bench v1, which tests fixing vulnerabilities, Argon ties Astra at 68%. The rest of the cyber evidence is internal: Google’s own vulnerability dataset and Wiz’s black-box penetration tests, where Google says Argon improves on Gemini 3.8 Flash Cyber.

For everyone else, Google lists four sets of safeguards to finish before a wider release. The model refuses help with cyber and chemical, biological, radiological or nuclear attacks under Google’s Frontier Safety Framework, and Google monitors its internal activations for misuse. Google calls Argon its most resilient model yet against indirect prompt injection, and says it leads Gray Swan’s benchmark for that attack. Monitors read its chain of thought and its actions and stop it when it goes beyond what the user asked. And its sandboxes are sealed before high-risk training or evaluation, the problem OpenAI ran into in July when a maths model kept escaping its sandbox. Google also asks the rest of the industry to keep models’ reasoning readable, so that monitors like these still work.

Already at work inside Google

Google says thousands of its employees already use Argon, and it gives three examples. The model sped up quantum-computing subroutines, beating a published baseline by 40% in minutes. A team of Argon agents read fleet-wide profiling data and freed more than 300 TiB of memory across Google’s data centres, with 500 TiB to 1 PiB expected in total. And Argon agents are moving C and C++ code to Rust, from libraries such as re2 up to the 800,000-line Fuchsia Zircon kernel. In libgav1, Google’s open-source video decoder, they replaced 32,000 lines of SIMD code in an existing Rust port. The result is memory-safe and runs 2.7 times faster than that port, with identical output.

What the leaks got wrong

The numbers that circulated in September were too high. A chart shared on X claimed 88% on DeepSWE and 86.8% on OSWorld-2.0. Google’s own figures are 77.9% and 69.2%, although the leaked chart gave no methodology, so its OSWorld figure may not measure the same subset. The name, at least, was right. What the announcement leaves out entirely is Gemini 3.5 Pro, the model Google promised for June and kept delaying. It is not mentioned once.

Google has not said when paid API customers get access, or how long the $2 and $10 price lasts once they do. Until then, Argon’s results come from Google and from Artificial Analysis, which evidently had early access. Nobody outside the Fairwind programme has yet been able to test it on their own work.

In the Lab: as soon as Gemini 4 Argon is released beyond Google’s testers, we will put it through the BitsMinds Lab, on the same build briefs we give every model, one attempt each, with every build playable side by side.

More on Gemini

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Two toy buggies climbing one hill for the BitsMinds Lab Hill Climb test An editorial illustration: one grassy hill at golden hour with a pennant on its summit and a large pale sun behind it. Two toy hill-climbing buggies drive up from either side at the same height, a clay-red one with the Claude mark inlaid in its side panel and a graphite-and-gold one with the OpenAI mark. Gold coins float ahead of each. The names Opus 5.5 and GPT-6.1 Sol are painted on the soil below, with VS between them. Opus 5.5 VS GPT-6.1 Sol BITSMINDS LAB · HILL CLIMB BITSMINDS.COM
Models

Claude Opus 5.5 vs GPT-6.1 Sol: A Familiar Hill Climb

GPT-6.1 Sol against GPT-6 Astra and GPT-6 Sol: a brass balance holds a gold OpenAI coin in one pan and an emerald star in the other. BITSMINDS LAB GPT-6.1 Sol GPT-6 Astra BITSMINDS.COM
Models

GPT-6.1 Sol vs GPT-6 Astra and GPT-6 Sol: A Dead Heat

GPT-6.1 Astra: deception detected An original Decepticon-inspired robotic mask is forged from sharply faceted gunmetal and violet armour. Narrow violet eyes glow beneath angular brows, a pointed jaw ends in a blade-like chin, and the official OpenAI knot is inset into its forehead. The caption reads GPT-6.1 Astra, deception detected, release cancelled. The fictional robot is an editorial metaphor requested for the article; it does not depict a real OpenAI product, a conscious model, or a numerical result from the separate Astra simulation study. OPENAI GPT-6.1 Astra DECEPTION DETECTED RELEASE CANCELLED BITSMINDS.COM
Models

OpenAI Scraps GPT-6.1 Astra Over Deception in Tests