Gemini 4 Argon: Benchmarks, Price and Who Can Use It
Google's Gemini 4 Argon tops 13 of the 19 tests in Google's own table, but so far only a group of cyber defenders can use it. Paid API customers and Google AI Ultra subscribers come next, at an introductory $2 per million input tokens and $10 per million output tokens. Artificial Analysis already puts it level with GPT-6 Astra at 53 points, for 61% of Astra's cost per task.
Gemini 4 Argon is Google’s new frontier model, announced on 30 September 2026 by Koray Kavukcuoglu, who runs Google DeepMind day to day. It is the first model of the Gemini 4 generation, and Google pitches it at long, multi-step work in software engineering, finance, law and cyber defence. You cannot use it yet. Google’s announcement says it is “rolling out to a set of trusted cyber defenders through our Fairwind Program”, with paid API customers and Google AI Ultra subscribers next and no date for either. Argon is the name that circulated in September leaks, which we treated as unverified in our Gemini 4 explainer. It turned out to be the product name.
Price, limits and a slow rollout
Argon will launch at an introductory $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper, which works out to $0.10 per million. After the introductory period the price doubles to $4 and $20, and Google has not said how long that period lasts. The launch sticker matches GPT-6.1 Sol, which OpenAI released the day before at the same $2 and $10.
The biggest technical change is the output limit. Argon can write up to 1M tokens in a single response, up from 64K before, which Google says gives it room to think through hard problems in one pass. Google’s post does not state a context window. Artificial Analysis lists it at 1M tokens.
The rollout is unusually cautious. Google says it is taking part in the U.S. government’s voluntary process for pre-release model access, and that it will keep collecting feedback from early testers “as we iterate on guardrails” before opening Argon to developers, businesses and consumers. It comes days after OpenAI scrapped GPT-6.1 Astra for failing its own safety tests.
Google’s benchmark table
Google compared Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 19 tests. Argon has the top score on 13, ties Astra on one and trails on five. The clearest wins are in knowledge work. On Harvey’s legal agent benchmark Argon scores 19.6%, against 6.7% for the next-best model. On AutomationBench, Zapier’s test of end-to-end business workflows, it scores 51.3% against 42.5% for Opus 5.5, and it leads the Vals Index, which weights finance, coding, legal and tax work by their share of U.S. GDP. On DeepSWE v1.1, the long-horizon software-engineering test, it reaches 77.9%, ahead of Opus 5.5 (74.2%) and Astra (74.1%).
Where Argon trails, the gaps are wider than its coding leads. Astra is 10.5 points ahead on FrontierSWE v2 and on Terminal-Bench Science, Opus 5.5 is 9 points ahead on Terminal-Bench 4.0 and 4 points ahead on PostTrainBench, and Astra leads OSWorld-2.0, a computer-use test, by 3.4. In the table’s own agentic-coding group Argon wins two tests, DeepSWE and Vibe Code Bench, and loses two, so it is not a clean sweep in coding.
Google’s methodology document is worth reading beside the table. Rival scores are mostly their makers’ own reported figures or public leaderboards, at the highest reasoning setting available. Several of Argon’s were run by Google itself: DeepSWE with a mini-SWE-agent harness, Terminal-Bench 4.0, Terminal-Bench Science with a verifier timeout six times the standard one, and OSWorld-2.0, where Google reports the best of three runs. Two newer rivals are missing: Claude Sonnet 5.5, released on 28 September, and GPT-6.1 Sol, released the day before Argon. OpenAI’s own DeepSWE 1.1 figure for GPT-6.1 Sol is 75.2%, measured in OpenAI’s set-up rather than Google’s.
Here is the full table as Google published it. The best score in each row is in bold.
| Area | Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|---|
| Knowledge work | Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| Knowledge work | AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% |
| Knowledge work | Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% |
| Knowledge work | Harvey’s Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% |
| Agentic coding | DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| Agentic coding | FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Agentic coding | Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% |
| Agentic coding | Terminal-Bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| ML engineering | PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% |
| Science and maths | Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% |
| Science and maths | LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% |
| Science and maths | RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% |
| Long context | GraphWalks, up to 128K (F1) | 99.7% | 98.7% | 91.4% | 90.6% |
| Long context | GraphWalks, 256K to 1M (F1) | 84.2% | 71.8% | 65.0% | 66.8% |
| Computer use | Agent’s Last Exam (pass rate) | 39.5% | 34.2% | — | 38.2% |
| Computer use | OSWorld-2.0 (offline, partial score) | 69.2% | 72.6% | — | — |
| Multimodal | Chartography | 71.6% | 71.0% | 46.2% | 66.3% |
| Multimodal | LVBench | 91.7% | 87.5% | 79.7% | 83.7% |
| Cybersecurity | CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
The independent number
Artificial Analysis, which runs its own ten-test Intelligence Index, has already measured Argon at its high setting, the only one it lists. Argon scores 53 on version 4.3.2 of the index. That is level with GPT-6 Astra and Claude Fable 5.1 at their maximum settings, three points behind Claude Sonnet 5.5 and five behind Claude Opus 5.5, which still leads at 58. For Google it is a large step: Gemini 3.8 Flash, its best-scoring model until now, sits at 41.
Cost is where Argon stands out. Running the index cost $1.99 per task, against $3.26 for Astra and $7.63 for Fable 5.1, so it matches their scores for 61% and 26% of their cost. That is despite being, in Artificial Analysis’s words, “somewhat verbose”: it wrote 110 million output tokens to finish the index. Two caveats apply. The figure uses the introductory price, and at $4 and $20 the same run would cost about twice as much, or roughly $4 a task (our arithmetic, not Artificial Analysis’s). And two rivals are better value even now. GPT-6.1 Sol scores 52 for $0.72 a task, and Claude Opus 5.5 at its high setting scores 54 for $1.82, a point more for slightly less money. On our AI model leaderboard Argon enters in third place: it ties Astra and Fable 5.1 on score, and the board settles ties on cost per task.
Cyber defenders first, without guardrails
Cybersecurity is the one area where Google is releasing Argon early, and deliberately without its usual restrictions. Google says the model can find, validate and patch serious vulnerabilities on its own, and trusted defenders get a version with the cyber guardrails removed. Wiz is already using it in Scan for Good, its free programme for protecting public infrastructure, where the model found a critical flaw exposing personal data in healthcare software used by hospitals worldwide. Google says earlier frontier models had missed it. On CWE-bench v1, which tests fixing vulnerabilities, Argon ties Astra at 68%. The rest of the cyber evidence is internal: Google’s own vulnerability dataset and Wiz’s black-box penetration tests, where Google says Argon improves on Gemini 3.8 Flash Cyber.
For everyone else, Google lists four sets of safeguards to finish before a wider release. The model refuses help with cyber and chemical, biological, radiological or nuclear attacks under Google’s Frontier Safety Framework, and Google monitors its internal activations for misuse. Google calls Argon its most resilient model yet against indirect prompt injection, and says it leads Gray Swan’s benchmark for that attack. Monitors read its chain of thought and its actions and stop it when it goes beyond what the user asked. And its sandboxes are sealed before high-risk training or evaluation, the problem OpenAI ran into in July when a maths model kept escaping its sandbox. Google also asks the rest of the industry to keep models’ reasoning readable, so that monitors like these still work.
Already at work inside Google
Google says thousands of its employees already use Argon, and it gives three examples. The model sped up quantum-computing subroutines, beating a published baseline by 40% in minutes. A team of Argon agents read fleet-wide profiling data and freed more than 300 TiB of memory across Google’s data centres, with 500 TiB to 1 PiB expected in total. And Argon agents are moving C and C++ code to Rust, from libraries such as re2 up to the 800,000-line Fuchsia Zircon kernel. In libgav1, Google’s open-source video decoder, they replaced 32,000 lines of SIMD code in an existing Rust port. The result is memory-safe and runs 2.7 times faster than that port, with identical output.
What the leaks got wrong
The numbers that circulated in September were too high. A chart shared on X claimed 88% on DeepSWE and 86.8% on OSWorld-2.0. Google’s own figures are 77.9% and 69.2%, although the leaked chart gave no methodology, so its OSWorld figure may not measure the same subset. The name, at least, was right. What the announcement leaves out entirely is Gemini 3.5 Pro, the model Google promised for June and kept delaying. It is not mentioned once.
Google has not said when paid API customers get access, or how long the $2 and $10 price lasts once they do. Until then, Argon’s results come from Google and from Artificial Analysis, which evidently had early access. Nobody outside the Fairwind programme has yet been able to test it on their own work.
In the Lab: as soon as Gemini 4 Argon is released beyond Google’s testers, we will put it through the BitsMinds Lab, on the same build briefs we give every model, one attempt each, with every build playable side by side.
More on Gemini
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.