The first people outside Google allowed to use Gemini 4 Argon are not developers or paying subscribers. They are cybersecurity teams, and Google is handing them the model without its cyber guardrails.
Google DeepMind announced Argon on 30 September 2026, calling it "our next era of frontier intelligence". It is the first new generation of Google's flagship model since Gemini 3 in November 2025, it can write up to 1 million tokens in a single answer, and it launches at an introductory USD 2 per million input tokens and USD 10 per million output tokens. By one independent measure it puts Google level with OpenAI's best. What it does not have is a public release date.
- What: Gemini 4 Argon, Google's new flagship model for long coding, legal, finance and security work, with a 1 million token output limit.
- Price: USD 2 per million input tokens and USD 10 per million output tokens at launch, rising to USD 4 and USD 20 after an introductory period.
- Why it matters: Google is back among the top three labs on an independent index, but access starts with vetted cyber defenders while Google takes part in a voluntary US government pre-release process.
What Google built
Argon is "built to sustain deep reasoning across complex, long-horizon workflows", Koray Kavukcuoglu, who runs Google DeepMind day to day, wrote in the announcement. The most concrete change is length: the output limit rises to 1 million tokens from 64,000, so the model can reason through hundreds of thousands of tokens in one pass instead of stopping partway through a long task.
Thousands of Google employees already use the model, the company said. A team of Argon agents studied profiling data from Google's server fleet and applied memory optimisations that free "over 300 TiB of memory" once rolled out. Other Argon agents are migrating C and C++ code to Rust, from core libraries up to more than 800,000 lines of the Fuchsia Zircon kernel, with human and automated review before release.
The benchmark claims are Google's own. It reports 77.9 percent on DeepSWE v1.1, a test of long software engineering tasks, 91.7 percent on LVBench for long video understanding, first place on Zapier's AutomationBench at 51.3 percent and a shared first place at 68 percent on CWE-bench v1, which measures how well a model repairs security flaws. In Google's comparisons, VentureBeat reported, Claude Opus 5.5 scored 74.2 percent and GPT-6 Astra 74.1 percent on DeepSWE, while GPT-6 Astra beat Argon on FrontierSWE v2 by 65.5 percent to 55.0 percent. These are vendor reported figures.
Who gets it, and at what price
Argon is "rolling out to a set of trusted cyber defenders" through Google's Fairwind Program, and Google said it is "actively engaged in the U.S. government's voluntary process for pre-release model access". The defenders and Google's internal teams get Argon "without cyber guardrails". Google said the security company Wiz used it to uncover a critical flaw in healthcare software used by hospitals worldwide, one that earlier frontier models had missed.
The wider release to developers, enterprises and consumers will start with paid API customers and Google AI Ultra subscribers, and Google promised it "as soon as possible", without a date. The introductory price is USD 2 per million input tokens and USD 10 per million output tokens, with cached input 95 percent cheaper than fresh input. After the introductory period it doubles to USD 4 and USD 20. Artificial Analysis, an independent benchmarking firm, said the discount runs for at least a month and that Google had not confirmed when it ends.
What outside testers and Google staff say
Artificial Analysis scored Argon at 53 on its Intelligence Index, a composite of its own tests: level with GPT-6 Astra, one point above GPT-6.1 Sol and 23 points above Gemini 3.1 Pro Preview, Google's previous model above the Flash class. At the discounted price a task cost USD 1.99, about 60 percent of GPT-6 Astra. The saving comes from cheaper tokens, not fewer of them: Argon used about 62,000 output tokens per task against 27,000 for GPT-6 Astra, and at standard prices the cost rises to USD 3.98. Its hallucination rate of 15 percent was the lowest of any model scoring 45 or more on the index, which means it more often says it does not know instead of guessing.
Inside Google the verdict is less tidy. Some employees say the model does well on benchmarks but less well when they put it to work, and that it struggles with some coding tasks, Bloomberg reported, as summarised by The Next Web. Google told Bloomberg it would be inaccurate to say Gemini 4 underperforms in areas such as coding.
The road from I/O
Argon ends an awkward stretch. At Google I/O on 19 May, Google released Gemini 3.5 Flash, made it the default model in the Gemini app and in AI Mode in Search, and promised Gemini 3.5 Pro "next month". On the same stage Sundar Pichai said the Gemini app had passed 900 million monthly active users. Gemini 3.5 Pro never shipped. Google kept releasing Flash models instead: 3.6 Flash on 21 July, 3.7 Flash on 13 August and 3.8 Flash on 2 September, according to the Gemini API changelog.
On 5 August Google reorganised Google DeepMind: Demis Hassabis moved to chairman and stepped back from daily management, Kavukcuoglu took over operations reporting to Pichai, and longtime chief scientist Jeff Dean was reported to be leaving, Fortune wrote, adding that 3.5 Pro had missed a June target and then a mid July date. Engadget read Argon's arrival as a sign that Google "clearly chose to focus on developing Gemini 4 instead".
Google has tied the wider release to feedback from early testers and to further work on guardrails. Until paid API customers and AI Ultra subscribers get in, Argon's lead is something most of Google's customers can read about but not yet try.





Comments
Comments are reviewed before they appear.
No comments yet. Be the first to share your thoughts.