Official Google key art for Gemini 4 Argon, the new flagship AI model

Google unveils Gemini 4 Argon, its first flagship AI model in almost a year

Google has unveiled Gemini 4 Argon, its first new flagship AI model in almost a year, ending months of delays that left it trailing OpenAI and Anthropic. But almost nobody can use it. The model is going first to a small group of trusted cybersecurity partners, and Google has given no date for a public release.

The launch post comes from Koray Kavukcuoglu, the executive who now runs Google DeepMind day to day, and it pitches Argon as a model already changing how Google builds software internally.

The headline change is the size of its answers. Argon can generate up to one million output tokens in a single run, up from 64,000 on earlier Gemini models. Google’s argument is that a model with room to reason across hundreds of thousands of tokens can solve a hard problem in one attempt instead of in fragments, which matters for long software engineering tasks and for working through long videos or large document collections.

Defenders get it first

Cybersecurity is the center of gravity for this launch. Argon rolls out first through Google’s Fairwind Program, which gives vetted cyber defenders early access to its most capable security models, and Google says it trained this one specifically for defensive security work.

It claims Argon can autonomously find, validate and patch critical software vulnerabilities. For trusted defenders and Google’s own teams, the company is releasing the model without its usual cyber guardrails so they can use its full defensive capability.

There is already a proof point. Cybersecurity firm Wiz is using Argon through its Scan for Good initiative, and Google says the model uncovered a critical vulnerability in healthcare software used by hospitals worldwide, one that previous frontier models had missed. Paid API customers and Google AI Ultra subscribers are next in line, but Google has not said when.

Google benchmark chart showing Gemini 4 Argon tying for first place on CWE-bench v1 at 68 percent
Source: Google

The numbers Google wants you to see

On the benchmarks Google chose to publish, Argon sets a new high of 77.9 percent on DeepSWE v1.1, a test of long-horizon software engineering, ahead of Claude Opus 5.5 at 74.2 percent and GPT-6 Astra at 74.1 percent.

It leads the Vals Index, which weights finance, coding, legal and tax work by their contribution to the economy, ranks first on Zapier’s AutomationBench at 51.3 percent, and reaches 91.7 percent on LVBench for long video understanding. On CWE-bench v1, which tests vulnerability remediation, it ties for first at 68 percent.

The fine print tells a more modest story. On two of the four coding benchmarks in Google’s own release materials, Argon trails its rivals. Google’s own description is carefully worded.

A company spokesperson called it the most performant model Google has built for complex workloads, comparable to OpenAI’s Astra and Anthropic’s Opus on key coding and cyber benchmarks. Comparable, not ahead, is doing a lot of work in that sentence.

Pricing is aggressive. Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, with cached inputs at 95 percent off, before stepping up to $4 and $20. The message underneath is that Google will compete on cost per unit of intelligence even where it cannot claim an outright capability lead.

The road to Argon

The launch follows a bruising stretch for Google’s AI division. Gemini 3.5 Pro, which CEO Sundar Pichai had said would arrive in June, will not ship at all. In the meantime the company overhauled its DeepMind lab, with founder Demis Hassabis stepping aside and several Gemini leaders leaving.

Google had not shipped a new flagship since the Gemini 3 series in November 2025, while OpenAI and Anthropic kept advancing with GPT-6 and new Claude models. Pichai pushed back in July on the idea that Google was losing ground, but the company has since shifted its messaging from cutting-edge capability to cost advantage.

Inside the company, the picture is less uniform. Some Googlers with direct access to Argon say it performs less well on real coding tasks than the benchmarks suggest, though others describe a large internal consensus that the model sits at the frontier. Google has disputed the skeptical read. The gap between benchmark marketing and day-to-day performance is the question every lab now faces, and Google has handed its critics the material to ask it.

What happens next

Google says it is taking part in the US government’s voluntary process for pre-release model access while it hardens safeguards around misuse, prompt injection and misalignment. Wider availability starts with paid API customers and AI Ultra subscribers, with developers, enterprises and consumers to follow. No dates have been announced.

The timing puts Google back in a race that has not waited for it. OpenAI launched Dots, always-on agents that keep working after you log off, and Manus launched Cue, personal agents with their own phone number and wallet. Google’s answer is a model built for the long, unglamorous work of engineering and defense. For now it is keeping that answer mostly to itself.

Nicole Catapano, a proficient news writer, covers AI, tech gadgets, and software products with over 6 years of experience. Her knack for simplifying complex tech topics is honed by her education in computer science.

Leave a Comment

Professor Derpy's Notes

I haven’t reviewed this story yet. Please check back later while I finish my highly scientific process of reading the headline three more times.

Join our newsletter

email subscription

Receive Latest AI Insights To Your Inbox