Gemini 4 Argon Hallucinates 15%. Its Accuracy Drops 13 Points.
Google's new flagship ties GPT-6 Astra on the industry's IQ test by refusing to guess instead of lying. That restraint has a price, and so does the model.

Ask Gemini 4 Argon something it doesn't know, and it tells you. Ask GPT-6 Astra the same question, and more than half the time you get a confident answer that isn't true.
That's the real story hiding under Tuesday's benchmark tie. Artificial Analysis ran its Intelligence Index on Google's new flagship and came back with a dead heat: Gemini 4 Argon scores 53, matching GPT-6 Astra. Every outlet is running that as a comeback headline. The number that matters is underneath it: on AA-Omniscience, Argon has a 15% hallucination rate, the lowest of any model in its class, compared with 51% for GPT-6 Astra and 54% for GPT-6.1 Sol.
Google didn't build a smarter model. It built a model that knows what it doesn't know.
Here's what nobody wants to sit with: Gemini 4 Argon scores 50% on accuracy, down 5 points from Gemini 3.1 Pro Preview and 13 points below Astra's 63%. That gap isn't rounding error. It's one in three hard questions where Astra lands the right answer and Argon simply declines to guess. Refusing to lie is not free. You're trading wrong-and-confident for right-less-often, and a leaderboard tie can't tell you which failure costs more when a contract closes at 2am.
The sale price is not the business model
At launch, Gemini 4 Argon costs $1.99 per Intelligence Index task, 60% of Astra's $3.26. That discount is a promotion with no expiration date. On standard pricing, Argon's cost per task rises to $3.98. Once the promo lapses, Argon costs about 20% more than Astra, not 40% less. The "60 percent of the cost" framing is a snapshot of a sale, not a pricing strategy Google committed to.

You can't use it yet, and Google explained why
Argon isn't rolling out to developers or paying customers first. Google is giving access to selected cybersecurity defenders through its Fairwind Program, with broader rollout only after defenders have patched system vulnerabilities the model identifies.
Translation: the people who find security holes get first crack at this model, before anyone with a credit card.
That's defensible. It also means the honesty-over-accuracy bet hasn't faced a real commercial customer base yet. "Selected users only" is a safety gate, not a verdict on demand.
For workflows where a wrong answer costs real money—legal review, financial modeling, security scans—a model that says "I don't know" is worth more than one that fills the gap with something plausible. For everything else, Astra's extra 13 accuracy points might just win on output alone.
What to watch
Whether Argon reaches Google AI Ultra subscribers and paid API customers on what timeline. Whether the $3.98 post-promo price holds against Astra once Google stops subsidizing it. Whether OpenAI trains Astra to refuse more often, dropping its own accuracy in the process. Whether Fairwind's cyber defenders surface something in the coming weeks that explains why Google wanted trusted eyes on this model first.
- Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved
- Techmeme: Artificial Analysis says Gemini 4 Argon matches GPT-6 Astra
- Why has Google not released Gemini 4 Argon yet?
- Google Launches Gemini 4 Argon, Returning to the AI Forefront with Restricted Access
- Gemini 4 Argon: our next era of frontier intelligence
- Techmeme: Artificial Analysis says Gemini 4 Argon (high) matches GPT-6 Astra (max) on its Intelligence Index and has a 15% hallucination rate, compared with 51% for Astra (Artificial Analysis)
- Techmeme on X: "Artificial Analysis says Gemini 4 Argon (high) matches GPT-6 Astra (max) on its Intelligence Index and has a 15% hallucination rate, compared with 51% for Astra (Artificial Analysis) (Visit Techmeme dot com for the link and full context!)" / X
- Gemini 4 Argon Matches GPT-6 Astra in Intelligence Index
- Gemini 4 Argon Ties GPT-6 Astra at 53 With 15% Hallucination Rate
- Artificial Analysis on X: "Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices. Google is now back to being one of the top three labs in intelligence achieved Gemini 4 Argon is @GoogleDeepMind’s first propr… / X
- Google Seemingly Delivers On The Internal RSI Hype With The All-New Gemini 4 Argon That Offers GPT-6 Astra-Level Capabilities
- Evaluating the Accuracy of Chatbots in Financial Literature
- AA-Omniscience Hallucination Rate Leaderboard & Scores — September 2026
- SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI
- Gemini 4 Matches GPT-6 Astra but Trails Opus 5.5
- Gemini 4 Argon vs GPT-6 Astra: Level on Index, 1.6x Cost
- Google’s Gemini 4 Argon Matches GPT-6 Astra in Tests at 40% Lower Launch Cost
- Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense - MarkTechPost
- Gemini 4 Argon Review: Benchmarks, Price and API Access · OmniaKey
- Gemini 4 Argon: Features, Pricing, Access & Alternatives
- Google announces Gemini 4 Argon as its new frontier model
- Gemini Enterprise Agent Platform
- Gemini 4 Argon: our next era of frontier intelligence
- Google AI Studio
- Gemini 4 Argon: Price, Access and Benchmarks