53 to 53. One Of Them Is Wrong Half The Time.
Gemini 4 Argon just tied GPT-6 Astra on the benchmark labs put in their decks. It hallucinates 15 percent of the time. Astra hallucinates 51 percent. The White House responded by renaming the category.

Two models walked into a leaderboard. They tied. That is the whole story if you only read the press release.
With high reasoning, Gemini 4 Argon scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra at 53 and edging GPT-6.1 Sol at 52. Google is already calling this a comeback. It is not wrong, exactly. It is just not the number that matters.
The number that matters is this: on AA-Omniscience, a benchmark that tests factual knowledge and honest handling of knowledge gaps, Argon's hallucination rate is 15 percent, while GPT-6 Astra comes in at 51 percent and GPT-6.1 Sol at 54 percent. Same intelligence score. One model is wrong, confidently, roughly every other time it opens its mouth.
The tie is real. The equivalence is not.
Argon is far more likely to admit it doesn't know an answer rather than guess wrong, but its accuracy reaches only 50 percent, 13 points below Astra's 63 percent. Astra guesses more and gets more right when it guesses. Different failure modes, not one model being simply better. A customer-support bot that says "I don't know" constantly is annoying. A customer-support bot that's confidently wrong half the time is a lawsuit.
That asymmetry is exactly why the Intelligence Index tie is theater. It collapses two completely different risk profiles into one number that goes in the sales deck.
At current discounted pricing, Gemini 4 Argon costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra's $3.26. Cheaper and more honest about its own ignorance. Also less accurate when it does commit to an answer. If your procurement team is buying off the top-line score alone, they're buying blind on the one axis that determines whether your product embarrasses you in public.

Washington's answer: change the label
While labs argued over decimal points, the White House solved a different problem entirely. President Donald Trump ordered all executive branch departments and agencies to use the term "super intelligence" instead of artificial intelligence.
Translation: stop measuring how often we're wrong, start admiring how smart we sound.
The timing is not subtle. The rebrand comes weeks before the midterm election, which appears increasingly shaped by Americans' concerns about rapid AI advancement. You don't rename a category unless the old name has become a liability. "Artificial" implies fake. "Super intelligence" implies trustworthy. Neither word has anything to do with whether the model answering your tax question is right.
Nobody in this story fixed the 51 percent hallucination rate. Google shipped a model that abstains more. OpenAI shipped Astra months ago and hasn't addressed the gap publicly. The administration renamed the discomfort. Three different responses to the same underlying fact, and only one of them is actually a fix.
If you're the enterprise buyer choosing between Argon and Astra for anything that talks to a customer, the leaderboard tie tells you nothing. The hallucination rate and the accuracy number sitting next to it tell you what you're actually signing up for.
What to watch: whether Artificial Analysis treats hallucination rate as a co-equal metric to the Intelligence Index instead of a footnote. Whether OpenAI responds to a now-public 51 percent figure with anything beyond a blog post. And whether any federal agency actually adopts "SI" in a document that matters, versus quietly ignoring an order with no enforcement teeth.
- Artificial Analysis: Gemini 4 Argon, Google is back as one of the top three labs
- Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead
- Trump tells federal agencies to use the term 'Super Intelligence' instead of artificial intelligence
- Trump tries to rename AI 'super intelligence' as polls show him sinking on key issue
- Techmeme: Artificial Analysis says Gemini 4 Argon (high) matches GPT-6 Astra (max)
- Gemini 4 Argon Ties GPT-6 Astra at 53 With 15% Hallucination Rate
- Techmeme on X: "Artificial Analysis says Gemini 4 Argon (high) matches GPT-6 Astra (max) on its Intelligence Index and has a 15% hallucination rate, compared with 51% for Astra (Artificial Analysis) (Visit Techmeme dot com for the link and full context!)" / X
- Gemini 4 Argon Matches GPT-6 Astra in Intelligence Index
- Gemini 4 Matches GPT-6 Astra but Trails Opus 5.5
- Google Seemingly Delivers On The Internal RSI Hype With The All-New Gemini 4 Argon That Offers GPT-6 Astra-Level Capabilities
- Can a Hallucinating Model help in Reducing Human "Hallucination"?
- Trump tells federal agencies to use the term 'Super Intelligence ' instead of artificial intelligence
- Trump Signs ‘Super Intelligence’ Order: What Actually Changes for AI
- Trump orders federal agencies to replace ‘AI’ with ‘Super Intelligence’
- Trump orders federal use of 'super intelligence' over 'AI'
- Superintelligence
- Trump Orders US Federal Government to Replace ‘AI’ With ‘Super Intelligence’ or ‘SI’
- scoop trump admin openai partner unleash artificial intelligence federal government
- Artificial Analysis on X: "Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices. Google is now back to being one of the top three labs in intelligence achieved Gemini 4 Argon is @GoogleDeepMind’s first propr… / X
- Evaluating ChatGPT and Google Gemini Performance and Implications in Turkish Dental Education
- Evaluating the Accuracy of Gemini 2.0 Advanced and ChatGPT 4o in Cataract Knowledge: A Performance Analysis Using Brazilian Council of Ophthalmology Board Exam Questions
- Comparative Evaluation of AI Models Such as ChatGPT 3.5, ChatGPT 4.0, and Google Gemini in Neuroradiology Diagnostics