◆ NOISE IN → SIGNAL OUT◆ READALCHEMIST.COM◆ FREE / NO PAYWALL◆ NOISE IN → SIGNAL OUT◆ READALCHEMIST.COM◆ FREE / NO PAYWALL
THE DIGITAL ALCHEMIST
AIIMPACT 90

53 to 53. One Of Them Is Wrong Half The Time.

Gemini 4 Argon just tied GPT-6 Astra on the benchmark labs put in their decks. It hallucinates 15 percent of the time. Astra hallucinates 51 percent. The White House responded by renaming the category.

2026-10-015 MIN READ#Gemini 4 Argon · #GPT-6 Astra · #Artificial Analysis · #hallucination rate · #AI policy · #Google DeepMind · #OpenAI
white house press briefing room podium by DC_Rebecca (BY) via Openverse
white house press briefing room podium by DC_Rebecca (BY) via Openverse

Two models walked into a leaderboard. They tied. That is the whole story if you only read the press release.

With high reasoning, Gemini 4 Argon scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra at 53 and edging GPT-6.1 Sol at 52. Google is already calling this a comeback. It is not wrong, exactly. It is just not the number that matters.

The number that matters is this: on AA-Omniscience, a benchmark that tests factual knowledge and honest handling of knowledge gaps, Argon's hallucination rate is 15 percent, while GPT-6 Astra comes in at 51 percent and GPT-6.1 Sol at 54 percent. Same intelligence score. One model is wrong, confidently, roughly every other time it opens its mouth.

Hallucination rate on AA-Omniscience
20%40%60%15%Gemini 4 Argon51%GPT-6 Astra54%GPT-6.1 Sol
Source: Artificial Analysis independent evaluation of Gemini 4 Argon, GPT-6 Astra, GPT-6.1 Sol

The tie is real. The equivalence is not.

Argon is far more likely to admit it doesn't know an answer rather than guess wrong, but its accuracy reaches only 50 percent, 13 points below Astra's 63 percent. Astra guesses more and gets more right when it guesses. Different failure modes, not one model being simply better. A customer-support bot that says "I don't know" constantly is annoying. A customer-support bot that's confidently wrong half the time is a lawsuit.

That asymmetry is exactly why the Intelligence Index tie is theater. It collapses two completely different risk profiles into one number that goes in the sales deck.

At current discounted pricing, Gemini 4 Argon costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra's $3.26. Cheaper and more honest about its own ignorance. Also less accurate when it does commit to an answer. If your procurement team is buying off the top-line score alone, they're buying blind on the one axis that determines whether your product embarrasses you in public.

The Digital Alchemist
The Digital Alchemist

Washington's answer: change the label

While labs argued over decimal points, the White House solved a different problem entirely. President Donald Trump ordered all executive branch departments and agencies to use the term "super intelligence" instead of artificial intelligence.

Translation: stop measuring how often we're wrong, start admiring how smart we sound.

The timing is not subtle. The rebrand comes weeks before the midterm election, which appears increasingly shaped by Americans' concerns about rapid AI advancement. You don't rename a category unless the old name has become a liability. "Artificial" implies fake. "Super intelligence" implies trustworthy. Neither word has anything to do with whether the model answering your tax question is right.

Nobody in this story fixed the 51 percent hallucination rate. Google shipped a model that abstains more. OpenAI shipped Astra months ago and hasn't addressed the gap publicly. The administration renamed the discomfort. Three different responses to the same underlying fact, and only one of them is actually a fix.

If you're the enterprise buyer choosing between Argon and Astra for anything that talks to a customer, the leaderboard tie tells you nothing. The hallucination rate and the accuracy number sitting next to it tell you what you're actually signing up for.

What to watch: whether Artificial Analysis treats hallucination rate as a co-equal metric to the Intelligence Index instead of a footnote. Whether OpenAI responds to a now-public 51 percent figure with anything beyond a blog post. And whether any federal agency actually adopts "SI" in a document that matters, versus quietly ignoring an order with no enforcement teeth.

The tie vs. the tradeoff
53IntelligenceIndex (both)1.99Argon cost pertask3.26Astra cost pertask50Argon accuracyvs Astra's 63%
Source: Artificial Analysis
Sources
  1. Artificial Analysis: Gemini 4 Argon, Google is back as one of the top three labs
  2. Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead
  3. Trump tells federal agencies to use the term 'Super Intelligence' instead of artificial intelligence
  4. Trump tries to rename AI 'super intelligence' as polls show him sinking on key issue
  5. Techmeme: Artificial Analysis says Gemini 4 Argon (high) matches GPT-6 Astra (max)
  6. Gemini 4 Argon Ties GPT-6 Astra at 53 With 15% Hallucination Rate
  7. Techmeme on X: "Artificial Analysis says Gemini 4 Argon (high) matches GPT-6 Astra (max) on its Intelligence Index and has a 15% hallucination rate, compared with 51% for Astra (Artificial Analysis) (Visit Techmeme dot com for the link and full context!)" / X
  8. Gemini 4 Argon Matches GPT-6 Astra in Intelligence Index
  9. Gemini 4 Matches GPT-6 Astra but Trails Opus 5.5
  10. Google Seemingly Delivers On The Internal RSI Hype With The All-New Gemini 4 Argon That Offers GPT-6 Astra-Level Capabilities
  11. Can a Hallucinating Model help in Reducing Human "Hallucination"?
  12. Trump tells federal agencies to use the term 'Super Intelligence ' instead of artificial intelligence
  13. Trump Signs ‘Super Intelligence’ Order: What Actually Changes for AI
  14. Trump orders federal agencies to replace ‘AI’ with ‘Super Intelligence’
  15. Trump orders federal use of 'super intelligence' over 'AI'
  16. Superintelligence
  17. Trump Orders US Federal Government to Replace ‘AI’ With ‘Super Intelligence’ or ‘SI’
  18. scoop trump admin openai partner unleash artificial intelligence federal government
  19. Artificial Analysis on X: "Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices. Google is now back to being one of the top three labs in intelligence achieved Gemini 4 Argon is @GoogleDeepMind’s first propr… / X
  20. Evaluating ChatGPT and Google Gemini Performance and Implications in Turkish Dental Education
  21. Evaluating the Accuracy of Gemini 2.0 Advanced and ChatGPT 4o in Cataract Knowledge: A Performance Analysis Using Brazilian Council of Ophthalmology Board Exam Questions
  22. Comparative Evaluation of AI Models Such as ChatGPT 3.5, ChatGPT 4.0, and Google Gemini in Neuroradiology Diagnostics
← back to the feed
NVDA 230.86 ▲ 1.09%AAPL 330.32 ▼ 0.81%MSFT 512.80 ▼ 0.02%GOOGL 338.24 ▼ 1.70%AMZN 248.23 ▼ 0.37%META 725.93 ▲ 0.10%TSLA 354.11 ▼ 0.20%AMD 615.73 ▲ 0.65%AVGO 343.64 ▼ 2.15%PLTR 190.04 ▲ 1.60%COIN 189.29 ▲ 1.54%MSTR 160.50 ▲ 4.84%NVDA 230.86 ▲ 1.09%AAPL 330.32 ▼ 0.81%MSFT 512.80 ▼ 0.02%GOOGL 338.24 ▼ 1.70%AMZN 248.23 ▼ 0.37%META 725.93 ▲ 0.10%TSLA 354.11 ▼ 0.20%AMD 615.73 ▲ 0.65%AVGO 343.64 ▼ 2.15%PLTR 190.04 ▲ 1.60%COIN 189.29 ▲ 1.54%MSTR 160.50 ▲ 4.84%