All News
gemini-4-argonbenchmarkshallucinationsartificial-analysisai-news

Gemini 4 Argon's 15% hallucination rate is real. 'Solved' is not.

Artificial Analysis measured Gemini 4 Argon's hallucination rate at 15%, the lowest among top-tier models. But the same run shows its accuracy fell to 50%.

Vlad MakarovVlad Makarovreviewed and published
3 min read
Gemini 4 Argon's 15% hallucination rate is real. 'Solved' is not.

On September 30, a post on r/singularity titled "Gemini 4 Argon solved hallucinations." climbed to roughly a thousand upvotes and a couple hundred comments. Reddit scores are an interest indicator and nothing more, but the framing traveled. The number underneath it is real, and it was measured by someone other than Google. The word "solved" is the story.

What the benchmark actually measured

AA-Omniscience is a knowledge-and-hallucination benchmark built by Artificial Analysis: 6,000 questions across six domains and 42 economically relevant topics, drawn from academic and industry sources. The firm states that all evaluations are conducted independently by it, so this is a third-party measurement rather than a vendor figure, and there is a public dataset and an arXiv paper behind it.

The Hallucination Rate is narrower than it sounds. It divides incorrect answers by incorrect plus partial answers plus not attempted. Lower is better, and what it tracks is how often a model guesses wrongly when it should have abstained — not how often it is wrong overall.

Argon at high reasoning recorded 15%. Artificial Analysis called that the lowest of any model scoring 45 or above on its Intelligence Index, against 51% for GPT-6 Astra (max) and 54% for GPT-6.1 Sol (max).

The trade the headline dropped

Here is the part that fits badly in a one-line post. The same run puts Argon's AA-Omniscience accuracy at 50%, a five-point drop from Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra at max effort (63%):

  • Hallucination rate: 15%
  • AA-Omniscience accuracy: 50%
  • AA-Omniscience Index: 42, level with GPT-6.1 Sol (42) and just under GPT-6 Astra (43)

A lower hallucination rate sitting next to lower accuracy is the signature of a model that abstains more, not one that knows more — and the benchmark rewards abstention by construction. Argon is also not the least-hallucinating model anywhere; small models beat it outright, including MiniCPM5-1B at 1% and G9v3-3B at 12%. Artificial Analysis's own gloss is careful: Argon "is much more likely to acknowledge when it does not know an answer rather than guess incorrectly."

What would settle it

Independent replication on other knowledge benchmarks, and real-world error rates from deployments rather than a single run. On the broader Intelligence Index the same day, Argon tied GPT-6 Astra at 53, which is a genuine result — and Google is not releasing it publicly yet, holding it behind the Fairwind Program for cyber defenders first.

We covered the same measurement gap when GPT-6.1 Sol launched, and public sentiment has been skeptical for a while. Until replication lands, the honest read is the boring one: more abstention, less accuracy, one number in a headline.

Related Articles

Scroll down

to load the next article