AI hallucination statistics 2026: rates across 108 models, and why the tests disagree

In brief

AI hallucination statistics for October 2026: on Vectara's leaderboard, models add unsupported claims to 1.8% to 24.2% of summaries, with a median of 9.6% across 108 models. On AA-Omniscience, frontier models give a wrong answer to 15% to 97% of the hard questions they miss. When Vectara made its test harder, 39 of 44 unchanged models scored worse.

The OpenAI, Claude and Gemini logos on plinths in front of curved chrome mirrors that warp their reflections

AI models add unsupported claims to 1.8% to 24.2% of summaries, and guess wrong on up to 96.5% of the hard questions they miss

We opened every AI hallucination statistic below at its primary source on 6 October 2026. The median and the matched-model count are our own arithmetic on Vectara's published tables.

StatisticFigureSourceDate
Lowest rate when summarising a document, 108 models1.8% (Ant Group Finix S1 32B)Vectara hallucination leaderboard22 Sep 2026
Median rate when summarising a document9.6%Whizi, from Vectara's table22 Sep 2026
Models adding unsupported claims to 10% or more of summaries50 of 108Whizi, from Vectara's table22 Sep 2026
Same models after Vectara made its test harder39 of 44 scored worse, median 4.1% to 7.95%Whizi, from Vectara's 2025 and 2026 tablesOct 2025 and Sep 2026
Lowest rate answering hard questions from memory, 30 frontier configurations15.1% (Gemini 4 Argon)Artificial Analysis AA-Omniscience6 Oct 2026
Highest rate on the same test96.5% (DeepSeek V4.1 Flash)Artificial Analysis AA-Omniscience6 Oct 2026
GPT-6 Astra on three different tests4.2%, 8.7% and 51.3%OpenAI, Vectara, Artificial AnalysisSep and Oct 2026
AI news answers with at least one significant issue45%EBU and BBC, 2,709 answers in 14 languagesOct 2025
Chatbot answers repeating a false news claimmore than 28%NewsGuard, 11 chatbotsJan 2026 audit
Court decisions involving AI-invented material2,149Damien Charlotin, AI Hallucination Cases database6 Oct 2026
Employees using AI at work who say it caused mistakes in their work56%KPMG and University of Melbourne, 48,000 peopleApr 2025
Legal research AI tools giving hallucinated answers17% to 33%Stanford RegLabMay 2024
Knowing more doesn't mean guessing less
02550751000204060accuracy on hard questions, %misses answered wrongly, %Gemini 4 ArgonQwen3.8 MaxGrok 4.7Claude Sonnet 5.5GPT-6 AstraClaude Opus 5.5Claude Fable 5.1GPT-6 LunaMistral Medium 3.5DeepSeek V4.1 Flash

The chart plots AA-Omniscience as Artificial Analysis showed it on 6 October 2026, at the highest reasoning setting it lists for each model. Across the bottom is accuracy on 6,000 hard questions. Up the side is the hallucination rate: of the questions a model didn't get right, the share it answered wrongly instead of saying it didn't know. Anthropic's Claude Fable 5.1 knows the most, at 67.2% correct. It still answers 72.6% of its misses with something false. Google's Gemini 4 Argon gets fewer right, 49.9%, and lands at 15.1% because it declines far more often. Knowing and guessing are separate skills.

What is the AI hallucination rate by model?

On Vectara's leaderboard, the AI hallucination rate by model runs from 1.8% for Ant Group's Finix S1 32B to 24.2% for Mistral's Ministral 3 3B, and the median model sits at 9.6%. Vectara gives every model the same set of more than 7,700 articles and counts how often its summary claims something that the article doesn't say. Not one of them scores zero.

ModelLabSummaries with an unsupported claim
GPT-5.4 nanoOpenAI3.1%
Gemini 2.5 Flash LiteGoogle3.3%
Llama 3.3 70BMeta4.1%
GPT-6 SolOpenAI6.5%
DeepSeek V4 ProDeepSeek8.6%
GPT-6 AstraOpenAI8.7%
Gemini 3.1 Pro (preview)Google10.4%
Claude Sonnet 4.6Anthropic10.6%
Kimi K2.6Moonshot AI10.8%
Claude Opus 4.7Anthropic12.0%
GPT-5.6 SolOpenAI12.4%
Grok 4.1 Fast (reasoning)xAI19.2%
Mistral Medium (2508)Mistral22.7%

Every row here is from Vectara's board of 22 September 2026, and the full list of all 108 models sits on Vectara's GitHub page. A lab's small model is often better at this than its flagship is: OpenAI's GPT-5.4 nano scores 3.1%, while GPT-5.6 Sol scores 12.4%. Newer isn't always lower, either. Anthropic's Claude Opus 4.7, at 12.0%, sits above the older Claude Sonnet 4.6. Vectara is open about what its test can't do. Its FAQ concedes that a model which copied sentences straight out of the article would get a perfect score.

The same model can carry three very different numbers. OpenAI's launch page for GPT-6 Astra, dated 3 September 2026, reports 4.2% on an internal benchmark. Vectara measures it at 8.7%, and Artificial Analysis puts it at 44.8% to 51.3%, depending on which reasoning setting it runs at. None of the three is wrong. Each of them asks a different question.

Why do AI hallucination benchmarks disagree?

AI hallucination benchmarks disagree because each one tests a different failure, on a different task, with a different denominator. Read the definition before you quote the number.

BenchmarkWhat the model is asked to doWhat the rate countsLatest range
Vectara leaderboardSummarise an article it's givenSummaries with a claim the article doesn't support1.8% to 24.2%, Sep 2026
Artificial Analysis AA-OmniscienceAnswer 6,000 hard questions from memoryMissed questions answered wrongly instead of declined15.1% to 96.5% across 30 frontier configurations, Oct 2026
OpenAI SimpleQA, in its September 2025 paperAnswer short fact questions from memoryAll questions answered wrongly26% (gpt-5-thinking-mini) to 75% (o4-mini), Sep 2025
OpenAI internal benchmarkAnswer ChatGPT conversations users flagged as wrongResponses with any factual error4.2% (GPT-6 Astra) and 12.2% (GPT-5.6 Sol), Sep 2026
NewsGuard AI False Claims MonitorRespond to prompts about false news storiesResponses that repeat the false claimmore than 28%, Jan 2026
EBU and BBC News Integrity studyAnswer news questionsAnswers with a significant issue, judged by journalists30% to 76% by assistant, Oct 2025

OpenAI's researchers argued in September 2025 that the way models are scored makes this worse. Their write-up says "standard training and evaluation procedures reward guessing over acknowledging uncertainty." The SimpleQA figures they used to show it are stark. OpenAI's older o4-mini got 24% of the questions right and 75% of them wrong, and it declined just 1%. The newer gpt-5-thinking-mini got 22% right, but it declined 52% and was wrong on only 26%. On a scoreboard that only counts what's right, the guesser wins.

That's why a single "hallucination rate" for a model doesn't exist. A reader who wants to see two labs disagree can send one question to two of Whizi's 280+ models in the side-by-side view, and why AI gets things wrong covers how to check the answer that follows.

Are AI hallucination rates going down over time?

On any one fixed test, AI hallucination rates have mostly fallen. But the tests keep getting harder, so a headline number can go up while the models get better. Vectara's leaderboard shows the second effect clearly.

Between October 2025 and September 2026, Vectara replaced its dataset with more than 7,700 articles, some as long as 24,000 words, and moved its judge from HHEM-2.1 to HHEM-2.3. We matched the 44 models listed under the same name on both boards, leaving out any name the old board carried in two versions. Thirty-nine scored worse on the new test. Four scored better, and DeepSeek V3.1 didn't move. The median for the 44 went from 4.1% to 7.95%, with every model the same release as before.

Same models, harder test: 39 of 44 scored worse
Oct 2025 testSep 2026 test051015% of summariesGPT-5 (high)1.4%15.1%10.8 times highergpt-oss-120b2.4%14.2%5.9 times higherClaude Sonnet 4.55.5%12.0%2.2 times higherGemini 2.5 FlashLite2.9%3.3%almost flatDeepSeek V3.15.5%5.5%unchangedAya Expanse 8B12.2%9.5%one of four that fell

OpenAI's GPT-5 at high reasoning effort made the biggest jump, from 1.4% to 15.1%. So a 2025 figure and a 2026 figure from the same leaderboard can't be read as a trend. Quote the date and the dataset with the number.

OpenAI's own figures mostly point down, with one reversal. Its April 2025 system card showed that o3 hallucinated on 33% of PersonQA questions, which is double the 16% of the older o1, and that o4-mini did so on 48%. In August 2025 OpenAI said GPT-5 with thinking was about 80% less likely than o3 to make a factual error. In September 2026 its GPT-6 Astra page put the model at 4.2%, against 12.2% for GPT-5.6 Sol. The system card says that this test uses conversations that users had flagged, so its rates run far above what people see in everyday use.

The launch figure itself moved. Fortune reported on 4 September 2026 that OpenAI changed the Astra and GPT-5.6 Sol numbers to 2% and 9.4% after the page went live, and then put back the 4.2% and 12.2% it had first shown. OpenAI told Fortune that the edits were there to make sure the numbers were its best estimate.

Independent audits moved both ways. NewsGuard found chatbots repeated false news claims 35% of the time in August 2025, up from 18% a year earlier, as their refusal rate fell from 31% to zero. Its January 2026 audit put the rate above 28%. The BBC's own comparison went the other way: answers with a significant issue fell from 51% in its first study to 37% in the round the EBU published in October 2025.

Artificial Analysis found in November 2025 that all but three models were more likely to hallucinate than answer correctly on its hard questions. In October 2026, 15 of the 30 frontier configurations it shows by default still answer more than half their misses wrongly. The median is 49.8%.

How often do AI assistants get the news wrong?

AI assistants gave a significantly flawed answer to 45% of news questions in the EBU and BBC study published on 21 October 2025, run by 22 public broadcasters in 18 countries. The EBU calls it one of the largest evaluations of its kind.

Journalists reviewed 2,709 answers from ChatGPT, Copilot, Gemini and Perplexity in 14 languages. Sourcing caused the most trouble. In 31% of the answers there was a serious sourcing problem, such as a source that was missing, wrong or misattributed. One in five had a major accuracy issue, including invented or outdated details. Google's Gemini fared worst, with significant issues in 76% of its answers. Copilot followed at 37%, ChatGPT at 36% and Perplexity at 30%.

Columbia's Tow Center ran 1,600 queries through eight AI search tools and asked each one to name the article a quote came from. The tools gave incorrect answers to more than 60% of queries, per the March 2025 write-up. Perplexity was wrong on 37%. xAI's Grok 3 was wrong on 94%.

What do AI hallucinations cost in the real world?

The clearest cost shows up in court. On 6 October 2026, Damien Charlotin's database listed 2,149 decisions in which a court dealt with AI-invented material, and 1,473 of those were in the United States.

Most involve people without a lawyer: 1,233 entries list a self-represented litigant. Lawyers account for 854 and judges for 33. Charlotin's data shows 852 decisions dated in all of 2025 and 1,208 from January to September 2026, so this year has already passed last year. Where a filing names the tool, ChatGPT appears in 139 of 235 cases. Of the 192 penalties stated in US dollars, six reached $50,000 or more.

The first famous case set a low price for it. In Mata v. Avianca, a federal judge in New York fined two lawyers and their firm $5,000 on 22 June 2023. They had cited six court decisions that ChatGPT made up.

Paid legal tools help without fixing it. Stanford RegLab's preregistered test, released in May 2024, found the AI research tools from LexisNexis and Thomson Reuters each hallucinated between 17% and 33% of the time.

Consultants aren't immune. Deloitte Australia agreed in October 2025 to repay the final instalment of an A$440,000 government contract, the Associated Press reported. A University of Sydney researcher had found up to 20 errors in its 237-page report, including an invented quote from a federal court judgment. The revised report disclosed that Azure OpenAI was used to write it.

At work the cost is quieter. In a KPMG and University of Melbourne survey of 48,000 people in 47 countries, published in April 2025, 66% of employees who use AI said they rely on its output without evaluating it. And 56% said they'd made mistakes in their work because of AI.

How we compiled these numbers

We opened every figure on this page at its primary source on 6 October 2026, and these are the sources we used:

  • Vectara's GitHub leaderboard, on its current and its previous dataset
  • Artificial Analysis's AA-Omniscience page and its November 2025 launch post
  • OpenAI's paper, system cards and launch pages
  • NewsGuard's AI False Claims Monitor
  • The EBU and BBC News Integrity report, and Columbia's Tow Center
  • Damien Charlotin's database and the Mata v. Avianca docket
  • Stanford RegLab's paper and the KPMG and University of Melbourne survey
  • Fortune and the Associated Press, for two events the companies didn't publish themselves

Some figures are our own arithmetic. The 9.6% median and the count of 50 models come from Vectara's table of 22 September 2026. The matched comparison pairs models listed under the same name on the October 2025 and September 2026 boards. We left out any name with two versions on the old board or no clear version, which drops GPT-4o, DeepSeek R1, DeepSeek V3, Command R Plus and Mistral Large. Llama 3.3 70B is matched across the Instruct build on the old board and the Turbo build on the new one. The 49.8% median covers the 30 configurations Artificial Analysis shows by default, and some are one model at two reasoning settings. The 2025 and 2026 court totals add up Charlotin's quarterly counts.

We dropped what we couldn't trace. Many statistics pages repeat a claim that AI hallucinations cost businesses $67.4 billion in 2024. The copies we found credit the AllAboutAI website, and we found no primary study behind it. We also left out two OpenAI browsing figures that appear only as chart labels in the GPT-5 system card, because we couldn't match each label to its model with confidence.

How to cite these AI hallucination statistics

Every figure on this page is free to use with a link back to this page. A plain credit line works:

  • Source: Whizi Research, "AI hallucination statistics 2026: rates across 108 models, and why the tests disagree", whizi.io/resources/ai-hallucination-statistics, updated October 2026.

When you quote a rate, put the benchmark and the date beside it. A Vectara figure and an AA-Omniscience figure for the same model can differ almost sixfold. Our ChatGPT statistics, AI chatbot market share, LLM API pricing statistics and AI subscription statistics pages follow the same sourcing rules.

Workflow checklist
  • Name the benchmark: a summary test, a memory test and a news audit measure different failures
  • Give the date and the dataset, since Vectara's 2025 and 2026 boards aren't comparable
  • Check the denominator: all answers, only the misses, or every summary
  • Say who ran the test, the vendor or an independent lab
  • Open the primary source, and drop any figure you can't trace to one
Common questions

Frequently asked questions

What is ChatGPT's hallucination rate?

OpenAI's GPT-6 Astra and GPT-6 Sol score anywhere from 4.2% to 60.1%, and it depends on the test. On Vectara's September 2026 leaderboard, GPT-6 Sol adds unsupported claims to 6.5% of its summaries and GPT-6 Astra to 8.7% of them. OpenAI reports 4.2% for Astra on its own benchmark of flagged conversations. On AA-Omniscience in October 2026, Astra answers 44.8% to 51.3% of the hard questions it misses with a wrong answer, and GPT-6 Sol does so with 60.1% of them.

Which AI model hallucinates the least?

Ant Group's Finix S1 32B is the model that hallucinates least on Vectara's summary test, at 1.8% in September 2026, and OpenAI's GPT-5.4 nano is the lowest from one of the major labs, at 3.1%. When the questions are hard and have to be answered from memory, Google's Gemini 4 Argon has the lowest rate of the 30 frontier configurations that Artificial Analysis shows, at 15.1% in October 2026.

How many court cases involve AI hallucinations?

Damien Charlotin's database listed 2,149 court decisions involving AI-invented material on 6 October 2026, and 1,473 of them were in the United States. A self-represented litigant appears in 1,233 of the entries, a lawyer in 854 and a judge in 33. The decisions dated from January to September 2026 already outnumber those from all of 2025, by 1,208 to 852.

Do AI models hallucinate less than they used to?

On a fixed test, mostly yes. OpenAI reported that GPT-5 and GPT-6 Astra each made fewer factual errors than the models before them, and the BBC's news audits fell from 51% to 37% of answers with a significant issue. The tests keep getting harder, though, so a rate from 2025 and a rate from 2026 rarely compare. When Vectara changed its dataset, 39 of 44 models that hadn't changed scored worse.

Still have a question?

Type it here. After you sign up, Whizi answers it first thing.