
AI models add unsupported claims to 1.8% to 24.2% of summaries, and guess wrong on up to 96.5% of the hard questions they miss
We opened every AI hallucination statistic below at its primary source on 6 October 2026. The median and the matched-model count are our own arithmetic on Vectara's published tables.
| Statistic | Figure | Source | Date |
|---|---|---|---|
| Lowest rate when summarising a document, 108 models | 1.8% (Ant Group Finix S1 32B) | Vectara hallucination leaderboard | 22 Sep 2026 |
| Median rate when summarising a document | 9.6% | Whizi, from Vectara's table | 22 Sep 2026 |
| Models adding unsupported claims to 10% or more of summaries | 50 of 108 | Whizi, from Vectara's table | 22 Sep 2026 |
| Same models after Vectara made its test harder | 39 of 44 scored worse, median 4.1% to 7.95% | Whizi, from Vectara's 2025 and 2026 tables | Oct 2025 and Sep 2026 |
| Lowest rate answering hard questions from memory, 30 frontier configurations | 15.1% (Gemini 4 Argon) | Artificial Analysis AA-Omniscience | 6 Oct 2026 |
| Highest rate on the same test | 96.5% (DeepSeek V4.1 Flash) | Artificial Analysis AA-Omniscience | 6 Oct 2026 |
| GPT-6 Astra on three different tests | 4.2%, 8.7% and 51.3% | OpenAI, Vectara, Artificial Analysis | Sep and Oct 2026 |
| AI news answers with at least one significant issue | 45% | EBU and BBC, 2,709 answers in 14 languages | Oct 2025 |
| Chatbot answers repeating a false news claim | more than 28% | NewsGuard, 11 chatbots | Jan 2026 audit |
| Court decisions involving AI-invented material | 2,149 | Damien Charlotin, AI Hallucination Cases database | 6 Oct 2026 |
| Employees using AI at work who say it caused mistakes in their work | 56% | KPMG and University of Melbourne, 48,000 people | Apr 2025 |
| Legal research AI tools giving hallucinated answers | 17% to 33% | Stanford RegLab | May 2024 |
The chart plots AA-Omniscience as Artificial Analysis showed it on 6 October 2026, at the highest reasoning setting it lists for each model. Across the bottom is accuracy on 6,000 hard questions. Up the side is the hallucination rate: of the questions a model didn't get right, the share it answered wrongly instead of saying it didn't know. Anthropic's Claude Fable 5.1 knows the most, at 67.2% correct. It still answers 72.6% of its misses with something false. Google's Gemini 4 Argon gets fewer right, 49.9%, and lands at 15.1% because it declines far more often. Knowing and guessing are separate skills.
What is the AI hallucination rate by model?
On Vectara's leaderboard, the AI hallucination rate by model runs from 1.8% for Ant Group's Finix S1 32B to 24.2% for Mistral's Ministral 3 3B, and the median model sits at 9.6%. Vectara gives every model the same set of more than 7,700 articles and counts how often its summary claims something that the article doesn't say. Not one of them scores zero.
| Model | Lab | Summaries with an unsupported claim |
|---|---|---|
| GPT-5.4 nano | OpenAI | 3.1% |
| Gemini 2.5 Flash Lite | 3.3% | |
| Llama 3.3 70B | Meta | 4.1% |
| GPT-6 Sol | OpenAI | 6.5% |
| DeepSeek V4 Pro | DeepSeek | 8.6% |
| GPT-6 Astra | OpenAI | 8.7% |
| Gemini 3.1 Pro (preview) | 10.4% | |
| Claude Sonnet 4.6 | Anthropic | 10.6% |
| Kimi K2.6 | Moonshot AI | 10.8% |
| Claude Opus 4.7 | Anthropic | 12.0% |
| GPT-5.6 Sol | OpenAI | 12.4% |
| Grok 4.1 Fast (reasoning) | xAI | 19.2% |
| Mistral Medium (2508) | Mistral | 22.7% |
Every row here is from Vectara's board of 22 September 2026, and the full list of all 108 models sits on Vectara's GitHub page. A lab's small model is often better at this than its flagship is: OpenAI's GPT-5.4 nano scores 3.1%, while GPT-5.6 Sol scores 12.4%. Newer isn't always lower, either. Anthropic's Claude Opus 4.7, at 12.0%, sits above the older Claude Sonnet 4.6. Vectara is open about what its test can't do. Its FAQ concedes that a model which copied sentences straight out of the article would get a perfect score.
The same model can carry three very different numbers. OpenAI's launch page for GPT-6 Astra, dated 3 September 2026, reports 4.2% on an internal benchmark. Vectara measures it at 8.7%, and Artificial Analysis puts it at 44.8% to 51.3%, depending on which reasoning setting it runs at. None of the three is wrong. Each of them asks a different question.
Why do AI hallucination benchmarks disagree?
AI hallucination benchmarks disagree because each one tests a different failure, on a different task, with a different denominator. Read the definition before you quote the number.
| Benchmark | What the model is asked to do | What the rate counts | Latest range |
|---|---|---|---|
| Vectara leaderboard | Summarise an article it's given | Summaries with a claim the article doesn't support | 1.8% to 24.2%, Sep 2026 |
| Artificial Analysis AA-Omniscience | Answer 6,000 hard questions from memory | Missed questions answered wrongly instead of declined | 15.1% to 96.5% across 30 frontier configurations, Oct 2026 |
| OpenAI SimpleQA, in its September 2025 paper | Answer short fact questions from memory | All questions answered wrongly | 26% (gpt-5-thinking-mini) to 75% (o4-mini), Sep 2025 |
| OpenAI internal benchmark | Answer ChatGPT conversations users flagged as wrong | Responses with any factual error | 4.2% (GPT-6 Astra) and 12.2% (GPT-5.6 Sol), Sep 2026 |
| NewsGuard AI False Claims Monitor | Respond to prompts about false news stories | Responses that repeat the false claim | more than 28%, Jan 2026 |
| EBU and BBC News Integrity study | Answer news questions | Answers with a significant issue, judged by journalists | 30% to 76% by assistant, Oct 2025 |
OpenAI's researchers argued in September 2025 that the way models are scored makes this worse. Their write-up says "standard training and evaluation procedures reward guessing over acknowledging uncertainty." The SimpleQA figures they used to show it are stark. OpenAI's older o4-mini got 24% of the questions right and 75% of them wrong, and it declined just 1%. The newer gpt-5-thinking-mini got 22% right, but it declined 52% and was wrong on only 26%. On a scoreboard that only counts what's right, the guesser wins.
That's why a single "hallucination rate" for a model doesn't exist. A reader who wants to see two labs disagree can send one question to two of Whizi's 280+ models in the side-by-side view, and why AI gets things wrong covers how to check the answer that follows.
Are AI hallucination rates going down over time?
On any one fixed test, AI hallucination rates have mostly fallen. But the tests keep getting harder, so a headline number can go up while the models get better. Vectara's leaderboard shows the second effect clearly.
Between October 2025 and September 2026, Vectara replaced its dataset with more than 7,700 articles, some as long as 24,000 words, and moved its judge from HHEM-2.1 to HHEM-2.3. We matched the 44 models listed under the same name on both boards, leaving out any name the old board carried in two versions. Thirty-nine scored worse on the new test. Four scored better, and DeepSeek V3.1 didn't move. The median for the 44 went from 4.1% to 7.95%, with every model the same release as before.
OpenAI's GPT-5 at high reasoning effort made the biggest jump, from 1.4% to 15.1%. So a 2025 figure and a 2026 figure from the same leaderboard can't be read as a trend. Quote the date and the dataset with the number.
OpenAI's own figures mostly point down, with one reversal. Its April 2025 system card showed that o3 hallucinated on 33% of PersonQA questions, which is double the 16% of the older o1, and that o4-mini did so on 48%. In August 2025 OpenAI said GPT-5 with thinking was about 80% less likely than o3 to make a factual error. In September 2026 its GPT-6 Astra page put the model at 4.2%, against 12.2% for GPT-5.6 Sol. The system card says that this test uses conversations that users had flagged, so its rates run far above what people see in everyday use.
The launch figure itself moved. Fortune reported on 4 September 2026 that OpenAI changed the Astra and GPT-5.6 Sol numbers to 2% and 9.4% after the page went live, and then put back the 4.2% and 12.2% it had first shown. OpenAI told Fortune that the edits were there to make sure the numbers were its best estimate.
Independent audits moved both ways. NewsGuard found chatbots repeated false news claims 35% of the time in August 2025, up from 18% a year earlier, as their refusal rate fell from 31% to zero. Its January 2026 audit put the rate above 28%. The BBC's own comparison went the other way: answers with a significant issue fell from 51% in its first study to 37% in the round the EBU published in October 2025.
Artificial Analysis found in November 2025 that all but three models were more likely to hallucinate than answer correctly on its hard questions. In October 2026, 15 of the 30 frontier configurations it shows by default still answer more than half their misses wrongly. The median is 49.8%.
How often do AI assistants get the news wrong?
AI assistants gave a significantly flawed answer to 45% of news questions in the EBU and BBC study published on 21 October 2025, run by 22 public broadcasters in 18 countries. The EBU calls it one of the largest evaluations of its kind.
Journalists reviewed 2,709 answers from ChatGPT, Copilot, Gemini and Perplexity in 14 languages. Sourcing caused the most trouble. In 31% of the answers there was a serious sourcing problem, such as a source that was missing, wrong or misattributed. One in five had a major accuracy issue, including invented or outdated details. Google's Gemini fared worst, with significant issues in 76% of its answers. Copilot followed at 37%, ChatGPT at 36% and Perplexity at 30%.
Columbia's Tow Center ran 1,600 queries through eight AI search tools and asked each one to name the article a quote came from. The tools gave incorrect answers to more than 60% of queries, per the March 2025 write-up. Perplexity was wrong on 37%. xAI's Grok 3 was wrong on 94%.
What do AI hallucinations cost in the real world?
The clearest cost shows up in court. On 6 October 2026, Damien Charlotin's database listed 2,149 decisions in which a court dealt with AI-invented material, and 1,473 of those were in the United States.
Most involve people without a lawyer: 1,233 entries list a self-represented litigant. Lawyers account for 854 and judges for 33. Charlotin's data shows 852 decisions dated in all of 2025 and 1,208 from January to September 2026, so this year has already passed last year. Where a filing names the tool, ChatGPT appears in 139 of 235 cases. Of the 192 penalties stated in US dollars, six reached $50,000 or more.
The first famous case set a low price for it. In Mata v. Avianca, a federal judge in New York fined two lawyers and their firm $5,000 on 22 June 2023. They had cited six court decisions that ChatGPT made up.
Paid legal tools help without fixing it. Stanford RegLab's preregistered test, released in May 2024, found the AI research tools from LexisNexis and Thomson Reuters each hallucinated between 17% and 33% of the time.
Consultants aren't immune. Deloitte Australia agreed in October 2025 to repay the final instalment of an A$440,000 government contract, the Associated Press reported. A University of Sydney researcher had found up to 20 errors in its 237-page report, including an invented quote from a federal court judgment. The revised report disclosed that Azure OpenAI was used to write it.
At work the cost is quieter. In a KPMG and University of Melbourne survey of 48,000 people in 47 countries, published in April 2025, 66% of employees who use AI said they rely on its output without evaluating it. And 56% said they'd made mistakes in their work because of AI.
How we compiled these numbers
We opened every figure on this page at its primary source on 6 October 2026, and these are the sources we used:
- Vectara's GitHub leaderboard, on its current and its previous dataset
- Artificial Analysis's AA-Omniscience page and its November 2025 launch post
- OpenAI's paper, system cards and launch pages
- NewsGuard's AI False Claims Monitor
- The EBU and BBC News Integrity report, and Columbia's Tow Center
- Damien Charlotin's database and the Mata v. Avianca docket
- Stanford RegLab's paper and the KPMG and University of Melbourne survey
- Fortune and the Associated Press, for two events the companies didn't publish themselves
Some figures are our own arithmetic. The 9.6% median and the count of 50 models come from Vectara's table of 22 September 2026. The matched comparison pairs models listed under the same name on the October 2025 and September 2026 boards. We left out any name with two versions on the old board or no clear version, which drops GPT-4o, DeepSeek R1, DeepSeek V3, Command R Plus and Mistral Large. Llama 3.3 70B is matched across the Instruct build on the old board and the Turbo build on the new one. The 49.8% median covers the 30 configurations Artificial Analysis shows by default, and some are one model at two reasoning settings. The 2025 and 2026 court totals add up Charlotin's quarterly counts.
We dropped what we couldn't trace. Many statistics pages repeat a claim that AI hallucinations cost businesses $67.4 billion in 2024. The copies we found credit the AllAboutAI website, and we found no primary study behind it. We also left out two OpenAI browsing figures that appear only as chart labels in the GPT-5 system card, because we couldn't match each label to its model with confidence.
How to cite these AI hallucination statistics
Every figure on this page is free to use with a link back to this page. A plain credit line works:
- Source: Whizi Research, "AI hallucination statistics 2026: rates across 108 models, and why the tests disagree", whizi.io/resources/ai-hallucination-statistics, updated October 2026.
When you quote a rate, put the benchmark and the date beside it. A Vectara figure and an AA-Omniscience figure for the same model can differ almost sixfold. Our ChatGPT statistics, AI chatbot market share, LLM API pricing statistics and AI subscription statistics pages follow the same sourcing rules.
- Name the benchmark: a summary test, a memory test and a news audit measure different failures
- Give the date and the dataset, since Vectara's 2025 and 2026 boards aren't comparable
- Check the denominator: all answers, only the misses, or every summary
- Say who ran the test, the vendor or an independent lab
- Open the primary source, and drop any figure you can't trace to one
Frequently asked questions
What is ChatGPT's hallucination rate?
OpenAI's GPT-6 Astra and GPT-6 Sol score anywhere from 4.2% to 60.1%, and it depends on the test. On Vectara's September 2026 leaderboard, GPT-6 Sol adds unsupported claims to 6.5% of its summaries and GPT-6 Astra to 8.7% of them. OpenAI reports 4.2% for Astra on its own benchmark of flagged conversations. On AA-Omniscience in October 2026, Astra answers 44.8% to 51.3% of the hard questions it misses with a wrong answer, and GPT-6 Sol does so with 60.1% of them.
Which AI model hallucinates the least?
Ant Group's Finix S1 32B is the model that hallucinates least on Vectara's summary test, at 1.8% in September 2026, and OpenAI's GPT-5.4 nano is the lowest from one of the major labs, at 3.1%. When the questions are hard and have to be answered from memory, Google's Gemini 4 Argon has the lowest rate of the 30 frontier configurations that Artificial Analysis shows, at 15.1% in October 2026.
How many court cases involve AI hallucinations?
Damien Charlotin's database listed 2,149 court decisions involving AI-invented material on 6 October 2026, and 1,473 of them were in the United States. A self-represented litigant appears in 1,233 of the entries, a lawyer in 854 and a judge in 33. The decisions dated from January to September 2026 already outnumber those from all of 2025, by 1,208 to 852.
Do AI models hallucinate less than they used to?
On a fixed test, mostly yes. OpenAI reported that GPT-5 and GPT-6 Astra each made fewer factual errors than the models before them, and the BBC's news audits fell from 51% to 37% of answers with a significant issue. The tests keep getting harder, though, so a rate from 2025 and a rate from 2026 rarely compare. When Vectara changed its dataset, 39 of 44 models that hadn't changed scored worse.