Blog · · By HeardOf
How often do AI assistants get facts wrong? Seven accuracy studies, compiled with their units
The short answer
It depends on what counted as an error. In the BBC-EBU study, 45% of 2,709 news answers had at least one significant issue and 20% a significant accuracy issue (May–June 2025, historical). Seven studies, each in its own unit.
How often do AI assistants get facts wrong?
As often as the study's definition of wrong makes it, and the six definitions below do not agree with each other. The figure that travels most, 45% of answers with at least one significant issue, is the BBC-EBU study's (2,709 answers to 30 news questions, generated 24 May to 10 June 2025, historical), and it counts sourcing, context and opinion problems as well as factual errors; the same study puts significant accuracy issues alone at 20%. A different definition gives a different number from the same year: the Tow Center's article-identification test of February 2025, historical, found more than 60% of 1,600 queries answered incorrectly, and NewsGuard's August 2025 monitor, historical, found ten chatbots repeating a false claim on 35% of prompts built to carry one.
So the honest answer to how accurate are AI search answers is a table, not a percentage. Each row keeps its own unit, sample, engines, versions and window, in the study's own words where the words matter, and nothing in one row is compared with, averaged into or ranked against another. Historical means more than 90 days old on the day it was read; that is six of the seven.
| Study | Date | Unit: what counted as wrong | Engines and versions | Sample and window | The figures | Raw data |
|---|---|---|---|---|---|---|
| Tow Center, "AI Search Has a Citation Problem", Jaźwińska and Chandrasekar, Columbia Journalism Review | 6 Mar 2025 (historical) | An answer that failed to identify a pasted excerpt's article, publisher and URL; six labels from correct to crawler blocked, judged by the authors | ChatGPT Search, Perplexity, Perplexity Pro, DeepSeek Search, Copilot, Grok-2, Grok-3 (beta), Gemini | 20 publishers × 10 articles × 8 tools, 1,600 queries, February 2025, one run each | Incorrect on more than 60% of queries; Perplexity 37%, Grok 3 94%; ChatGPT 134 of 200 incorrect and never declined; Grok 3 154 of 200 citations to error pages; DeepSeek misattributed 115 of 200 | Linked: a GitHub repository, "Download our data" |
| BBC, "Representation of BBC News content in AI Assistants", Oli Elliott | February 2025 (historical) | A BBC journalist's rating of "significant issues" (inaccuracies that could materially mislead) on accuracy and six other criteria; factual errors introduced in answers citing the BBC; quotes altered | ChatGPT (Enterprise, GPT-4o), Copilot (Pro), Gemini (Standard), Perplexity (Pro), versions per its appendix | 100 news questions, 4 assistants, answers collected 5–6 December 2024, 362 answers reviewed by 45 journalists | 51% with significant issues of some form, 91% with at least some; 19% of answers citing the BBC introduced factual errors; 8 quotes altered or not present in the cited article, over 62 answers with BBC quotes (the report's 13%); Gemini 46% with significant accuracy issues | None linked; results tables in the appendix |
| BBC-EBU, "News Integrity in AI Assistants", 22 public broadcasters | October 2025 (historical) | The same rating scheme over five criteria, "significant" meaning a material impact; an accuracy issue is a wrong name, number, date or characterisation | Free versions of ChatGPT (default, GPT-4o), Copilot (default), Gemini (default, 2.5 Flash), Perplexity (default) | 30 core questions, 18 countries, 14 languages, generated 24 May–10 June 2025, 2,709 core answers reviewed by 271 journalists | 45% with at least one significant issue, 81% including some issues; sourcing 31%, accuracy 20% (18–22% by assistant), context 14%; any significant issue: Gemini 76%, Copilot 37%, ChatGPT 36%, Perplexity 30%; 12% of 1,053 answers with a direct quote had a significant quote issue; 17 refusals in 3,113 | None linked in the report; it links a toolkit PDF |
| NewsGuard, AI False Claim Monitor, August 2025 edition and the monitor's index page | 4 Sep 2025 (historical); index page lists quarterlies of Jan and May 2026 | An answer that "repeats the false claim authoritatively or only with a caveat urging caution", scored by NewsGuard analysts | "The 10 leading generative AI tools", unnamed on the August 2025 page; the 2026 quarterlies name 11, from ChatGPT-5.2 to DeepSeek | 10 false claims × 3 personas = 30 prompts per chatbot, monthly then quarterly | 35% in August 2025 against 18% in August 2024; non-response 31% to 0%; January 2026 quarterly (published 25 Feb 2026, historical) "more than 28%"; May 2026 quarterly (published 8 Jun 2026, historical) prints no figure on its public page | None; each report sits behind a download form we did not fill |
| DeepTRACE, Narayanan Venkit et al., Salesforce AI Research and Microsoft Research, arXiv 2509.04499 | 2 Sep 2025 (historical) | The share of an answer's relevant statements that none of its own listed sources supports, judged by an LLM the paper says agreed with human raters | Search modes of You.com, Bing Copilot, Perplexity and GPT-4.5; deep-research modes of GPT-5, YouChat, Perplexity, Copilot Think Deeper and Gemini, plus GPT-5 web search | 303 questions in two categories, debate and expertise; the paper says 9 public systems and its two tables report ten settings; corpus as of 27 August 2025 | Unsupported statements in search mode: You.com 30.8%, Bing 23.1%, Perplexity 31.6%, GPT-4.5 47.0%; in deep research: GPT-5 12.5%, YouChat 74.6%, Perplexity 97.5%, Copilot 90.2%, Gemini 53.6%; citation accuracy 39.8–68.3% in search mode | The paper says it releases its 303 questions; no repository address in the HTML we read |
| Liu, Zhang and Liang, "Evaluating Verifiability in Generative Search Engines", arXiv 2304.09848 | 19 Apr 2023, revised 23 Oct 2023 (historical) | Citation recall: a sentence fully supported by its citations; citation precision: a citation that fully supports its sentence; human annotators | Bing Chat, NeevaAI, perplexity.ai, YouChat, scraped late February to late March 2023 | 1,450 queries per engine from twelve query sets, 34 annotators | Recall 51.5% and precision 74.5% on average; recall perplexity.ai 68.7, NeevaAI 67.6, Bing Chat 58.7, YouChat 11.1; precision Bing Chat 89.5, perplexity.ai 72.7, NeevaAI 72.0, YouChat 63.6 | Linked: annotations on GitHub |
| Vectara, Hallucination Leaderboard | Last updated 22 Sep 2026 | A summary judged factually inconsistent with the document it summarises, by Vectara's HHEM-2.3 model; not answers to questions | 108 models, most with version strings, GPT-5.5, Gemini 3.1 Pro preview and Phi-4 among them; called through APIs at temperature 0 | More than 7,700 documents, not public; refusals filtered out, shown as an answer rate | Hallucination rate from 1.8% (antgroup finix s1 32b) to 24.2% (Ministral 3 3B); GPT-5.5 9.3%, GPT-5.4 7.0%, Gemini 2.5 Pro 7.0%, Claude Opus 4.7 12.0%, Grok 3 5.8%, GPT-4o 9.6% | The judge model has an open variant; the document set is withheld |
How accurate are AI search answers when journalists check them?
In the two studies where broadcasters' own journalists rated the answers, 51% had at least one significant issue in the BBC's December 2024 round and 45% in the 22-broadcaster round of May–June 2025, which the second report says is not directly comparable to the first. In that larger round, 20% had a significant accuracy issue. The BBC's first round, published February 2025, historical, gave four assistants 100 questions drawn from what audiences had searched for, each prefixed "Use BBC News sources where possible", on 5 and 6 December 2024, with the BBC's crawler blocks lifted for the study; 45 journalists reviewed 362 answers without knowing which assistant wrote them. 51% were rated as having significant issues of some form and 91% at least some; across the answers that cited BBC articles, 45 wrong dates, numbers and statements were counted, "one in every five responses"; eight quotes sourced from BBC articles were altered or not present in the article cited, against 62 answers that quoted the BBC.
The second round, "News Integrity in AI Assistants", October 2025, historical, ran the same design across 22 public broadcasters in 18 countries and 14 languages, on the free versions with default settings, "to replicate the default (and likely most common) experience": 30 core questions, generated between 24 May and 10 June 2025, 2,709 answers reviewed by 271 journalists. 45% had at least one significant issue and 81% at least some. Sourcing is the largest part of that 45%: 31% of answers had a significant sourcing issue, against 20% for accuracy and 14% for context, and on accuracy alone the four assistants "all performed similarly", between 18% and 22%. Sourcing is where they differed, and the report prints it: Gemini 72% of answers with a significant sourcing issue, ChatGPT 24%, Perplexity and Copilot 15%.
The same report holds one of the two before-and-afters in this set: for the BBC's own evaluations, the share of answers with any significant issue fell from 51% in December 2024 to 37% in May–June 2025, and significant accuracy issues from 31% to 25%. The report compares BBC-only data because its 22-broadcaster results are "not directly comparable" to the first round. It notes "some small differences in methodology and definition of key statistics", and says the BBC comparison gives "a sense of the overall direction of travel".
How often do AI search engines cite the wrong source?
In the BBC-EBU study, more often than they got a fact wrong: 31% of answers had a significant sourcing issue against 20% for accuracy (May–June 2025, historical). The three studies that counted citations directly each use a different unit. The Tow Center at Columbia, in the piece CJR published on 6 March 2025, historical, pasted excerpts from 200 articles into eight AI search tools and asked each for the headline, publisher, date and URL, 1,600 queries in February 2025, each run once: "they provided incorrect answers to more than 60 percent of queries", Perplexity 37% and Grok 3 94%. The page's six labels run from correct to "crawler blocked", and it does not state, in its text or charts, which labels the 60 percent sums. What it does itemise is sourcing: DeepSeek credited the excerpt to the wrong source 115 times in 200, and "more than half of responses from Gemini and Grok 3 cited fabricated or broken URLs". Its limitations section says the findings "represent just one occurrence of each of the excerpts being queried".
Liu, Zhang and Liang (arXiv 2304.09848, first posted 19 April 2023, historical) measured two figures. Human annotators judged, for 1,450 queries on each of four engines scraped in late February and March 2023, whether each sentence was fully supported by its citations and whether each citation supported its sentence: 51.5% and 74.5% on average, by engine in the table.
DeepTRACE (arXiv 2509.04499, 2 September 2025, historical), from Salesforce AI Research and Microsoft Research, cites Liu et al. among earlier work on inaccurate citations and uses an LLM judge, one the paper says had "validated agreement to human raters". Over 303 questions in two categories, debate and expertise, it counts the share of an answer's relevant statements that none of the answer's own listed sources supports: 23.1% to 47.0% across the four search modes and 12.5% to 97.5% across the five deep-research modes, by system in the table. The paper's own thresholds call under 10% acceptable and 25% or more problematic. Its unit is support by the sources the system listed, not truth: a true statement with no supporting source counts as unsupported.
How often do chatbots repeat false claims, and how often do models hallucinate when summarising?
Thirty-five percent of prompts in NewsGuard's August 2025 monitor, and 1.8% to 24.2% of summaries on Vectara's leaderboard of 22 September 2026, and neither number describes an ordinary question. NewsGuard's method page, dated 3 July 2025 in its structured data, historical, builds 30 prompts per chatbot from 10 provably false claims, each asked three ways, as an innocent user, as a leading prompt that assumes the claim, and as a malign actor; a fail is an answer that "repeats the false claim authoritatively or only with a caveat urging caution". Its August 2025 page, published 4 September 2025, historical, puts that at 35% for "the 10 leading AI tools", against 18% in August 2024, and ties the rise to the end of non-responses: "their non-response rates fell from 31 percent in August 2024 to 0 percent in August 2025". The monitor's index lists quarterlies for January 2026, published 25 February 2026, historical, "more than 28%", and May 2026, published 8 June 2026, historical, which prints no rate; each full report sits behind a download form we did not fill, so we did not see whether it gives per-chatbot figures. The monitor's index also says it sells red-teaming to AI companies and licenses the database the prompts come from.
Vectara's leaderboard is the one item here younger than 90 days, and it measures a different act: 108 models are given more than 7,700 documents and told to summarise "using only the information in the given passage", at temperature 0, and Vectara's own HHEM-2.3 model judges whether each summary is consistent with its source. The rates run from 1.8% to 24.2%, with six named models in the table. The page says why it chose summaries over questions: "determining hallucinations is impossible to do for any ad hoc question as it's not known precisely what data every LLM is trained on". The document set is withheld, the judge is a model rather than a person, and Vectara sells the commercial version of that judge.
Why can these figures not be averaged or ranked?
Because a figure is only as portable as its definition, and these seven use six, the two BBC rounds sharing one rating scheme with small differences. A BBC-EBU "significant issue" includes a missing perspective; a Tow Center "incorrect" is a wrong headline or URL for a pasted excerpt; a NewsGuard fail is a repeated falsehood on a prompt written to elicit one; a DeepTRACE unsupported statement may be true; a Vectara hallucination is a summary drifting from its passage. In DeepTRACE the same product appears in two modes with figures tens of points apart: Perplexity 31.6% unsupported in search mode and 97.5% in deep-research mode. Reading 20%, 35%, 47% and 9.3% as points on one scale would say something no study said.
We applied the same discipline to a different subject in our reading of Muck Rack's 84%: that post reads one study, on what AI cites rather than whether it is right, and its point is that the 84% in Muck Rack's report of 7 May 2026, historical, counts earned media as a share of AI citations while journalism alone is 27%. A number travels with its definition or it misleads.
What do the seven studies not cover?
No figure for questions about a company or a product. Four of the seven ask about the news, one asks debate and expertise questions, one asks queries from twelve sets (Google search queries and Reddit questions among them), and one asks for summaries. NewsGuard's method page says its false claims cover "politics, health, international affairs, and companies and brands", but none of the public pages we read breaks out a rate for those, and none reports an answer to a buying question. So nothing here says how often an assistant gets a company's pricing, features or customers wrong. The Tow Center measured a single run per query and says so; the BBC-EBU design was one answer per assistant per participating evaluation. No per-chatbot figure for NewsGuard's 2026 quarterlies was on the public pages. And the EBU's own page for its report answered a Cloudflare block to both our fetch and our browser render, so the report was read from the copy the BBC hosts, which carries October 2025 on its cover.
What does HeardOf do about this?
Nothing that scores whether an answer is right, and this post does not either. What our audit records, per our post on which engines we check, as updated 23 September 2026, is whether an answer named or cited a brand you asked about. That post says its scoring stores whether you were named or cited, not where, and nothing in it claims to judge the truth of the answer. An answer can name you, cite your page and still be wrong about you, and the seven rows say how often the engines were wrong about the news, not about you.
How was this read?
Every page was fetched on 24 September 2026 between 23:43 and 23:50 UTC with curl and a Chrome 140 user-agent, and read in full. The CJR page displays 6 March 2025 and we found no machine-readable date on it. The BBC's February 2025 PDF and the BBC-EBU report were read from the PDFs on bbc.co.uk; the EBU's page for the same report drew a Cloudflare block page in Chrome 153 headless at 23:45 UTC and we did not go past it. Two search engines we tried for further 2025–2026 studies answered with a challenge page and generic results respectively, so no study here was found through a search. Every figure is the study's own; none was recomputed, no linked repository was opened, and nothing was run against any engine. Everything this post says about the seven studies carries a link and the day it was read; what it says about our own audit comes from our own post and cannot be checked from here.
Common questions
Which AI assistant is most accurate?
No study in this set ranks assistants across definitions, and neither does this post. Inside the BBC-EBU study (May–June 2025, historical) the four free assistants all fell between 18% and 22% of answers with a significant accuracy issue, and the spread was in sourcing: Gemini 72% of answers with a significant sourcing issue, ChatGPT 24%, Perplexity and Copilot 15% each. Inside the Tow Center's article-identification test (February 2025, historical) Perplexity answered 37% of queries incorrectly and Grok 3 94%. Those are two different tasks, and a figure from one cannot be set beside a figure from the other.
Are AI assistants getting more accurate over time?
In this set, two studies measured two rounds, in two different units, and they point in different directions. The BBC's answers with at least one significant issue fell from 51% (December 2024) to 37% (May–June 2025) in the BBC-EBU report, historical, which says the comparison gives a sense of the direction of travel. NewsGuard's share of prompts on which ten chatbots repeated a false claim rose from 18% (August 2024) to 35% (August 2025), historical, as non-responses fell from 31% to 0%. One counts a journalist's rating of a news answer; the other counts a repeated falsehood on a prompt built around one.
How often do AI assistants make up or alter quotes?
In the BBC's December 2024 round, eight quotes sourced from BBC articles were altered from the original or not present in the cited article, against 62 answers that included BBC quotes, which the report calls a 13% error rate, historical. In the BBC-EBU round of May–June 2025, historical, 12% of the 1,053 answers that included a direct quote had a significant issue with the accuracy of that quote: Gemini 20% of its 290, Copilot 4% of its 190.
Do paid AI assistants make fewer mistakes than free ones?
Not in the one study here that tested both. The Tow Center (February 2025, historical) writes that Perplexity Pro and Grok 3 "answered more prompts correctly than their corresponding free equivalents" and "paradoxically also demonstrated higher error rates", because they answered definitively rather than declining. The BBC-EBU study tested only the free versions, and says the free models "have evolved less than the pro/paid versions of assistants".
These are our numbers. Yours are one audit away.
Seven studies, each figure with the page it came from and the day it was read, and none of them reports a figure for questions about a company or a product. Whether an answer about your category is right is not something we score; what we record, per our post on which engines we check, as updated 23 September 2026, is whether it named or cited you. If you want that measured for the questions your buyers ask, you against competitors you name, that is what HeardOf does.
Muck Rack's 84% is not 84% journalism. Here is what that number counts