Is ChatGPT accurate? The number everyone quotes is not its score
The widely repeated finding that 45 per cent of AI news answers contain a significant issue is an average across four assistants. The one study that broke the figures out by name put ChatGPT lowest — which is not the same as putting it right.

The figure in circulation is that AI assistants get the news wrong 45 per cent of the time. It comes from real research, it is being quoted accurately as far as it goes, and it is not ChatGPT's score.
That study, published in October 2025, was coordinated by the European Broadcasting Union and led by the BBC. Twenty-two public service media organisations across 18 countries and 14 languages had journalists evaluate more than 3,000 responses from four assistants: ChatGPT, Copilot, Gemini and Perplexity. It found that 45 per cent of answers contained at least one significant issue, that 31 per cent had serious sourcing problems, and that 20 per cent contained major accuracy problems such as hallucinated details or outdated information. On a looser threshold, 81 per cent had some form of error.
The 45 per cent is the average across all four. The study published only one per-assistant number: Gemini, at 76 per cent, which the EBU described as more than double the others. ChatGPT, Copilot and Perplexity were not broken out. So anyone stating that ChatGPT is wrong 45 per cent of the time is attributing an aggregate — one that a single badly performing system pulled upward — to a specific product.
An earlier BBC study does allow that comparison. In February 2025, BBC journalists put 100 news questions to the same four assistants and scored the answers against seven criteria including accuracy, attribution and impartiality. Fifty-one per cent of answers had significant issues. On errors in how each assistant used BBC source material, the ranking was Gemini 34 per cent, Copilot 27 per cent, Perplexity 17 per cent, and ChatGPT 15 per cent.
Lowest of four. Which is worth stating precisely, because it is a narrower claim than it looks: that 15 per cent measures errors in the handling of BBC sources, not overall accuracy, and the study it comes from still found half of all answers significantly flawed.
The more useful finding is not the rate but where the failures cluster. The dominant problem in both studies is sourcing and attribution rather than fluency. Nineteen per cent of answers citing BBC content introduced factual errors — wrong statements, numbers, dates. Thirteen per cent of quotes attributed to BBC articles had been altered from the original or did not appear in the cited article at all.
That is the failure mode to understand. These systems do not produce obviously broken text you can spot by reading. They produce fluent, confident prose in which the attribution is wrong, the date has drifted, or the quotation has been smoothed into something nobody said. Nothing in the output signals which sentence is the unreliable one.
Peter Archer, the BBC's programme director for generative AI, said that despite some improvements there are still significant issues with these assistants. Jean Philip De Tender, the EBU's media director, described the problems as systemic, cross-border and multilingual.
So the practical answer depends on the task, and the studies above test the hardest version of it: open questions about current events, answered from whatever the system can retrieve. That is where sourcing breaks. Summarising a document you supply is a different task with a different failure rate, and neither study measures it.
If you take one operational rule from the research, take this one: the assistant's confidence carries no information about whether the attribution is real. Open the cited source. In roughly one case in eight, the quote will not be there.
Key takeaways
See full context →The 45 per cent figure is an average across four assistants, not ChatGPT's score. Only Gemini's individual number was published, at 76 per cent.
The one study breaking it out by name put ChatGPT lowest on source-handling errors at 15 per cent, against Gemini's 34 per cent.
Failures cluster in attribution, not fluency: 13 per cent of quotes attributed to BBC articles were altered or absent.
Source map
Explore all sources →✓What we know
- The October 2025 EBU-BBC study evaluated more than 3,000 responses across 22 organisations, 18 countries and 14 languages.
- It found 45 per cent of answers had at least one significant issue, 31 per cent serious sourcing problems and 20 per cent major accuracy problems.
- The only per-assistant figure published was Gemini's, at 76 per cent — described by the EBU as more than double the others.
- The February 2025 BBC study of 100 news questions found 51 per cent of answers had significant issues.
- On errors in handling BBC source material that study ranked Gemini 34 per cent, Copilot 27 per cent, Perplexity 17 per cent and ChatGPT 15 per cent.
- Nineteen per cent of answers citing BBC content introduced factual errors, and 13 per cent of quotes were altered or absent from the cited article.
?What remains unclear
See full context- ChatGPT, Copilot and Perplexity were not given individual overall scores in the October 2025 study, so no current per-assistant headline rate exists for them.
- The 15 per cent figure measures errors in the use of BBC source material specifically, not overall factual accuracy.
- Neither study measures summarisation of a document supplied by the user, which is a different task with a different failure rate.
- Both studies test a fixed point in time; model versions change and the figures may not describe current systems.
- How the error rates vary by topic, language or question type is not broken out in the public summaries.
Every factual claim, and what supports it
Each statement in this article is listed with how it is classified and which of the sources below establish it. A verified fact is corroborated by two or more independent sources; a reported claim rests on fewer, or on a single party’s account.
The October 2025 EBU-BBC study found 45 per cent of AI answers contained at least one significant issue.
Headline finding published by the EBU and reported independently.That study evaluated more than 3,000 responses across 22 organisations, 18 countries and 14 languages.
Scope stated in the EBU announcement.31 per cent showed serious sourcing problems and 20 per cent major accuracy problems.
Breakdown published with the headline figure.Gemini was the only assistant given an individual figure, at 76 per cent.
The EBU published Gemini's rate and described it as more than double the others; no equivalent figure was given for ChatGPT, Copilot or Perplexity.The February 2025 BBC study found 51 per cent of answers had significant issues across 100 news questions.
Published finding of the earlier BBC research.On errors handling BBC source material the ranking was Gemini 34, Copilot 27, Perplexity 17 and ChatGPT 15 per cent.
Per-assistant breakdown from the February 2025 study; measures source handling, not overall accuracy.19 per cent of answers citing BBC content introduced factual errors.
Incorrect statements, numbers and dates, per the BBC study.13 per cent of quotes attributed to BBC articles were altered or absent from the cited article.
Quotation-integrity finding of the February 2025 study.Neither study measures summarisation of a document supplied by the user.
Both test open questions about current events; no summarisation condition is described. Stated as a limit of the evidence.
How the story developed
- BBC journalists evaluate 100 news questions across four assistants; 51 per cent of answers show significant issues.
- EBU and BBC publish the larger study: 3,000+ responses, 45 per cent with a significant issue, Gemini at 76 per cent.
Why this framing matters
This is a case where a real finding is being misattributed rather than invented. The 45 per cent is accurate to its study and wrong as a statement about ChatGPT, because it averages four systems and only one of them was named. Reporting the aggregate as a product's score is the same error the studies themselves identify in AI answers: a number carried away from the thing it described. The article gives each figure with what it measured and declines to convert the 15 per cent into an accuracy rate.
5 sources reviewed
Every source used in this summary, grouped by its role in the reporting chain.
- 1European Broadcasting UnionPrimary · 2025-10Primary
- 2BBCPrimary · 2025-02Primary
- 3Digital Content NextIndependent · 2025-02-24Independent
- 4The RegisterIndependent · 2025-10-24Independent
- 5Journalism.co.ukIndependent · 2025-10Independent
How we verified this story
Every percentage is reported with its study, its date and the population it describes, because the two studies use different samples and thresholds and are routinely mixed. The per-assistant ranking is taken only from the February 2025 study, which published it; the October 2025 study is explicitly noted as not breaking out ChatGPT. No overall accuracy rate is stated for any assistant, since neither study reports one in that form. Both named officials are quoted through their organisations' own announcements.
Updates and corrections
Preview page created.
Source context and unresolved questions updated.
Join the discussion
Help build a clearer picture. Add context, challenge a claim or share a source.


