Is ChatGPT accurate? The number everyone quotes is not its score

The widely repeated finding that 45 per cent of AI news answers contain a significant issue is an average across four assistants. The one study that broke the figures out by name put ChatGPT lowest — which is not the same as putting it right.

Pre-launch editorial previewDemonstration content · Human reviewer identity pending · Advertising disabled
Verified○ Context added5 Sources
+1
Printed pages fanned across a desk in raking light, one lifted with a passage marked in the margin
AI-generated illustration — not a photograph of the events describedPrinted pages fanned across a desk in raking light, one lifted with a passage marked in the margin AI-generated illustration · Open News

The figure in circulation is that AI assistants get the news wrong 45 per cent of the time. It comes from real research, it is being quoted accurately as far as it goes, and it is not ChatGPT's score.

That study, published in October 2025, was coordinated by the European Broadcasting Union and led by the BBC. Twenty-two public service media organisations across 18 countries and 14 languages had journalists evaluate more than 3,000 responses from four assistants: ChatGPT, Copilot, Gemini and Perplexity. It found that 45 per cent of answers contained at least one significant issue, that 31 per cent had serious sourcing problems, and that 20 per cent contained major accuracy problems such as hallucinated details or outdated information. On a looser threshold, 81 per cent had some form of error.

The 45 per cent is the average across all four. The study published only one per-assistant number: Gemini, at 76 per cent, which the EBU described as more than double the others. ChatGPT, Copilot and Perplexity were not broken out. So anyone stating that ChatGPT is wrong 45 per cent of the time is attributing an aggregate — one that a single badly performing system pulled upward — to a specific product.

An earlier BBC study does allow that comparison. In February 2025, BBC journalists put 100 news questions to the same four assistants and scored the answers against seven criteria including accuracy, attribution and impartiality. Fifty-one per cent of answers had significant issues. On errors in how each assistant used BBC source material, the ranking was Gemini 34 per cent, Copilot 27 per cent, Perplexity 17 per cent, and ChatGPT 15 per cent.

Lowest of four. Which is worth stating precisely, because it is a narrower claim than it looks: that 15 per cent measures errors in the handling of BBC sources, not overall accuracy, and the study it comes from still found half of all answers significantly flawed.

The more useful finding is not the rate but where the failures cluster. The dominant problem in both studies is sourcing and attribution rather than fluency. Nineteen per cent of answers citing BBC content introduced factual errors — wrong statements, numbers, dates. Thirteen per cent of quotes attributed to BBC articles had been altered from the original or did not appear in the cited article at all.

That is the failure mode to understand. These systems do not produce obviously broken text you can spot by reading. They produce fluent, confident prose in which the attribution is wrong, the date has drifted, or the quotation has been smoothed into something nobody said. Nothing in the output signals which sentence is the unreliable one.

Peter Archer, the BBC's programme director for generative AI, said that despite some improvements there are still significant issues with these assistants. Jean Philip De Tender, the EBU's media director, described the problems as systemic, cross-border and multilingual.

So the practical answer depends on the task, and the studies above test the hardest version of it: open questions about current events, answered from whatever the system can retrieve. That is where sourcing breaks. Summarising a document you supply is a different task with a different failure rate, and neither study measures it.

If you take one operational rule from the research, take this one: the assistant's confidence carries no information about whether the attribution is real. Open the cited source. In roughly one case in eight, the quote will not be there.

Key takeaways

See full context
THE CORE

The widely repeated finding that 45 per cent of AI news answers contain a significant issue is an average across four assistants. The one study that broke the figures out by name put ChatGPT lowest — which is not the same as putting it right.

FACTS CHECKED

6 facts cross-checked

Compared across 5 primary and independent sources.

?
WHAT’S NEXT

We are tracking 5 open questions

We’ll update this page as stronger evidence emerges.

Source map

Explore all sources
Newswire1
TV / Digital1
Official1
Experts1
Documents1

What we know

  • Multiple independent sources support the central development.
  • The timeline reflects the latest verified update.
  • Confirmed facts are separated from analysis and projections.

?What remains unclear

See full context
  • 5 material questions still need stronger evidence.
  • Forecasts may change as official information is released.
TIMELINE

How the story developed

  1. Initial evidence set assembled

    Primary material and independent reporting were grouped for comparison.

  2. Context and open questions added

    The preview was updated to separate supported points from unresolved claims.

EDITORIAL CONTEXT

Why this framing matters

This page focuses on the evidence shared across sources, identifies where reporting diverges, and avoids treating forecasts as established facts. It is intended to complement—not replace—the original reporting.

▢ Discuss 64
SOURCE MAP

5 sources reviewed

Every source used in this summary, grouped by its role in the reporting chain.

  1. 1European Broadcasting UnionPrimary / official sourcePrimary
  2. 2BBCIndependent reportingCross-check
  3. 3Digital Content NextIndependent reportingCross-check
  4. 4The RegisterIndependent reportingCross-check
  5. 5Journalism.co.ukIndependent reportingCross-check

Sources are listed for transparency. Open News summarizes and links; it does not copy full source articles.

How we verified this story

Open News compared the claims above across primary documents and independent reporting. Status labels reflect the strength and agreement of the available evidence.

REVISION HISTORY

Updates and corrections

  1. Preview page created.

  2. Source context and unresolved questions updated.

Join the discussion

Help build a clearer picture. Add context, challenge a claim or share a source.

G
0/1000
3 comments
AM
Amit M.

The reporting agrees on the direction, but the exact timeline still depends on local infrastructure and permitting.

Reuters — Full report
YS
Yael S.Context

Important context: the public commitments are not the same as completed capacity. The implementation gap is still material.

Comments are screened before publication. Community rules · Report a violation