LIFEWOOD
Ready100
AI search

When an AI Gets Your Brand Wrong

Short answer. In the largest published study of answer reliability, professional journalists found significant issues in 45% of AI assistant answers about news, and serious sourcing…

Lifewood Data Technology · August 2026 · 7 min read

Download PDF

Short answer. In the largest published study of answer reliability, professional journalists found significant issues in 45% of AI assistant answers about news, and serious sourcing problems in 31% — missing, misleading or incorrect attribution (European Broadcasting Union and BBC, October 2025). There is no reason to think answers about your company are handled more carefully. There is also no correction desk: no engine offers a takedown route, a submission endpoint or a ticket queue for a factual error, so the only mechanism available is changing what the retrieval layer finds — and then waiting.

An answer engine that gets your pricing wrong fails in exactly the same way as one that gets an election wrong. Nothing in the pipeline distinguishes a brand fact from a news fact.

This piece covers what the evidence says about the failure rate, the four distinct ways an answer goes wrong about you, why you cannot simply have it corrected, what a monitoring programme that would actually catch it looks like, and the regulatory backdrop.


How often do assistants get it wrong?

The strongest published measurement was produced by the organisations with the most expertise in checking factual claims.

Finding Rate
Answers with at least one significant issue 45%
Answers with serious sourcing problems — missing, misleading or incorrect attribution 31%
Answers with major accuracy issues, including hallucinated details and outdated information 20%
Gemini, significant issues — more than double the other assistants, largely due to poor sourcing 76%

Source: European Broadcasting Union and BBC, October 2025 — over 3,000 responses evaluated by professional journalists at 22 public service media organisations across 18 countries and 14 languages, testing ChatGPT, Copilot, Gemini and Perplexity.

Two details make this transferable to brand claims. The failures were consistent across languages and territories, so it is not an artefact of one market. And the dominant failure mode was sourcing rather than invention — the assistant attached a claim to the wrong source, or to no source, more often than it fabricated the claim outright.

The caveat is real and worth stating plainly: this study is about news, not brands. No equivalent study of brand-fact accuracy exists at that scale, and model behaviour has moved since October 2025. It is the closest rigorous proxy available, and the failure mechanisms are shared.


What are the four ways an answer goes wrong about you?

Each has a different fix, and conflating them is why most "AI reputation" work goes nowhere.

Failure What it looks like Root cause What actually addresses it
Stale Old pricing, discontinued product, former executive The cited source is out of date, or the model's memory predates the change Update or retire the source page; correct the third-party record
Misattributed A true fact credited to the wrong company, or your fact credited elsewhere Sourcing failure — the 31% category Make the correct source unambiguous and easier to attribute than the wrong one
Conflated Answers merge you with a similarly-named organisation Entity resolution failure Entity identity work: canonical home, stable identifiers, consistent naming
Fabricated A claim that appears in no source Generation failure, most common where sources are thin Publish a checkable source for the true version; thin coverage invites invention

The last two are the ones brands consistently misdiagnose as content problems. Conflation is not fixed by publishing more — publishing more from an unresolvable entity adds noise. Fabrication concentrates where coverage is thin, which means the fix is supplying a source, not disputing the output.


Why can't you simply have it corrected?

There is no correction desk. No major engine offers a takedown route for a factual error about a company, no submission endpoint, and no ticket queue. The available mechanism is indirect: change what the retrieval layer finds, and wait. That is slow for three structural reasons.

  • Retrieval mode moves in days to weeks; memory mode moves when a model is retrained. If the wrong claim sits in trained memory, publishing corrections does not reach it this quarter, whatever you spend.
  • Roughly 85% of AI references point at third-party sources (Omnibound, AEO statistics compilation 2026). If the error originates in a directory entry, an old press release or a community thread, correcting your own site does not touch it.
  • Sources churn violently. With roughly 79% of ChatGPT's cited sources changing overnight (Parse, 693,509 answers, March–April 2026), an error can disappear and return without anything having been fixed — which makes it very easy to declare a false victory.

What would a monitoring programme that catches this look like?

  1. Include accuracy questions in the tracked set, not just visibility questions. "What does [company] do?", "How much does [product] cost?", "Where is [company] based?", "Who founded [company]?" Visibility tracking answers whether you were named. It does not answer whether what was said is true.
  2. Score the answer text, not just the mention. Correct, incomplete, misattributed, conflated, fabricated. A mention-rate dashboard scores a wrong answer as a success.
  3. Run it per engine and per language. The EBU/BBC failures were consistent across 14 languages, and Gemini failed at more than twice the rate of the others. An averaged score hides exactly the engine you need to act on.
  4. Repeat enough to distinguish an error from a draw. Given the churn, one wrong answer is a sample. A claim appearing in a third of runs is a problem.
  5. Trace every error to its source. Which cited URL carried the wrong claim. Without that step, correction is guesswork.
  6. Log the raw text. You cannot demonstrate that an error was corrected without the before.

The cheapest high-value addition to any tracked set is simply "What does [company] do?", run monthly across every engine. It catches conflation, staleness and fabrication in one line, and almost nobody tracks it, because it is not a visibility metric.


How do you reduce the exposure you carry?

You cannot control the generation step. You can control how much work the retrieval step has to do.

  1. Make the true version easy to find, easy to attribute and clearly dated. Sourcing failure is the largest error category, so an unambiguous, current, well-structured source is the direct countermeasure.
  2. Resolve your entity. One canonical home, a stable identifier, consistent naming, sameAs links to external records. Conflation is an identity failure, not a content failure.
  3. Fix stale third-party records first. Directories, databases, encyclopedic references and old coverage outlive the facts they contain, and are cited more often than your own pages.
  4. Publish the awkward facts. Pricing, limits, what the product is not for. Where you leave a gap, retrieval fills it from somewhere less accurate, or invents it.
  5. Date everything, machine-readably. Staleness is only detectable if the source says when it was true.
  6. Do not block retrieval crawlers. A site that cannot be fetched cannot correct the record — and per Digital Applied's 2026 crawler-access analysis, new Cloudflare domains block three major AI bots by default, without distinguishing training from retrieval.

None of this compels an output. Every countermeasure changes the odds on a system nobody outside the engine operates, and memory-mode errors in particular are not addressable on a business timeline.


What does the regulatory backdrop change?

Two markers. The European Commission's transparency obligations for general-purpose AI under the EU AI Act became fully enforceable on 2 August 2026. And Press Gazette's publisher AI tracker recorded that, as of 31 May 2026, nine organisations had active suits against Perplexity over alleged copyright or trademark infringement, including CNN, The New York Times, News Corp, Encyclopedia Britannica and Reddit.

Neither gives a company a route to correct a specific answer. What they indicate is a direction of travel: obligations on providers are increasing, and organisations with resources are using litigation where no product mechanism exists. For most companies neither is a practical remedy, which leaves an internal monitoring capability as the only reliable option.


How Lifewood approaches this

Lifewood runs accuracy questions inside the same tracked prompt sets it uses for visibility, scored on the answer text rather than the mention, per engine and per language, with raw runs retained so a corrected error can be demonstrated against its before. Errors are classified as stale, misattributed, conflated or fabricated before any work is commissioned, because those four route to entirely different fixes and the third and fourth are not content problems at all.

The remediation work is mostly outside the client's own site: third-party records, entity identity, and publishing checkable sources for the facts that retrieval is currently inventing. Across 50+ languages and 40+ delivery centres across 30+ countries, that monitoring is run in-market, because an answer that is accurate in English is routinely wrong in Japanese.

See AEO services, GEO services, AI data validation and reducing LLM hallucinations.


Sources and further reading

  • European Broadcasting Union and BBC, international study of AI assistants and news, October 2025 — 3,000+ responses, 22 organisations, 18 countries, 14 languages.
  • Parse, AI citation volatility by industry — 693,509 answers, March–April 2026.
  • Omnibound, Answer Engine Optimization statistics 2026 — third-party citation share.
  • European Commission, EU AI Act transparency obligations for general-purpose AI, effective 2 August 2026.
  • Press Gazette, publisher AI deals and lawsuits tracker.
  • Digital Applied, AI crawler access control: the 2026 decision matrix.

Frequently asked questions

In the largest published study, professional journalists found significant issues in 45% of AI answers about news, with 31% showing serious sourcing problems and 20% containing major accuracy issues. The EBU/BBC study covered over 3,000 responses across 22 organisations, 18 countries and 14 languages, and the failures were consistent across languages. It measured news, not brand facts, and no equivalent brand study exists at that scale.

Gemini performed worst, with significant issues in 76% of responses — more than double the other assistants tested — driven mainly by poor sourcing. ChatGPT, Copilot and Perplexity performed notably better, though all four showed substantial error rates. The study was published in October 2025 and model behaviour changes.

There is no correction desk, takedown route or submission endpoint on any major engine. The only available mechanism is indirect: make the accurate source easier to find and attribute than the wrong one, correct the third-party records, and wait for retrieval to reflect it. No correction can be guaranteed.

That is entity resolution failure rather than a content problem. If a machine cannot resolve which organisation a name refers to, publishing more content from the ambiguous entity adds noise. The fix is a canonical entity home, a stable identifier, consistent naming, and `sameAs` links to external records.

Thin coverage. Fabrication concentrates on questions where no good source exists — pricing, limits, what the product is not for — because the generation step has nothing to retrieve. Publishing a checkable, dated source for the true version is a more effective response than disputing the output.

Add accuracy questions to the tracked set — what the company does, what it costs, where it is based, who runs it — score the answer text rather than the mention, run it per engine and per language, repeat enough that one wrong answer is not mistaken for a pattern, and trace each error to the source URL that carried it.

On retrieval surfaces, days to weeks once the underlying source is fixed, though the churn makes any single observation unreliable. If the error sits in a model's trained memory rather than in retrieved documents, it does not change until a new model ships, regardless of what you publish.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team