Skip to main content
AEO/GEO

When an AI Gets Your Brand Wrong

August 2026 · 8 min read · Updated September 2026

Short answer. In the largest published study of answer reliability, professional journalists found significant issues in 45% of AI assistant answers about news, and serious sourcing problems in 31% — missing, misleading or incorrect attribution (European Broadcasting Union and BBC, October 2025). There is no reason to think answers about your company are handled more carefully, and no engine offers a takedown route or a ticket queue for a factual error.

An answer engine that gets your pricing wrong fails the same way one that gets an election wrong fails: nothing in the pipeline distinguishes a brand fact from a news fact. This piece covers the failure rate, the four ways an answer goes wrong about a company, why correction is not straightforward, what a monitoring programme that actually catches it looks like, and the regulatory backdrop.

Key takeaways

  • The EBU/BBC study evaluated over 3,000 AI assistant responses across 22 public service media organisations, 18 countries and 14 languages, and found significant issues in 45% of answers.
  • Gemini showed significant issues in 76% of responses in that study, more than double the other three assistants tested.
  • No major AI answer engine offers a correction desk, takedown route or submission endpoint for a factual error about a company.
  • Brand-answer errors fall into four distinct types — stale, misattributed, conflated and fabricated — and each needs a different fix.
  • The European Commission's transparency obligations for general-purpose AI under the EU AI Act became fully enforceable on 2 August 2026.

How often do assistants get it wrong?

The strongest published measurement of AI-assistant accuracy comes from the journalism organisations with the most experience checking factual claims, and it found errors in roughly half of tested answers.

Finding Rate
Answers with at least one significant issue 45%
Answers with serious sourcing problems — missing, misleading or incorrect attribution 31%
Answers with major accuracy issues, including hallucinated details and outdated information 20%
Gemini, significant issues — more than double the other assistants, largely due to poor sourcing 76%

Source: European Broadcasting Union and BBC, October 2025 — over 3,000 responses evaluated by professional journalists at 22 public service media organisations across 18 countries and 14 languages, testing ChatGPT, Copilot, Gemini and Perplexity.

Two details make this transferable to brand claims. The failures were consistent across languages and territories, so it is not an artefact of one market. And the dominant failure mode was sourcing rather than invention — the assistant attached a claim to the wrong source, or to no source, more often than it fabricated the claim outright.

The caveat is worth stating plainly: this study measured news, not brands. No equivalent study of brand-fact accuracy exists at this scale, and model behaviour has moved since October 2025. It remains the closest rigorous proxy available, because the underlying failure mechanisms — thin sourcing, entity confusion, stale citations — are shared between news and brand claims, a pattern explored further in how AI search engines decide which brands to mention and cite.

What are the four ways an answer goes wrong about you?

Each of the four failure types has a different root cause and a different fix, and conflating them is why most "AI reputation" work goes nowhere.

Failure What it looks like Root cause What actually addresses it
Stale Old pricing, discontinued product, former executive The cited source is out of date, or the model's memory predates the change Update or retire the source page; correct the third-party record
Misattributed A true fact credited to the wrong company, or your fact credited elsewhere Sourcing failure — the 31% category above Make the correct source unambiguous and easier to attribute than the wrong one
Conflated Answers merge you with a similarly named organisation Entity resolution failure Entity identity work: a canonical home, stable identifiers, consistent naming
Fabricated A claim that appears in no source Generation failure, most common where sources are thin Publish a checkable source for the true version; thin coverage invites invention

Entity resolution is the process by which a machine decides which real-world organisation a name refers to. Conflation is not fixed by publishing more content — publishing more from an unresolvable entity adds noise rather than clarity, a problem covered in more depth in structured data and entity identity for AEO. Fabrication concentrates where coverage is thin, which means the fix is supplying a source, not disputing the output.

Why can't you simply have it corrected?

There is no correction desk on any major answer engine, so the only lever available is changing what the retrieval layer finds and waiting for it to take effect.

No major engine offers a takedown route for a factual error about a company, a submission endpoint, or a ticket queue. That indirect mechanism is slow for three structural reasons.

  • Retrieval mode can update in days to weeks once a source is fixed. Memory mode — facts baked into the model during training — only changes when a model is retrained, so a correction published today may not reach it this quarter regardless of spend.
  • A large share of AI citations point at third-party sources rather than a company's own site, a dynamic explored in why third-party brand mentions matter for GEO and AI search. If the error originates in a directory entry, an old press release or a community thread, correcting your own site does not touch it.
  • Cited sources churn heavily from one run to the next. An error can disappear and reappear without anything having actually been fixed, which makes it easy to mistake churn for a correction.

What would a monitoring programme that catches this look like?

A programme that actually catches brand-answer errors tracks the accuracy of the answer text itself, not just whether the brand was named.

  1. Include accuracy questions in the tracked set, not just visibility questions — "What does [company] do?", "How much does [product] cost?", "Where is [company] based?", "Who founded [company]?" A visibility metric answers whether you were named; it does not answer whether what was said is true, a distinction covered in how to measure AI visibility without fooling yourself.
  2. Score the answer text, not just the mention — correct, incomplete, misattributed, conflated, fabricated. A mention-rate dashboard scores a wrong answer as a success.
  3. Run it per engine and per language. The EBU/BBC failures were consistent across 14 languages, and Gemini failed at more than twice the rate of the other assistants tested; an averaged score hides exactly the engine that needs action.
  4. Repeat enough to distinguish an error from a one-off. Given how much citation sets churn, one wrong answer is a single sample; a claim appearing across a third of runs is a pattern.
  5. Trace every error to its source, meaning the specific cited URL that carried the wrong claim, because without that step correction is guesswork.
  6. Log the raw answer text, since demonstrating that an error was corrected requires the before as well as the after.

The cheapest high-value addition to any tracked set is simply "What does [company] do?", run monthly across every engine — it catches conflation, staleness and fabrication in one line, and almost nobody tracks it because it is not a visibility metric.

How do you reduce the exposure you carry?

A company cannot control the generation step of an answer engine, but it can reduce how much guessing the retrieval step has to do.

  1. Make the true version easy to find, easy to attribute and clearly dated. Sourcing failure is the largest error category above, so an unambiguous, current, well-structured source is the direct countermeasure, as covered in what actually gets you cited by AI answer engines.
  2. Resolve the entity: one canonical home, a stable identifier, consistent naming, and sameAs links to external records. Conflation is an identity failure, not a content failure.
  3. Fix stale third-party records first. Directories, databases, encyclopedic references and old coverage outlive the facts they contain, and are frequently cited more often than a company's own pages.
  4. Publish the awkward facts — pricing, limits, what the product is not for. Where a company leaves a gap, retrieval fills it from somewhere less accurate, or invents it.
  5. Date everything, machine-readably, since staleness is only detectable if the source states when it was true.
  6. Keep retrieval crawlers unblocked. A site that cannot be fetched cannot correct the record, and default crawler-blocking rules on some hosting platforms do not distinguish AI training bots from the retrieval bots that answer engines use to check facts in real time.

None of this compels a specific output. Every countermeasure changes the odds on a system nobody outside the engine operator controls, and memory-mode errors in particular are not addressable on a business timeline.

What does the regulatory backdrop change?

Regulation and litigation are moving, but neither currently gives a company a direct route to correct one wrong answer.

The European Commission's transparency obligations for general-purpose AI under the EU AI Act became fully enforceable on 2 August 2026. Separately, Press Gazette's tracker of publisher deals and lawsuits against AI companies has recorded a growing list of suits against Perplexity over alleged copyright or content-use disputes, brought by organisations including News Corp, The New York Times, Encyclopedia Britannica and Yomiuri Shimbun.

What these two developments indicate is a direction of travel: obligations on providers are increasing, and organisations with resources are using litigation where no product mechanism exists. For most companies neither is a practical remedy today, which leaves an internal monitoring capability as the more reliable option in the meantime.

How does Lifewood approach brand-accuracy monitoring?

Lifewood runs accuracy questions inside the same tracked prompt sets used for visibility tracking, scored on the answer text rather than the mention alone.

Runs are scored per engine and per language, with raw output retained so a corrected error can be demonstrated against its before. Errors are classified as stale, misattributed, conflated or fabricated before any remediation work is commissioned, because those four types route to entirely different fixes, and two of them are not content problems at all.

The remediation work itself is mostly outside a client's own site — third-party records, entity identity, and publishing checkable sources for facts that retrieval is currently inventing. Across 100+ languages and 40+ delivery centres across 30+ countries, that monitoring is run in-market, because an answer that reads as accurate in English is routinely wrong in another language, a gap covered further in AEO in the markets Google does not own. Programmes like this sit within Lifewood's AEO and GEO services.

Frequently asked questions

In the largest published study, professional journalists found significant issues in 45% of AI answers about news, with 31% showing serious sourcing problems and 20% containing major accuracy issues. The EBU/BBC study covered over 3,000 responses across 22 organisations, 18 countries and 14 languages. It measured news, not brand facts, and no equivalent brand study exists at that scale.

Gemini performed worst, with significant issues in 76% of responses — more than double the other assistants tested — driven mainly by poor sourcing. ChatGPT, Copilot and Perplexity performed notably better, though all four showed substantial error rates. The study was published in October 2025, and model behaviour changes over time.

There is no correction desk, takedown route or submission endpoint on any major engine. The only available mechanism is indirect: make the accurate source easier to find and attribute than the wrong one, correct the third-party records, and wait for retrieval to reflect it. No correction can be guaranteed.

That is entity resolution failure rather than a content problem. If a machine cannot resolve which organisation a name refers to, publishing more content from the ambiguous entity adds noise. The fix is a canonical entity home, a stable identifier, consistent naming, and `sameAs` links to external records.

Thin coverage. Fabrication concentrates on questions where no good source exists — pricing, limits, what the product is not for — because the generation step has nothing to retrieve. Publishing a checkable, dated source for the true version is more effective than disputing the output.

Add accuracy questions to the tracked set — what the company does, what it costs, where it is based, who runs it — score the answer text rather than the mention, run it per engine and per language, repeat enough that one wrong answer is not mistaken for a pattern, and trace each error to the source URL that carried it.

Sources and further reading

  1. European Broadcasting Union and BBC, News Integrity in AI Assistants, October 2025
  2. European Commission, AI Act transparency obligations for general-purpose AI
  3. Press Gazette, publisher AI deals and lawsuits tracker

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team