Short answer. In the largest published study of answer reliability, professional journalists found significant issues in 45% of AI assistant answers about news, and serious sourcing problems in 31% — missing, misleading or incorrect attribution (European Broadcasting Union and BBC, October 2025). There is no reason to think answers about your company are handled more carefully, and no engine offers a takedown route or a ticket queue for a factual error.
An answer engine that gets your pricing wrong fails the same way one that gets an election wrong fails: nothing in the pipeline distinguishes a brand fact from a news fact. This piece covers the failure rate, the four ways an answer goes wrong about a company, why correction is not straightforward, what a monitoring programme that actually catches it looks like, and the regulatory backdrop.
Key takeaways
- The EBU/BBC study evaluated over 3,000 AI assistant responses across 22 public service media organisations, 18 countries and 14 languages, and found significant issues in 45% of answers.
- Gemini showed significant issues in 76% of responses in that study, more than double the other three assistants tested.
- No major AI answer engine offers a correction desk, takedown route or submission endpoint for a factual error about a company.
- Brand-answer errors fall into four distinct types — stale, misattributed, conflated and fabricated — and each needs a different fix.
- The European Commission's transparency obligations for general-purpose AI under the EU AI Act became fully enforceable on 2 August 2026.
How often do assistants get it wrong?
The strongest published measurement of AI-assistant accuracy comes from the journalism organisations with the most experience checking factual claims, and it found errors in roughly half of tested answers.
| Finding | Rate |
|---|---|
| Answers with at least one significant issue | 45% |
| Answers with serious sourcing problems — missing, misleading or incorrect attribution | 31% |
| Answers with major accuracy issues, including hallucinated details and outdated information | 20% |
| Gemini, significant issues — more than double the other assistants, largely due to poor sourcing | 76% |
Source: European Broadcasting Union and BBC, October 2025 — over 3,000 responses evaluated by professional journalists at 22 public service media organisations across 18 countries and 14 languages, testing ChatGPT, Copilot, Gemini and Perplexity.
Two details make this transferable to brand claims. The failures were consistent across languages and territories, so it is not an artefact of one market. And the dominant failure mode was sourcing rather than invention — the assistant attached a claim to the wrong source, or to no source, more often than it fabricated the claim outright.
The caveat is worth stating plainly: this study measured news, not brands. No equivalent study of brand-fact accuracy exists at this scale, and model behaviour has moved since October 2025. It remains the closest rigorous proxy available, because the underlying failure mechanisms — thin sourcing, entity confusion, stale citations — are shared between news and brand claims, a pattern explored further in how AI search engines decide which brands to mention and cite.
What are the four ways an answer goes wrong about you?
Each of the four failure types has a different root cause and a different fix, and conflating them is why most "AI reputation" work goes nowhere.
| Failure | What it looks like | Root cause | What actually addresses it |
|---|---|---|---|
| Stale | Old pricing, discontinued product, former executive | The cited source is out of date, or the model's memory predates the change | Update or retire the source page; correct the third-party record |
| Misattributed | A true fact credited to the wrong company, or your fact credited elsewhere | Sourcing failure — the 31% category above | Make the correct source unambiguous and easier to attribute than the wrong one |
| Conflated | Answers merge you with a similarly named organisation | Entity resolution failure | Entity identity work: a canonical home, stable identifiers, consistent naming |
| Fabricated | A claim that appears in no source | Generation failure, most common where sources are thin | Publish a checkable source for the true version; thin coverage invites invention |
Entity resolution is the process by which a machine decides which real-world organisation a name refers to. Conflation is not fixed by publishing more content — publishing more from an unresolvable entity adds noise rather than clarity, a problem covered in more depth in structured data and entity identity for AEO. Fabrication concentrates where coverage is thin, which means the fix is supplying a source, not disputing the output.
Why can't you simply have it corrected?
There is no correction desk on any major answer engine, so the only lever available is changing what the retrieval layer finds and waiting for it to take effect.
No major engine offers a takedown route for a factual error about a company, a submission endpoint, or a ticket queue. That indirect mechanism is slow for three structural reasons.
- Retrieval mode can update in days to weeks once a source is fixed. Memory mode — facts baked into the model during training — only changes when a model is retrained, so a correction published today may not reach it this quarter regardless of spend.
- A large share of AI citations point at third-party sources rather than a company's own site, a dynamic explored in why third-party brand mentions matter for GEO and AI search. If the error originates in a directory entry, an old press release or a community thread, correcting your own site does not touch it.
- Cited sources churn heavily from one run to the next. An error can disappear and reappear without anything having actually been fixed, which makes it easy to mistake churn for a correction.
What would a monitoring programme that catches this look like?
A programme that actually catches brand-answer errors tracks the accuracy of the answer text itself, not just whether the brand was named.
- Include accuracy questions in the tracked set, not just visibility questions — "What does [company] do?", "How much does [product] cost?", "Where is [company] based?", "Who founded [company]?" A visibility metric answers whether you were named; it does not answer whether what was said is true, a distinction covered in how to measure AI visibility without fooling yourself.
- Score the answer text, not just the mention — correct, incomplete, misattributed, conflated, fabricated. A mention-rate dashboard scores a wrong answer as a success.
- Run it per engine and per language. The EBU/BBC failures were consistent across 14 languages, and Gemini failed at more than twice the rate of the other assistants tested; an averaged score hides exactly the engine that needs action.
- Repeat enough to distinguish an error from a one-off. Given how much citation sets churn, one wrong answer is a single sample; a claim appearing across a third of runs is a pattern.
- Trace every error to its source, meaning the specific cited URL that carried the wrong claim, because without that step correction is guesswork.
- Log the raw answer text, since demonstrating that an error was corrected requires the before as well as the after.
The cheapest high-value addition to any tracked set is simply "What does [company] do?", run monthly across every engine — it catches conflation, staleness and fabrication in one line, and almost nobody tracks it because it is not a visibility metric.
How do you reduce the exposure you carry?
A company cannot control the generation step of an answer engine, but it can reduce how much guessing the retrieval step has to do.
- Make the true version easy to find, easy to attribute and clearly dated. Sourcing failure is the largest error category above, so an unambiguous, current, well-structured source is the direct countermeasure, as covered in what actually gets you cited by AI answer engines.
- Resolve the entity: one canonical home, a stable identifier, consistent naming, and
sameAslinks to external records. Conflation is an identity failure, not a content failure. - Fix stale third-party records first. Directories, databases, encyclopedic references and old coverage outlive the facts they contain, and are frequently cited more often than a company's own pages.
- Publish the awkward facts — pricing, limits, what the product is not for. Where a company leaves a gap, retrieval fills it from somewhere less accurate, or invents it.
- Date everything, machine-readably, since staleness is only detectable if the source states when it was true.
- Keep retrieval crawlers unblocked. A site that cannot be fetched cannot correct the record, and default crawler-blocking rules on some hosting platforms do not distinguish AI training bots from the retrieval bots that answer engines use to check facts in real time.
None of this compels a specific output. Every countermeasure changes the odds on a system nobody outside the engine operator controls, and memory-mode errors in particular are not addressable on a business timeline.
What does the regulatory backdrop change?
Regulation and litigation are moving, but neither currently gives a company a direct route to correct one wrong answer.
The European Commission's transparency obligations for general-purpose AI under the EU AI Act became fully enforceable on 2 August 2026. Separately, Press Gazette's tracker of publisher deals and lawsuits against AI companies has recorded a growing list of suits against Perplexity over alleged copyright or content-use disputes, brought by organisations including News Corp, The New York Times, Encyclopedia Britannica and Yomiuri Shimbun.
What these two developments indicate is a direction of travel: obligations on providers are increasing, and organisations with resources are using litigation where no product mechanism exists. For most companies neither is a practical remedy today, which leaves an internal monitoring capability as the more reliable option in the meantime.
How does Lifewood approach brand-accuracy monitoring?
Lifewood runs accuracy questions inside the same tracked prompt sets used for visibility tracking, scored on the answer text rather than the mention alone.
Runs are scored per engine and per language, with raw output retained so a corrected error can be demonstrated against its before. Errors are classified as stale, misattributed, conflated or fabricated before any remediation work is commissioned, because those four types route to entirely different fixes, and two of them are not content problems at all.
The remediation work itself is mostly outside a client's own site — third-party records, entity identity, and publishing checkable sources for facts that retrieval is currently inventing. Across 100+ languages and 40+ delivery centres across 30+ countries, that monitoring is run in-market, because an answer that reads as accurate in English is routinely wrong in another language, a gap covered further in AEO in the markets Google does not own. Programmes like this sit within Lifewood's AEO and GEO services.