Short answer. Ask an answer engine the same question two days running and most of its sources change. In the Parse study of 693,509 answers, ChatGPT repeated only 21.2% of its cited domains on a repeat ask and Google AI Overviews 31.5%. A measurement that survives that churn is a rate estimated from repeated runs, reported per engine and per market, with retrieval and memory modes kept apart — never a position, never a single blended score, and never a screenshot.
Key takeaways
- ChatGPT repeats only 21.2% of its cited domains and Google AI Overviews 31.5% when the same question is asked twice, per a Parse study of 693,509 answers (26 March–25 April 2026).
- Day-over-day source churn reaches 88.3% on Gemini, 79.2% on ChatGPT, 75.9% on Google AI Mode and 44.4% on Perplexity, and 84% of a question's sources appear on only one engine, per GetMentions (June 2026).
- A defensible measurement is a rate from repeated runs across a frozen question set, reported per engine and per market with retrieval and memory kept separate — never a rank position or a single blended score.
- In an unsettled AI-answer category, a modest, consistent mention rate can already be a leading position, since Semrush's tracking of 1,094 US categories found only 15.2% had a clear ChatGPT owner.
- Control questions in an adjacent category you do not serve are the cheapest way to tell a genuine visibility change from the underlying models shifting.
How noisy is the system, really?
Two 2026 studies, run by different teams with different methods in different months, agree closely on how unstable AI-generated answers are from one ask to the next.
| Measurement | Finding | Source |
|---|---|---|
| Cited domains repeated on a second ask | ChatGPT 21.2%, Google AI Overviews 31.5% | Parse, 693,509 answers across 16,143 ChatGPT and 15,805 Google prompts, 26 March – 25 April 2026 |
| Same, widened to a one-week window | ChatGPT 26.7%, Google AI Overviews 36.8% | Parse, as above |
| Day-over-day source churn | Gemini 88.3%, ChatGPT 79.2%, Google AI Mode 75.9%, Perplexity 44.4% | GetMentions, 530,875 citations across 181,225 URLs and 2,398 queries, seven consecutive days, June 2026 |
| Sources cited on all seven consecutive days | Perplexity 11.1%, Google AI Mode 2.6%, ChatGPT 1.1%, Gemini 0.4% | GetMentions, as above |
| Sources cited by only one of the four engines | 84% | GetMentions, as above |
Three numbers carry the argument: roughly 79% of ChatGPT's sources change overnight, around 1% survive a week, and 84% of a question's sources appear on only one engine. A screenshot of an AI answer is one draw from a very fat-tailed distribution, not evidence of anything durable.
What does the noise floor rule out?
It rules out any measurement built on a single observation or a single number, because both are indistinguishable from noise at these churn rates.
- The one-off audit. "We asked ChatGPT ten questions and you appeared twice." At these churn rates that is compatible with almost any underlying visibility rate.
- Position language. "You rank third for this prompt." With roughly 1% seven-day source persistence, position is not a property the system has.
- Week-on-week movement charts from small prompt sets. Movement appears every week whether or not anything changed, and a chart that always moves cannot detect change.
- A single blended AI visibility score. With 84% of sources unique to one engine, averaging deletes the only actionable structure in the data.
- Attributing a single week's change to a single action. Causal claims need a control, and most programmes have none.
What does a defensible measurement look like?
A defensible measurement treats visibility as a sampling problem, which is a solved problem, rather than reaching for a better tool.
- Fix the question set before you start and freeze the wording. A rephrased question is a new series, not a continuation. If the text can be edited later, every trend line is unfalsifiable.
- Write questions as buyers ask them. "Who can run AI visibility across Japanese and Korean?" is a question. "multilingual AEO services" is a keyword, and nobody types that into an assistant.
- Separate retrieval from memory and never average them. With search on, the engine cites URLs and responds to what you publish within days to weeks. With search off, it answers from training data and changes only when a new model ships. Blending them makes a quarter of correct retrieval work look like failure.
- Run repeatedly and report a rate. The unit is "named in 34% of runs across 40 questions over four weeks", not "ranked third". Rates are estimable under noise; positions are not.
- Measure each engine and each market separately. Engine by market is a matrix, and one number is not a summary of it.
- Include control questions you do not intend to win — an adjacent category you do not serve. If you start appearing there, that is mispositioning, not progress, and it is the cheapest way to tell a real movement from the models changing underneath your benchmark.
- Record mentions separately from links. Being described accurately without a link still shapes the buyer's view, and on memory-mode answers it is the only outcome available.
- Paraphrase deliberately. Repeats estimate stability; paraphrases show whether visibility survives a change of wording.
- Log the raw answers. When a number moves, the only way to explain it is to read what changed.
How many runs is enough?
There is no universal number, but precision improves with the square root of the observation count, and an observation is one question, one run, one engine.
| Design | Observations per engine | What it supports |
|---|---|---|
| 10 questions checked once | 10 | A demonstration. Not enough to distinguish 20% visibility from 40% at the observed churn |
| 40 questions, twice weekly, four weeks | 320 | Enough to see a genuine step-change and ignore ordinary week-to-week noise |
Sampling cost is not equal across engines: at 44.4% daily churn against Gemini's 88.3% (GetMentions, June 2026), the same confidence costs far less sampling on Perplexity. An honest first report reads: "Across 40 buyer questions run eight times on four engines, we were named in 6% of ChatGPT runs, 0% of Gemini runs, 11% of Perplexity runs and 2% of Google AI Mode runs. Confidence is lowest on Gemini, because its churn is highest." That is actionable; "you are invisible in AI search" is not.
What does a good number look like in your category?
A good number depends entirely on whether anyone owns the category yet, so the benchmark is the best-performing competitor rather than an absolute target. Semrush, with Kevin Indig for Growth Memo, tracked 1,094 US categories in ChatGPT between January and June 2026: 15.2% had a clear owner, 31.2% an emerging leader and 53.7% were unsettled, and clear owners held the top spot in 90.4% of month-over-month comparisons. In an unsettled category — most are — a modest, consistent rate is a leading position. Where a category has an owner, displacement is slow and second-source presence is the more realistic goal.
How do you tell a real measurement from a demo?
A real measurement can answer a follow-up question about sample size and mode; a demo cannot.
| Claim you will hear | Why it fails | What to ask instead |
|---|---|---|
| "You rank #3 for this prompt" | Answers have no stable positions; around 1% of sources survive a week | "What is the rate across repeated runs?" |
| "Your AI visibility score is 42" | Blends engines that mostly do not share sources | "Show me the score per engine, per market" |
| "We saw a 12% lift this week" | Inside the noise floor of a small prompt set | "How many observations, and what is the confidence?" |
| "We can guarantee citations" | No engine offers submission or placement | "What outcome are you contractually promising?" |
| "You appear in 0% of AI answers" | Usually a single run, often memory mode | "Which mode, how many runs, which questions?" |
| "We optimised and citations rose" | No control questions, no counterfactual | "What did the control set do in that period?" |
What should a weekly or monthly report show?
A useful report shows the overall share of answer and its change from the prior period, alongside enough raw evidence to audit the summary rather than trust it. That includes the strongest and weakest topic clusters, new and lost citations, competitor gains, any answers that described the brand incorrectly, and the actions planned next.
Separate movement caused by your own work from platform volatility wherever you can — a major engine update shifts visibility across many brands at once, and will otherwise be read as a result of your programme. Keep visibility framed as an intermediate metric: where analytics allow, track referral traffic, lead source and sales conversations that mention an AI recommendation, without asserting a causal line from mention rate to revenue that the data does not support. Measurement has hard limits: it cannot make the system deterministic, cannot fix memory-mode absence on a reporting cycle, and attribution to revenue stays hard since AI referrals often arrive un-attributed. Every figure above carries a date because benchmarks age fast.
How does Lifewood approach this?
Lifewood runs AI visibility as a measured programme rather than a set of recommendations, on client sites and on its own. The instrument is in-house: fixed prompt sets with frozen wording, a pre-work baseline, retrieval and memory reported separately, control questions in every set, and raw run files retained so a number that moves can be explained rather than guessed at.
For multi-market programmes, the constraint is authorship rather than tooling. Lifewood delivers across 100+ languages from 40+ delivery centres across 30+ countries, so prompt sets are written in-market rather than translated — the question a buyer asks in Vietnamese is rarely the English question rendered in Vietnamese. That authorship pipeline is one output of the wider multilingual content pipeline built for answer engines, and it connects to the broader discipline of answer engine optimization and generative engine optimization, of which share-of-answer tracking is one part.
Measurement on its own only tells you where you stand; closing the gap is a separate, ongoing effort covered in what actually gets you cited by AI answer engines and in the comparison of GEO, SEO and AEO. Programmes that skip the measurement step tend to reach for one of the 7 reasons AI isn't citing a brand without evidence for which one applies, and teams evaluating vendors for this work often start from a buyer's checklist for AEO and GEO services.