Short answer. Ask an answer engine the same question two days running and most of its sources change. In the Parse study of 693,509 answers, ChatGPT repeated only 21.2% of its cited domains on a repeat ask and Google AI Overviews 31.5%. A measurement that survives that churn is a rate estimated from repeated runs, reported per engine and per market, with retrieval and memory modes kept apart — never a position, never a single blended score, and never a screenshot.
Most of what is sold as AI visibility reporting is a demonstration dressed as a baseline. The screenshot, the one-off audit, the week-on-week movement chart: each reads noise as signal, because the system is far noisier than the reports assume.
This piece sets out the noise floor, what it rules out, what a defensible measurement looks like, how many runs is enough, and how to tell a real measurement from a demo.
How noisy is the system, really?
Two 2026 studies — different teams, different methods, different months — agree closely.
| Measurement | Finding | Source |
|---|---|---|
| Cited domains repeated on a second ask | ChatGPT 21.2%, Google AI Overviews 31.5% | Parse, 693,509 answers across 16,143 ChatGPT and 15,805 Google prompts, 26 March – 25 April 2026 |
| Same, widened to a one-week window | ChatGPT 26.7%, Google AI Overviews 36.8% | Parse, as above |
| Day-over-day source churn | Gemini 88.3%, ChatGPT 79.2%, Google AI Mode 75.9%, Perplexity 44.4% | GetMentions, 530,875 citations across 181,225 URLs and 2,398 queries, seven consecutive days, June 2026 |
| Sources cited on all seven consecutive days | Perplexity 11.1%, Google AI Mode 2.6%, ChatGPT 1.1%, Gemini 0.4% | GetMentions, as above |
| Sources cited by only one of the four engines | 84% | GetMentions, as above |
Three numbers carry the argument: roughly 79% of ChatGPT's sources change overnight, around 1% survive a week, and 84% of a question's sources appear on only one engine.
A screenshot of an AI answer is not evidence. It is one draw from a very fat-tailed distribution.
What does that rule out?
The one-off audit. "We asked ChatGPT ten questions and you appeared twice." At these churn rates that is compatible with almost any underlying visibility rate.
Position language. "You rank third for this prompt." With roughly 1% seven-day source persistence, position is not a property the system has.
Week-on-week movement charts from small prompt sets. Movement appears every week whether or not anything changed, and a chart that always moves cannot detect change.
A single blended AI visibility score. With 84% of sources unique to one engine, averaging deletes the only actionable structure in the data.
Attributing a single week's change to a single action. Causal claims need a control, and most programmes have none.
What is share of answer?
Share of answer is how often a brand appears across a defined set of AI-generated answers, relative to the total opportunity or to competitors. There is no standardised industry formula, which makes documenting your own version the load-bearing step.
Share of answer = Answers naming the brand ÷ Total answers across the fixed prompt set
Cited share = Answers linking or attributing the brand ÷ Answers naming the brand
The simplest version is the percentage of tracked prompts where the brand is mentioned; weighted versions add prominence, position within the answer, or whether a link was given. Consistency matters more than picking the perfect formula: hold the prompt set, engines, geography, language and scoring rules stable enough that two periods are comparable at all.
Mention rate alone is not sufficient. Track alongside it the owned-domain citation rate, the third-party citation rate, competitor mentions, the pages actually cited, and a qualitative field for whether the description was accurate, incomplete or wrong. A mention-rate dashboard scores a confidently wrong answer as a success.
What does a defensible measurement look like?
The fix is not a better tool. It is treating this as sampling, which is a solved problem.
- Fix the question set before you start and freeze the wording. A rephrased question is a new series, not a continuation. If the text can be edited later, every trend line is unfalsifiable.
- Write questions as buyers ask them. "Who can run AI visibility across Japanese and Korean?" is a question. "multilingual AEO services" is a keyword, and nobody types that into an assistant.
- Separate retrieval from memory and never average them. With search on, the engine cites URLs and responds to what you publish within days to weeks. With search off, it answers from training data and changes only when a new model ships. Blending them makes a quarter of correct retrieval work look like failure.
- Run repeatedly and report a rate. The unit is "named in 34% of runs across 40 questions over four weeks", not "ranked third". Rates are estimable under noise; positions are not.
- Measure each engine and each market separately. Engine by market is a matrix, and one number is not a summary of it.
- Include control questions you do not intend to win — an adjacent category you do not serve. If you start appearing there, that is mispositioning, not progress, and it is the cheapest way to tell a real movement from the models changing underneath your benchmark.
- Record mentions separately from links. Being described accurately without a link still shapes the buyer's view, and on memory-mode answers it is the only outcome available.
- Paraphrase deliberately. Repeats estimate stability; paraphrases show whether visibility survives a change of wording.
- Log the raw answers. When a number moves, the only way to explain it is to read what changed.
How many runs is enough?
There is no universal number, but the shape of the answer is straightforward. You are estimating a proportion under high variance, so precision improves with the square root of the observation count — and an observation is one question, one run, one engine.
| Design | Observations per engine | What it supports |
|---|---|---|
| 10 questions checked once | 10 | A demonstration. Not enough to distinguish 20% visibility from 40% at the observed churn |
| 40 questions, twice weekly, four weeks | 320 | Enough to see a genuine step-change and ignore ordinary week-to-week noise |
Sampling cost is not equal across engines: at 44.4% daily churn against Gemini's 88.3% (GetMentions, June 2026), the same confidence costs far less sampling on Perplexity.
An honest first report reads: "Across 40 buyer questions run eight times on four engines, we were named in 6% of ChatGPT runs, 0% of Gemini runs, 11% of Perplexity runs and 2% of Google AI Mode runs. Confidence is lowest on Gemini, because its churn is highest." That is actionable. "You are invisible in AI search" is not.
What does a good number look like in your category?
Absolute figures mean little without knowing whether anyone owns the category at all. Semrush, with Kevin Indig for Growth Memo, tracked 1,094 US categories in ChatGPT between January and June 2026: 15.2% had a clear owner, 31.2% an emerging leader and 53.7% were unsettled, and clear owners held the top spot in 90.4% of month-over-month comparisons. So in an unsettled category — most are — a modest, consistent rate is a leading position, and the benchmark is the best-performing competitor rather than an absolute target. Where a category has an owner, displacement is slow and second-source presence is the more realistic goal.
How do you tell a real measurement from a demo?
| Claim you will hear | Why it fails | What to ask instead |
|---|---|---|
| "You rank #3 for this prompt" | Answers have no stable positions; around 1% of sources survive a week | "What is the rate across repeated runs?" |
| "Your AI visibility score is 42" | Blends engines that mostly do not share sources | "Show me the score per engine, per market" |
| "We saw a 12% lift this week" | Inside the noise floor of a small prompt set | "How many observations, and what is the confidence?" |
| "We can guarantee citations" | No engine offers submission or placement | "What outcome are you contractually promising?" |
| "You appear in 0% of AI answers" | Usually a single run, often memory mode | "Which mode, how many runs, which questions?" |
| "We optimised and citations rose" | No control questions, no counterfactual | "What did the control set do in that period?" |
What should a weekly or monthly report show?
Overall share of answer and the change from the prior period; strongest and weakest topic clusters; new and lost citations; competitor gains; any answers that described the brand incorrectly; and the actions planned next. Include prompt-level raw evidence so a stakeholder can audit the summary rather than trust it.
Separate movement caused by your work from platform volatility wherever you can — a major engine update shifts visibility across many brands at once, and will otherwise be read as a result. Keep visibility framed as an intermediate metric: where analytics allow, track referral traffic, lead source and sales conversations that mention an AI recommendation, without asserting a causal line from mention rate to revenue that the data does not support.
Measurement has hard limits. It cannot make the system deterministic — you can estimate a rate precisely, but not reproduce a specific answer. It cannot fix memory-mode absence on a reporting cycle. Attribution to revenue stays hard, since AI referrals often arrive un-attributed. And benchmarks age fast, which is why every figure above carries a date.
How Lifewood approaches this
Lifewood runs AI visibility as a measured programme rather than a set of recommendations, on client sites and on its own. The instrument is in-house: fixed prompt sets with frozen wording, a pre-work baseline, retrieval and memory reported separately, control questions in every set, and raw run files retained so a number that moves can be explained rather than guessed at. For multi-market programmes the constraint is authorship rather than tooling: 50+ languages and 40+ delivery centres across 30+ countries mean prompt sets are written in-market rather than translated, because the question a buyer asks in Vietnamese is rarely the English question rendered in Vietnamese.
See AEO services, GEO services, what gets you cited by AI answer engines and GEO vs AEO vs traditional SEO.
Sources and further reading
- Parse, AI citation volatility by industry — 693,509 answers, 26 March – 25 April 2026.
- GetMentions, AI citation volatility — 530,875 citations, 2,398 queries, June 2026.
- Semrush with Kevin Indig, Growth Memo, AI visibility is a topic-level game, January–June 2026.

