Skip to main content
AEO/GEO

How to Measure AI Visibility Without Fooling Yourself

August 2026 · 9 min read · Updated September 2026

Short answer. Ask an answer engine the same question two days running and most of its sources change. In the Parse study of 693,509 answers, ChatGPT repeated only 21.2% of its cited domains on a repeat ask and Google AI Overviews 31.5%. A measurement that survives that churn is a rate estimated from repeated runs, reported per engine and per market, with retrieval and memory modes kept apart — never a position, never a single blended score, and never a screenshot.

Key takeaways

  • ChatGPT repeats only 21.2% of its cited domains and Google AI Overviews 31.5% when the same question is asked twice, per a Parse study of 693,509 answers (26 March–25 April 2026).
  • Day-over-day source churn reaches 88.3% on Gemini, 79.2% on ChatGPT, 75.9% on Google AI Mode and 44.4% on Perplexity, and 84% of a question's sources appear on only one engine, per GetMentions (June 2026).
  • A defensible measurement is a rate from repeated runs across a frozen question set, reported per engine and per market with retrieval and memory kept separate — never a rank position or a single blended score.
  • In an unsettled AI-answer category, a modest, consistent mention rate can already be a leading position, since Semrush's tracking of 1,094 US categories found only 15.2% had a clear ChatGPT owner.
  • Control questions in an adjacent category you do not serve are the cheapest way to tell a genuine visibility change from the underlying models shifting.

How noisy is the system, really?

Two 2026 studies, run by different teams with different methods in different months, agree closely on how unstable AI-generated answers are from one ask to the next.

Measurement Finding Source
Cited domains repeated on a second ask ChatGPT 21.2%, Google AI Overviews 31.5% Parse, 693,509 answers across 16,143 ChatGPT and 15,805 Google prompts, 26 March – 25 April 2026
Same, widened to a one-week window ChatGPT 26.7%, Google AI Overviews 36.8% Parse, as above
Day-over-day source churn Gemini 88.3%, ChatGPT 79.2%, Google AI Mode 75.9%, Perplexity 44.4% GetMentions, 530,875 citations across 181,225 URLs and 2,398 queries, seven consecutive days, June 2026
Sources cited on all seven consecutive days Perplexity 11.1%, Google AI Mode 2.6%, ChatGPT 1.1%, Gemini 0.4% GetMentions, as above
Sources cited by only one of the four engines 84% GetMentions, as above

Three numbers carry the argument: roughly 79% of ChatGPT's sources change overnight, around 1% survive a week, and 84% of a question's sources appear on only one engine. A screenshot of an AI answer is one draw from a very fat-tailed distribution, not evidence of anything durable.

What does the noise floor rule out?

It rules out any measurement built on a single observation or a single number, because both are indistinguishable from noise at these churn rates.

  • The one-off audit. "We asked ChatGPT ten questions and you appeared twice." At these churn rates that is compatible with almost any underlying visibility rate.
  • Position language. "You rank third for this prompt." With roughly 1% seven-day source persistence, position is not a property the system has.
  • Week-on-week movement charts from small prompt sets. Movement appears every week whether or not anything changed, and a chart that always moves cannot detect change.
  • A single blended AI visibility score. With 84% of sources unique to one engine, averaging deletes the only actionable structure in the data.
  • Attributing a single week's change to a single action. Causal claims need a control, and most programmes have none.

What is share of answer?

Share of answer is how often a brand appears across a defined set of AI-generated answers, relative to the total opportunity or to competitors. There is no standardised industry formula, which makes documenting your own version the load-bearing step.

Share of answer = Answers naming the brand ÷ Total answers across the fixed prompt set
Cited share     = Answers linking or attributing the brand ÷ Answers naming the brand

The simplest version is the percentage of tracked prompts where the brand is mentioned; weighted versions add prominence, position within the answer, or whether a link was given. Consistency matters more than picking the perfect formula: hold the prompt set, engines, geography, language and scoring rules stable enough that two periods are comparable at all. Mention rate alone is not sufficient — track alongside it the owned-domain citation rate, the third-party citation rate, competitor mentions, the pages actually cited, and a qualitative field for whether the description was accurate, incomplete or wrong, since a mention-rate dashboard scores a confidently wrong answer as a success.

What does a defensible measurement look like?

A defensible measurement treats visibility as a sampling problem, which is a solved problem, rather than reaching for a better tool.

  1. Fix the question set before you start and freeze the wording. A rephrased question is a new series, not a continuation. If the text can be edited later, every trend line is unfalsifiable.
  2. Write questions as buyers ask them. "Who can run AI visibility across Japanese and Korean?" is a question. "multilingual AEO services" is a keyword, and nobody types that into an assistant.
  3. Separate retrieval from memory and never average them. With search on, the engine cites URLs and responds to what you publish within days to weeks. With search off, it answers from training data and changes only when a new model ships. Blending them makes a quarter of correct retrieval work look like failure.
  4. Run repeatedly and report a rate. The unit is "named in 34% of runs across 40 questions over four weeks", not "ranked third". Rates are estimable under noise; positions are not.
  5. Measure each engine and each market separately. Engine by market is a matrix, and one number is not a summary of it.
  6. Include control questions you do not intend to win — an adjacent category you do not serve. If you start appearing there, that is mispositioning, not progress, and it is the cheapest way to tell a real movement from the models changing underneath your benchmark.
  7. Record mentions separately from links. Being described accurately without a link still shapes the buyer's view, and on memory-mode answers it is the only outcome available.
  8. Paraphrase deliberately. Repeats estimate stability; paraphrases show whether visibility survives a change of wording.
  9. Log the raw answers. When a number moves, the only way to explain it is to read what changed.

How many runs is enough?

There is no universal number, but precision improves with the square root of the observation count, and an observation is one question, one run, one engine.

Design Observations per engine What it supports
10 questions checked once 10 A demonstration. Not enough to distinguish 20% visibility from 40% at the observed churn
40 questions, twice weekly, four weeks 320 Enough to see a genuine step-change and ignore ordinary week-to-week noise

Sampling cost is not equal across engines: at 44.4% daily churn against Gemini's 88.3% (GetMentions, June 2026), the same confidence costs far less sampling on Perplexity. An honest first report reads: "Across 40 buyer questions run eight times on four engines, we were named in 6% of ChatGPT runs, 0% of Gemini runs, 11% of Perplexity runs and 2% of Google AI Mode runs. Confidence is lowest on Gemini, because its churn is highest." That is actionable; "you are invisible in AI search" is not.

What does a good number look like in your category?

A good number depends entirely on whether anyone owns the category yet, so the benchmark is the best-performing competitor rather than an absolute target. Semrush, with Kevin Indig for Growth Memo, tracked 1,094 US categories in ChatGPT between January and June 2026: 15.2% had a clear owner, 31.2% an emerging leader and 53.7% were unsettled, and clear owners held the top spot in 90.4% of month-over-month comparisons. In an unsettled category — most are — a modest, consistent rate is a leading position. Where a category has an owner, displacement is slow and second-source presence is the more realistic goal.

How do you tell a real measurement from a demo?

A real measurement can answer a follow-up question about sample size and mode; a demo cannot.

Claim you will hear Why it fails What to ask instead
"You rank #3 for this prompt" Answers have no stable positions; around 1% of sources survive a week "What is the rate across repeated runs?"
"Your AI visibility score is 42" Blends engines that mostly do not share sources "Show me the score per engine, per market"
"We saw a 12% lift this week" Inside the noise floor of a small prompt set "How many observations, and what is the confidence?"
"We can guarantee citations" No engine offers submission or placement "What outcome are you contractually promising?"
"You appear in 0% of AI answers" Usually a single run, often memory mode "Which mode, how many runs, which questions?"
"We optimised and citations rose" No control questions, no counterfactual "What did the control set do in that period?"

What should a weekly or monthly report show?

A useful report shows the overall share of answer and its change from the prior period, alongside enough raw evidence to audit the summary rather than trust it. That includes the strongest and weakest topic clusters, new and lost citations, competitor gains, any answers that described the brand incorrectly, and the actions planned next.

Separate movement caused by your own work from platform volatility wherever you can — a major engine update shifts visibility across many brands at once, and will otherwise be read as a result of your programme. Keep visibility framed as an intermediate metric: where analytics allow, track referral traffic, lead source and sales conversations that mention an AI recommendation, without asserting a causal line from mention rate to revenue that the data does not support. Measurement has hard limits: it cannot make the system deterministic, cannot fix memory-mode absence on a reporting cycle, and attribution to revenue stays hard since AI referrals often arrive un-attributed. Every figure above carries a date because benchmarks age fast.

How does Lifewood approach this?

Lifewood runs AI visibility as a measured programme rather than a set of recommendations, on client sites and on its own. The instrument is in-house: fixed prompt sets with frozen wording, a pre-work baseline, retrieval and memory reported separately, control questions in every set, and raw run files retained so a number that moves can be explained rather than guessed at.

For multi-market programmes, the constraint is authorship rather than tooling. Lifewood delivers across 100+ languages from 40+ delivery centres across 30+ countries, so prompt sets are written in-market rather than translated — the question a buyer asks in Vietnamese is rarely the English question rendered in Vietnamese. That authorship pipeline is one output of the wider multilingual content pipeline built for answer engines, and it connects to the broader discipline of answer engine optimization and generative engine optimization, of which share-of-answer tracking is one part.

Measurement on its own only tells you where you stand; closing the gap is a separate, ongoing effort covered in what actually gets you cited by AI answer engines and in the comparison of GEO, SEO and AEO. Programmes that skip the measurement step tend to reach for one of the 7 reasons AI isn't citing a brand without evidence for which one applies, and teams evaluating vendors for this work often start from a buyer's checklist for AEO and GEO services.

Frequently asked questions

Fix a set of buyer questions, freeze the wording, run them repeatedly against each engine in both retrieval and memory modes, and report the share of runs in which the brand is named or cited. The output is a rate with a confidence range per engine and per market, not a position and not a single blended score.

No. There is no universal industry formula, so the scoring method has to be documented and then held constant. A simple version is the percentage of tracked prompts where the brand is mentioned; weighted versions add prominence, position in the answer, or whether a link was given.

Because they sample different questions at different times from a system with roughly 79% day-to-day source churn on ChatGPT, where 84% of a question's sources appear on only one engine (GetMentions, June 2026). Two honest tools with different prompt sets will legitimately produce different numbers.

Enough that the churn averages out. Ten prompts checked once is a demonstration. Roughly 40 buyer questions run twice weekly for a month gives a few hundred observations per engine — enough to separate a real step-change from ordinary noise.

Always. Retrieval mode responds to what you publish within days to weeks and can cite a URL. Memory mode answers from training data and changes only when a new model ships. Averaging them makes months of correct retrieval work look like failure.

Control questions are queries in an adjacent category you deliberately do not intend to win. If you begin appearing for them, that indicates mispositioning rather than progress. They are also the cheapest way to distinguish a genuine movement from the models changing beneath your benchmark.

On retrieval surfaces, published changes can register within days to weeks, but proving a change against this noise floor takes repeated runs across several weeks. On the memory surface, nothing you publish moves the number until a new model is trained.

Sources and further reading

  1. Parse — AI citation volatility by industry
  2. GetMentions — AI Citation Volatility: A 530,875-Citation Study
  3. Semrush with Kevin Indig, Growth Memo — AI visibility is a topic-level game

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team