Short answer. AI visibility tools can measure mention rate over repeated runs, the distribution of cited domains, competitor co-occurrence and per-engine differences, because all four are observations. They cannot measure position, causation, answer accuracy or revenue attribution, whatever the interface implies. Every tool samples a probabilistic system with roughly 79% day-to-day source churn on ChatGPT, using a prompt list whose composition alone can move the reported number by nearly 17 percentage points.
Key takeaways
- AI visibility tools reliably measure mention rate, cited-domain distribution, competitor co-occurrence, per-engine breakdowns and raw answer archives, because each is directly observed in the answer text.
- No AI visibility tool can measure rank, causation, answer accuracy or revenue attribution, because answers have no stable order, almost no product runs a control set, and AI referrals were 0.29% of search referrals in May 2026.
- ChatGPT changed 79.2% of its cited sources from one day to the next in a 530,875-citation study, and only 1.1% of its sources persisted across all seven consecutive days.
- Prompt archetype alone shifted mention rate by 16.7 percentage points across engines in a 22,295-answer B2B study, so two honest tools with different prompt lists will report different numbers for the same brand.
- Before buying, require an exportable prompt list, frozen wording, visible sample sizes per engine, per-market reporting, separated retrieval and memory modes, raw answer export and support for control prompts.
What are you actually buying when you buy an AI visibility tool?
You are buying a scheduler that runs a fixed list of prompts against a set of engines, parses the answers for brand mentions and cited URLs, and aggregates the results into rates and trend lines. The subscription is the cheap part of the programme; deciding what to ask and knowing what the answer means is the expensive part, and no dashboard does either.
An AI visibility tool is a software product that repeatedly submits a list of prompts to AI answer engines and reports how often, and alongside which sources, a brand appears in the returned answers.
The AI-visibility tooling market in 2026 includes Profound, Peec AI, AthenaHQ, Bluefish, Scrunch, Adobe LLM Optimizer and Semrush's AI Visibility Toolkit, with published entry pricing between roughly $99 and $295 per month for the tools that disclose it. That list comes from Scrunch's own 2026 comparison, which is worth reading with the conflict visible: Scrunch is one of the tools compared, and it ranks itself first. Bluefish and Adobe LLM Optimizer do not publish standard pricing.
Underneath the interfaces, almost all of them do the same four things:
- Run a list of prompts against a set of engines on a schedule, through APIs or automation.
- Parse the returned answers for brand mentions and cited URLs.
- Aggregate the results into rates, shares and trend lines.
- Attribute changes to content, competitors or time.
Steps one to three are engineering problems, and most vendors solve them adequately. Step four is where the claims outrun the evidence, because attributing a change requires a control and almost no product has one. The first 90 days of an AI visibility programme are mostly spent designing the questions the tool will run.
A tool can tell you what an engine said. It cannot tell you why, and very few are honest about the difference.
What can these tools measure well?
These tools measure five things well: mention rate over repeated runs, cited-domain distribution, competitor co-occurrence, per-engine breakdowns and the raw answer text. Every one of them is an observation, and none requires the product to know why anything happened.
Mention rate is the proportion of answers to a fixed prompt set, over a fixed period, in which a named brand appears. Turning it into a defensible share of answer metric is a separate question from whether the tool can observe it.
| Capability | Why it works | What to require |
|---|---|---|
| Mention rate over repeated runs | It is a proportion, and proportions are estimable under noise given enough samples | Sample size per engine, per period |
| Cited-domain distribution | Directly observed from the answer text | The full domain list, not a top-five summary |
| Competitor co-occurrence | Observed in the same answers, so genuinely comparative | Which brands appeared alongside yours, per prompt |
| Per-engine breakdown | Engines differ enormously and can be measured separately | Never accept a blended score |
| Answer text archive | Raw evidence, and the only way to explain a movement | Exportable raw answers, not screenshots |
What can they not measure, whatever the interface implies?
They cannot measure position, causation, memory-mode citations, cross-engine truth, answer accuracy or revenue attribution. Each requires something the underlying system does not provide: a stable ordering, a control group, a retrieval step, a shared source pool, a fact-check or a reliable referrer.
Source churn is the share of an engine's cited sources for a prompt that changes between one run and the next. The volatility figures behind the limits below come from two large studies that are independent of each other, both from mid-2026.
- Position. Answers do not have stable ranks. In GetMentions' seven-day study, only 1.1% of ChatGPT's cited sources persisted across all seven consecutive days. A tool showing "rank 3" has invented an ordering the underlying system does not have.
- Causation. Without a control set of questions you do not intend to win, a rise is indistinguishable from the models changing underneath the benchmark.
- Memory-mode presence, as a URL. With search off, an engine cites nothing. A product reporting "citations" from a non-retrieval answer is reporting mentions and should say so.
- Cross-engine truth. In the same study, 84% of the sources for a question were cited by only one of the four engines measured. An averaged score describes no engine that exists.
- Whether the answer was accurate. Mention-rate dashboards score a confidently wrong sentence about your pricing as a win.
- Revenue attribution. All AI chatbots combined sent 0.29% of search referrals in May 2026, according to Cloudflare Radar data, and in one April 2026 sample of 371,847 sessions, 35.7% of AI traffic arrived with no referrer at all.
GetMentions measured 530,875 citations across 181,225 URLs and 2,398 queries over seven consecutive days in June 2026, finding day-over-day source churn of 88.3% on Gemini, 79.2% on ChatGPT, 75.9% on Google AI Mode and 44.4% on Perplexity. Parse, analysing 693,509 answers between 26 March and 25 April 2026, found that asking the same question twice returned only 21.2% of the same cited domains on ChatGPT and 31.5% on Google AI Overviews. A single check therefore tells you almost nothing; measuring AI visibility without fooling yourself covers how many runs a series needs.
Why do two tools report different numbers for the same brand?
Mostly the prompt list. Prompt composition is the largest source of variation between vendors reporting on the same brand, and it is almost never visible in the interface.
Analyze, measuring 22,295 AI answers and 115,843 citation events across 460 B2B prompts and 37 organisations, found mention rate varied by prompt archetype from 41.2% for recommendation prompts on Perplexity, to 34.7% for comparison prompts on Google AI Mode, to 24.5% for research prompts on ChatGPT, a spread of 16.7 percentage points. Within engines the spread was still 12.2 points on Perplexity, 9.3 on Google AI Mode and 8.5 on ChatGPT.
Phrasing compounds it. Ehrlinspiel, Landwehr and Rudzki at Peec AI, working across 37,804 AI responses from 1,754 prompts on five engines in a paper published on SSRN on 10 June 2026, found keyword-style prompts produced up to 25% higher average brand visibility than conversational phrasing, ranking-style prompts about 20% higher, and prompt length effectively no effect at all. Visibility held steady while prompt variants stayed above roughly 0.50 cosine similarity to the original; variants drifting well below that threshold lost about half their observed visibility.
Two vendors, both honest and both competent, can report numbers sixteen points apart for the same brand purely from prompt-mix decisions. Any comparison of tool outputs that does not hold the prompt list constant is comparing prompt lists, not visibility.
What should you require before buying?
Require the full prompt list, frozen wording with an audit trail, visible sample sizes per engine, per-engine and per-market reporting by default, separated retrieval and memory modes, raw answer export, support for control prompts and a stated position on accuracy. A vendor that cannot supply the first three cannot supply an interpretable number.
- The full prompt list, exportable, with archetype labels. If you cannot see the questions, you cannot interpret the number.
- Frozen wording, with an audit trail when a prompt changes. A silently edited prompt breaks the series without breaking the chart.
- Sample size and run cadence per engine, shown in the interface. Against 79% churn, a rate without an N is decoration.
- Per-engine and per-market reporting as the default view, with any blended score available only as a secondary.
- Retrieval and memory modes separated and labelled.
- Raw answer export. When a number moves, the text is the only explanation available.
- Support for control prompts in a category you do not intend to win.
- A stated position on accuracy. Ask whether the product scores whether the sentence about you was correct. Most do not.
One disqualifier is quicker than all eight: ask the vendor what their tool would show if you changed nothing for three months. The correct answer is a rate fluctuating around a stable mean, with confidence intervals. If the answer implies a smooth line, the product is smoothing noise into a story. The same test belongs among the questions to ask before hiring AEO and GEO help.
Should you build, buy, or both?
For most organisations the honest answer is hybrid: buy a tool for the standard engines in English, build in-house measurement for regional engines, local languages and accuracy scoring. The one rule that makes a hybrid work is a single prompt registry used by both sides.
A prompt registry is the versioned, frozen list of questions, with their archetype and market labels, that every measurement run in a programme draws from. Two lists produce two numbers and a permanent argument about which is real.
| Option | Buy a tool | Build in-house | Hybrid |
|---|---|---|---|
| Best for | Standard engines, English, fast start | Regional engines, local languages, research needs | Most enterprises |
| Cost shape | $99 to $295 per month entry, plus interpretation labour | Engineering plus ongoing maintenance | Tool for the common case, in-house for the gaps |
| Main risk | A prompt list you did not design; blended scores | Underestimating maintenance and drift | Two sources of truth, if the prompt lists differ |
| Coverage gap | Regional engines, local languages, accuracy scoring | None inherent | Whatever the tool misses |
The same trade-off appears one level up, in the choice between in-house GEO and a GEO agency.
What are the limits of this assessment?
This assessment is a description of a category, not a review of any product, and its market list comes from a vendor inside that market. Capabilities in this category change quarterly, and any specific product claim would be stale before it was useful.
- The tools list is compiled by a vendor in the same market that places itself first. Read it as a starting point rather than a ranking.
- Nothing here is a product review. Capabilities in this category change quarterly.
- The volatility figures come from two studies, GetMentions and Parse, both large, both independent of each other, both from mid-2026.
- The prompt-mix figures come from two further studies, Analyze and Peec AI, and Peec AI is itself one of the tools in the category.
- No tool can deliver a guaranteed citation, because no engine sells one.
How does Lifewood approach AI visibility measurement?
Lifewood runs its own instrument rather than a purchased dashboard, because the prompt registry has to be authored per market rather than translated, and no tool in the category writes questions in the language a buyer actually asks them in. The registry is frozen at the start of a series, versioned when it changes, and reported per engine and per market with the retrieval and memory surfaces kept apart.
Raw answers are retained for every run, because a rate tells you something moved and only the text tells you why. Control questions in an adjacent category run alongside the real set, so a movement can be checked against the noise floor rather than asserted against it. Measurement of this kind is the first stage of Lifewood's answer engine optimisation service, and the same registry drives the optimisation work that follows.
Where a client already owns a tool, the sensible arrangement is usually hybrid: the tool covers the standard engines in English, in-house measurement covers regional engines, local languages and accuracy scoring, and both run from the same registry. The criteria for that arrangement are set out in what to look for in AEO and GEO services, and buyers comparing measurement-plus-optimisation vendors can start with the AEO and GEO provider comparison.