Short answer. Every tool in this category samples a probabilistic system with roughly 79% day-to-day source churn, using a prompt list whose composition can move the reported number by more than 16 percentage points. That is not a criticism of the tools — it is the specification they work against, and it determines which of their outputs you can trust. Mention rates, cited-domain distributions, competitor co-occurrence and per-engine breakdowns are measurable. Position, causation, accuracy and revenue attribution are not, whatever the interface implies.
Buying an AI visibility tool is the easy decision in this category, which is why it is usually made first. At entry pricing the subscription is not the expensive part of a programme. The expensive part is deciding what to ask and knowing what the answer means, and no dashboard does either.
This piece separates what these products can genuinely measure from what they present as measurement, and sets out what to require before signing.
What are you actually buying?
The AI-visibility tooling market in 2026 includes Profound, Peec AI, AthenaHQ, Bluefish, Scrunch, Adobe LLM Optimizer and Semrush's AI Visibility Toolkit, with entry pricing between roughly $99 and $295 per month. That list comes from Scrunch's own 2026 comparison, which is worth reading with the conflict visible: Scrunch is one of the tools compared, and it ranks itself first.
Underneath the interfaces, almost all of them do the same four things:
- Run a list of prompts against a set of engines on a schedule, through APIs or automation.
- Parse the returned answers for brand mentions and cited URLs.
- Aggregate the results into rates, shares and trend lines.
- Attribute changes to content, competitors or time.
Steps one to three are engineering problems, and most vendors solve them adequately. Step four is where the claims outrun the evidence, because attributing a change requires a control and almost no product has one.
A tool can tell you what an engine said. It cannot tell you why, and very few are honest about the difference.
What can these tools measure well?
| Capability | Why it works | What to require |
|---|---|---|
| Mention rate over repeated runs | It is a proportion, and proportions are estimable under noise given enough samples | Sample size per engine, per period |
| Cited-domain distribution | Directly observed from the answer text | The full domain list, not a top-five summary |
| Competitor co-occurrence | Observed in the same answers, so genuinely comparative | Which brands appeared alongside yours, per prompt |
| Per-engine breakdown | Engines differ enormously and can be measured separately | Never accept a blended score |
| Answer text archive | Raw evidence, and the only way to explain a movement | Exportable raw answers, not screenshots |
Every item on that list is an observation. None of them requires the product to know why anything happened.
What can they not measure, whatever the interface implies?
- Position. Answers do not have stable ranks. In GetMentions' seven-day study, only 1.1% of ChatGPT's cited sources persisted across all seven consecutive days. A tool showing "rank 3" has invented an ordering the underlying system does not have.
- Causation. Without a control set of questions you do not intend to win, a rise is indistinguishable from the models changing underneath the benchmark.
- Memory-mode presence, as a URL. With search off, an engine cites nothing. A product reporting "citations" from a non-retrieval answer is reporting mentions and should say so.
- Cross-engine truth. In the same study, 84% of the sources for a question were cited by only one of the four engines measured. An averaged score describes no engine that exists.
- Whether the answer was accurate. Mention-rate dashboards score a confidently wrong sentence about your pricing as a win.
- Revenue attribution. AI referrals are roughly 0.29% of search referrals and frequently arrive with no referrer at all.
The volatility figures behind those limits come from two large independent studies. GetMentions measured 530,875 citations across 181,225 URLs and 2,398 queries over seven consecutive days in June 2026, finding day-over-day source churn of 88.3% on Gemini, 79.2% on ChatGPT, 75.9% on Google AI Mode and 44.4% on Perplexity. Parse, analysing 693,509 answers between 26 March and 25 April 2026, found that asking the same question twice returned only 21.2% of the same cited domains on ChatGPT and 31.5% on Google AI Overviews.
Why do two tools report different numbers for the same brand?
Mostly the prompt list. This is the largest source of variation between vendors reporting on the same brand, and it is almost never visible in the interface.
Analyze, measuring 22,295 AI answers and 115,843 citation events across 460 B2B prompts and 37 organisations, found mention rate varied by prompt archetype from 41.2% for recommendation prompts on Perplexity, to 34.7% for comparison prompts on Google AI Mode, to 24.5% for research prompts on ChatGPT — a spread of 16.7 percentage points. Within engines the spread was still 12.2 points on Perplexity, 9.3 on Google AI Mode and 8.5 on ChatGPT.
Phrasing compounds it. Ehrlinspiel, Landwehr and Rudzki at Peec AI, working across 37,804 AI responses from 1,754 prompts on five engines and published 10 June 2026, found keyword-style prompts produced up to 25% higher average brand visibility than conversational phrasing, ranking-style prompts about 20% higher, and prompt length effectively no effect at all. Prompts drifting below roughly 0.50 cosine similarity lost about half their observed visibility.
Two vendors, both honest and both competent, can report numbers sixteen points apart for the same brand purely from prompt-mix decisions. Any comparison of tool outputs that does not hold the prompt list constant is comparing prompt lists, not visibility.
What should you require before buying?
- The full prompt list, exportable, with archetype labels. If you cannot see the questions, you cannot interpret the number.
- Frozen wording, with an audit trail when a prompt changes. A silently edited prompt breaks the series without breaking the chart.
- Sample size and run cadence per engine, shown in the interface. Against 79% churn, a rate without an N is decoration.
- Per-engine and per-market reporting as the default view, with any blended score available only as a secondary.
- Retrieval and memory modes separated and labelled.
- Raw answer export. When a number moves, the text is the only explanation available.
- Support for control prompts in a category you do not intend to win.
- A stated position on accuracy. Ask whether the product scores whether the sentence about you was correct. Most do not.
One disqualifier is quicker than all eight: ask the vendor what their tool would show if you changed nothing for three months. The correct answer is a rate fluctuating around a stable mean, with confidence intervals. If the answer implies a smooth line, the product is smoothing noise into a story.
Should you build, buy, or both?
| Buy a tool | Build in-house | Hybrid | |
|---|---|---|---|
| Best for | Standard engines, English, fast start | Regional engines, local languages, research needs | Most enterprises |
| Cost shape | $99–$295/month entry, plus interpretation labour | Engineering plus ongoing maintenance | Tool for the common case, in-house for the gaps |
| Main risk | A prompt list you did not design; blended scores | Underestimating maintenance and drift | Two sources of truth, if the prompt lists differ |
| Coverage gap | Regional engines, local languages, accuracy scoring | None inherent | Whatever the tool misses |
For most organisations the honest answer is hybrid, with one rule attached: one prompt registry, used by both. Two lists produce two numbers and a permanent argument about which is real.
Limits of this assessment
- The tools list is compiled by a vendor in the same market that places itself first. Read it as a starting point rather than a ranking.
- Nothing here is a product review. Capabilities in this category change quarterly, and any specific claim would be stale before it was useful.
- The volatility figures come from two studies, both large, both independent of each other, both from mid-2026.
- No tool can deliver a guaranteed citation, because no engine sells one.
How Lifewood approaches this
Lifewood runs its own instrument rather than a purchased dashboard, for one reason that matters commercially: the prompt registry has to be authored per market rather than translated, and no tool in the category writes questions in the language a buyer actually asks them in. The registry is frozen at the start of a series, versioned when it changes, and reported per engine and per market with the retrieval and memory surfaces kept apart.
Raw answers are retained for every run, because a rate tells you something moved and only the text tells you why. Control questions in an adjacent category run alongside the real set, so a movement can be checked against the noise floor rather than asserted against it.
Where a client already owns a tool, the sensible arrangement is usually hybrid: the tool covers the standard engines in English, in-house measurement covers regional engines, local languages and accuracy scoring, and both run from the same registry. See what to look for in AEO and GEO services and how to measure AI visibility without fooling yourself.
Sources and further reading
- Scrunch, AEO and GEO tools comparison 2026 — market composition and entry pricing. Published by one of the tools compared.
- GetMentions, AI citation volatility: a 530,875-citation study, June 2026 — churn, seven-day persistence, and the 84% single-engine figure.
- Parse, AI citation volatility by industry, 693,509 answers, March–April 2026.
- Analyze, State of AI search: prompt archetypes, 22,295 answers across 460 B2B prompts and 37 organisations.
- Ehrlinspiel, Landwehr & Rudzki (Peec AI), prompt variance study, SSRN, 10 June 2026, via Search Engine Journal.
- Technology Checker, search engine market share, August 2026 update — AI referral share.

