Short answer. An AI visibility audit measures how often, where and how accurately a brand appears across a controlled set of buyer-relevant prompts. A good audit does not rely on a few screenshots. It defines a prompt universe, runs comparable tests across ChatGPT, Gemini and other relevant platforms, records mentions and citations, benchmarks competitors, identifies the sources influencing answers, reviews brand accuracy and sentiment, and repeats the process over time to separate durable change from normal model variability.
What should an AI visibility audit answer?
Does the brand appear for the questions buyers actually ask?
Which competitors appear more often?
Is the brand directly cited, or only discussed through third-party sources?
Is the brand described accurately?
Which topics or buyer stages are strongest or weakest?
Which websites are repeatedly used to support recommendations?
Is visibility improving over time?
How do you build a representative prompt set?
The prompt set is the audit's foundation. It should be designed around customer decisions, not around prompts the brand is already likely to win.
- Prompt group
- Examples
- Category discovery
- Best providers for X; leading tools for Y
- Problem discovery
How do companies solve X?
- Use case
- Best X for healthcare / enterprise / multilingual teams
- Comparison
- A vs B; alternatives to A
- Trust
- Most secure providers; reputable companies
- Pricing
How much does X cost?
Implementation
How to choose / deploy X
Brand-specific
What is Brand A? Is Brand A good for X?
How should competitor benchmarking work?
Use the same prompts, engines and time window for every competitor. Count not only whether competitors appear, but the context of the appearance: are they recommended first, cited as a source, included only as an alternative or described inaccurately?
| Competitor metric | Why it matters |
|---|---|
| Mention rate | Basic visibility |
| Recommendation rate | Commercial consideration |
| Median shortlist position | Relative prominence |
| Citation share | Source authority |
| Accuracy | Quality of representation |
| Topic coverage | Where competitor is strong or weak |
What is the difference between a mention and a citation?
A mention means the brand appears in the generated answer. A citation means a source is explicitly referenced or linked. The two should be tracked separately because a brand can be recommended without its own website being cited, and a brand-owned page can be cited for information without the brand being recommended.
Bing's 2026 AI Performance dashboard makes this distinction concrete by reporting citations and cited pages without claiming they represent ranking or placement within an answer. Bing AI Performance
How do you calculate AI share of voice?
A simple share-of-voice metric is the percentage of total tracked brand mentions captured by each brand within the same prompt set. It is easy to understand, but should be paired with recommendation and citation metrics because not all mentions are equally valuable.
For volatile engines, repeated runs on a sample of prompts can help estimate how much variation is normal.
How should sentiment and accuracy be reviewed?
Sentiment alone can be misleading. More useful is a structured qualitative review: is the brand described correctly, are important limitations represented fairly and are outdated claims being repeated?
- Review dimension
- Example label
- Category accuracy
- Correct / partly correct / wrong
- Feature accuracy
- Current / outdated / unsupported
- Tone
- Positive / neutral / negative
- Recommendation context
- Best fit / alternative / warning
- Citation support
- Strong / weak / none
What does source analysis reveal?
Source analysis identifies the domains that repeatedly appear behind answers. Those domains can reveal why competitors are visible. For example, one category may be driven by software review sites, another by analyst research, and another by independent editorial comparisons.
Top cited domains by prompt group.
Owned versus third-party citation mix.
Sources mentioning competitors but not the brand.
Outdated pages that appear repeatedly.
Sources that describe the brand incorrectly.
How often should measurement repeat?
For an active GEO program, monthly measurement is usually frequent enough to detect trends without overreacting to day-to-day variability. High-change categories may justify weekly checks on a smaller core prompt set. The key is consistency: same prompt taxonomy, documented engine and comparable methodology.
What should an audit report include?
- Section
- Deliverable
- Executive summary
- Key visibility and competitor findings
- Prompt methodology
- Prompt list, grouping and run date
- Engine results
- Mentions, citations and recommendation rates
- Competitor benchmark
- Share of voice by topic
- Source map
- Domains influencing answers
- Accuracy review
- Incorrect or outdated descriptions
- Opportunity map
- Content, authority and technical gaps
- Measurement plan
- What to track next and how often
Key takeaways
- Define the buyer questions that matter.
- Group prompts by funnel stage and intent.
- Benchmark direct competitors.
- Track brand mentions and recommendation presence.
- Record owned and third-party citations.
- Review brand accuracy, context and sentiment.
- Analyze which source domains repeatedly influence answers.
- Repeat the same prompt set on a schedule and compare trends.
Sources and further reading
- Bing Webmaster Blog - AI Performance in Bing Webmaster Tools.
- OpenAI - Searching the web with ChatGPT.
- OpenAI - Publishers and Developers FAQ.
- Google Search Central - AI optimization guide.
- Princeton / KDD - GEO: Generative Engine Optimization.