Skip to main content
AEO/GEO

AI Visibility Audits: How to Measure Your Brand in ChatGPT and Gemini

Short answer. An AI visibility audit measures how often, where and how accurately a brand appears across a controlled set of buyer-relevant prompts. A good audit does not rely on a few…

Kelvin T. · August 2026 · 4 min read

Download PDF

Short answer. An AI visibility audit measures how often, where and how accurately a brand appears across a controlled set of buyer-relevant prompts. A good audit does not rely on a few screenshots. It defines a prompt universe, runs comparable tests across ChatGPT, Gemini and other relevant platforms, records mentions and citations, benchmarks competitors, identifies the sources influencing answers, reviews brand accuracy and sentiment, and repeats the process over time to separate durable change from normal model variability.


What should an AI visibility audit answer?

Does the brand appear for the questions buyers actually ask?

Which competitors appear more often?

Is the brand directly cited, or only discussed through third-party sources?

Is the brand described accurately?

Which topics or buyer stages are strongest or weakest?

Which websites are repeatedly used to support recommendations?

Is visibility improving over time?


How do you build a representative prompt set?

The prompt set is the audit's foundation. It should be designed around customer decisions, not around prompts the brand is already likely to win.

  • Prompt group
  • Examples
  • Category discovery
  • Best providers for X; leading tools for Y
  • Problem discovery

How do companies solve X?

  • Use case
  • Best X for healthcare / enterprise / multilingual teams
  • Comparison
  • A vs B; alternatives to A
  • Trust
  • Most secure providers; reputable companies
  • Pricing

How much does X cost?

Implementation

How to choose / deploy X

Brand-specific

What is Brand A? Is Brand A good for X?


How should competitor benchmarking work?

Use the same prompts, engines and time window for every competitor. Count not only whether competitors appear, but the context of the appearance: are they recommended first, cited as a source, included only as an alternative or described inaccurately?

Competitor metric Why it matters
Mention rate Basic visibility
Recommendation rate Commercial consideration
Median shortlist position Relative prominence
Citation share Source authority
Accuracy Quality of representation
Topic coverage Where competitor is strong or weak

What is the difference between a mention and a citation?

A mention means the brand appears in the generated answer. A citation means a source is explicitly referenced or linked. The two should be tracked separately because a brand can be recommended without its own website being cited, and a brand-owned page can be cited for information without the brand being recommended.

Bing's 2026 AI Performance dashboard makes this distinction concrete by reporting citations and cited pages without claiming they represent ranking or placement within an answer. Bing AI Performance


How do you calculate AI share of voice?

A simple share-of-voice metric is the percentage of total tracked brand mentions captured by each brand within the same prompt set. It is easy to understand, but should be paired with recommendation and citation metrics because not all mentions are equally valuable.

For volatile engines, repeated runs on a sample of prompts can help estimate how much variation is normal.


How should sentiment and accuracy be reviewed?

Sentiment alone can be misleading. More useful is a structured qualitative review: is the brand described correctly, are important limitations represented fairly and are outdated claims being repeated?

  • Review dimension
  • Example label
  • Category accuracy
  • Correct / partly correct / wrong
  • Feature accuracy
  • Current / outdated / unsupported
  • Tone
  • Positive / neutral / negative
  • Recommendation context
  • Best fit / alternative / warning
  • Citation support
  • Strong / weak / none

What does source analysis reveal?

Source analysis identifies the domains that repeatedly appear behind answers. Those domains can reveal why competitors are visible. For example, one category may be driven by software review sites, another by analyst research, and another by independent editorial comparisons.

Top cited domains by prompt group.

Owned versus third-party citation mix.

Sources mentioning competitors but not the brand.

Outdated pages that appear repeatedly.

Sources that describe the brand incorrectly.


How often should measurement repeat?

For an active GEO program, monthly measurement is usually frequent enough to detect trends without overreacting to day-to-day variability. High-change categories may justify weekly checks on a smaller core prompt set. The key is consistency: same prompt taxonomy, documented engine and comparable methodology.


What should an audit report include?

  • Section
  • Deliverable
  • Executive summary
  • Key visibility and competitor findings
  • Prompt methodology
  • Prompt list, grouping and run date
  • Engine results
  • Mentions, citations and recommendation rates
  • Competitor benchmark
  • Share of voice by topic
  • Source map
  • Domains influencing answers
  • Accuracy review
  • Incorrect or outdated descriptions
  • Opportunity map
  • Content, authority and technical gaps
  • Measurement plan
  • What to track next and how often

Key takeaways

  • Define the buyer questions that matter.
  • Group prompts by funnel stage and intent.
  • Benchmark direct competitors.
  • Track brand mentions and recommendation presence.
  • Record owned and third-party citations.
  • Review brand accuracy, context and sentiment.
  • Analyze which source domains repeatedly influence answers.
  • Repeat the same prompt set on a schedule and compare trends.

Sources and further reading

    1. Bing Webmaster Blog - AI Performance in Bing Webmaster Tools.
    1. OpenAI - Searching the web with ChatGPT.
    1. OpenAI - Publishers and Developers FAQ.
    1. Google Search Central - AI optimization guide.
    1. Princeton / KDD - GEO: Generative Engine Optimization.

Frequently asked questions

Enough to cover the main buyer journeys. Small programs may start with 30-50 prompts; enterprise programs often need 100 or more segmented by category, market or audience.

Some should, but the most important discovery prompts are usually unbranded.

No. It establishes a baseline. Recurring measurement is needed to understand direction and durability.

Sometimes. AI referral traffic can be measured when links generate visits, but many AI interactions are zero-click, so brand and pipeline impact may require broader attribution.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team