Skip to main content
AEO/GEO

AI Visibility Audits: How to Measure Your Brand in ChatGPT and Gemini

August 2026 · 6 min read · Updated September 2026

Short answer. An AI visibility audit measures how often, where and how accurately a brand appears across a controlled set of buyer-relevant prompts. A good audit does not rely on a few screenshots. It defines a prompt universe, runs comparable tests across ChatGPT, Gemini and other relevant platforms, records mentions and citations, benchmarks competitors, identifies the sources influencing answers, reviews brand accuracy and sentiment, and repeats the process over time to separate durable change from normal model variability.

Key takeaways

  • Define the buyer questions that matter before running any test.
  • Group prompts by funnel stage and intent, not by brand-favorable topics.
  • Benchmark direct competitors on the same prompts, engines and time window.
  • Track brand mentions and recommendation presence separately from citations.
  • Review brand accuracy, context and sentiment, not sentiment alone.
  • Analyze which source domains repeatedly influence answers.
  • Repeat the same prompt set on a schedule and compare trends.

What should an AI visibility audit answer?

A complete audit answers whether the brand appears for the questions buyers actually ask, how it compares with named competitors, whether mentions turn into citations, whether the description is accurate, and whether visibility is improving over time.

  • Does the brand appear for the questions buyers actually ask?
  • Which competitors appear more often?
  • Is the brand directly cited, or only discussed through third-party sources?
  • Is the brand described accurately?
  • Which topics or buyer stages are strongest or weakest?
  • Which websites are repeatedly used to support recommendations?
  • Is visibility improving over time?

How do you build a representative prompt set?

The prompt set is the audit's foundation, and it should be designed around customer decisions rather than prompts the brand is already likely to win.

Prompt group Examples
Category discovery Best providers for X; leading tools for Y
Problem discovery How do companies solve X?
Use case Best X for healthcare / enterprise / multilingual teams
Comparison A vs B; alternatives to A
Trust Most secure providers; reputable companies
Pricing How much does X cost?
Implementation How to choose / deploy X
Brand-specific What is Brand A? Is Brand A good for X?

How should competitor benchmarking work?

Competitor benchmarking uses the same prompts, engines and time window for every competitor and records the context of each appearance, not just whether it happened.

Competitor benchmarking is the practice of running an identical prompt set against rival brands so mention, recommendation and citation rates are directly comparable. Count whether competitors are recommended first, cited as a source, included only as an alternative, or described inaccurately.

Competitor metric Why it matters
Mention rate Basic visibility
Recommendation rate Commercial consideration
Median shortlist position Relative prominence
Citation share Source authority
Accuracy Quality of representation
Topic coverage Where competitor is strong or weak

What is the difference between a mention and a citation?

A mention means the brand appears in the generated answer, while a citation means a source is explicitly referenced or linked.

A mention is any appearance of the brand name inside an AI-generated answer, with or without a link. A citation is an explicit reference to a source, usually a link, that the model used to support the answer. The two are tracked separately because a brand can be recommended without its own website being cited, and a brand-owned page can be cited for information without the brand being recommended.

Bing's 2026 AI Performance dashboard makes this distinction concrete by reporting citations and cited pages without claiming they represent ranking or placement within an answer. This is consistent with how ChatGPT decides which sources to cite and why citation counts should never be read as a ranking signal on their own.

How do you calculate AI share of voice?

AI share of voice is the percentage of total tracked brand mentions captured by each brand within the same prompt set, and it should always be paired with recommendation and citation metrics.

Share of voice is easy to understand but treats every mention as equal, which it is not. Reading it alongside a dedicated share of answer metric gives a fuller picture of whether a mention actually drove the recommendation. For volatile engines, repeated runs on a sample of prompts help estimate how much variation is normal before a change is treated as a trend.

How should sentiment and accuracy be reviewed?

Sentiment alone can be misleading, so accuracy review pairs a sentiment label with a structured check of whether the brand is described correctly.

A more useful review asks whether the brand is described correctly, whether important limitations are represented fairly, and whether outdated claims are being repeated.

Review dimension Example label
Category accuracy Correct / partly correct / wrong
Feature accuracy Current / outdated / unsupported
Tone Positive / neutral / negative
Recommendation context Best fit / alternative / warning
Citation support Strong / weak / none

What does source analysis reveal?

Source analysis identifies the domains that repeatedly appear behind answers, which can explain why competitors are visible in a given category.

For example, one category may be driven by software review sites, another by analyst research, and another by independent editorial comparisons. A thorough pass reviews the top cited domains by prompt group, the owned-versus-third-party citation mix, sources that mention competitors but not the brand, outdated pages that appear repeatedly, and sources that describe the brand incorrectly. Understanding why third-party brand mentions matter for GEO helps prioritize which of those domains to pursue first.

How often should measurement repeat?

Monthly measurement is usually frequent enough to detect trends without overreacting to day-to-day model variability.

High-change categories may justify weekly checks on a smaller core prompt set. The key is consistency: the same prompt taxonomy, a documented engine list, and a comparable methodology each time, which is also how GEO KPIs beyond clicks and rankings should be tracked over a program's life.

What should an audit report include?

An audit report should move from an executive summary through methodology, results and benchmarking to a concrete plan for what to measure next.

Section Deliverable
Executive summary Key visibility and competitor findings
Prompt methodology Prompt list, grouping and run date
Engine results Mentions, citations and recommendation rates
Competitor benchmark Share of voice by topic
Source map Domains influencing answers
Accuracy review Incorrect or outdated descriptions
Opportunity map Content, authority and technical gaps
Measurement plan What to track next and how often

Teams comparing audit and monitoring vendors can review the AI visibility tools that measure these signals before committing to a single provider, and can read more on AEO and GEO methodology generally.

Frequently asked questions

Enough to cover the main buyer journeys. Small programs may start with 30-50 prompts; enterprise programs often need 100 or more segmented by category, market or audience.

Some should, but the most important discovery prompts are usually unbranded, since those are the ones a buyer types before they know which brands exist.

No. It establishes a baseline. Recurring measurement is needed to understand direction and durability, since a single run cannot separate a real change from normal model variability.

Sometimes. AI referral traffic can be measured when links generate visits, but many AI interactions are zero-click, so brand and pipeline impact may require broader attribution methods.

Sources and further reading

  1. Bing Webmaster Blog - AI Performance in Bing Webmaster Tools
  2. OpenAI - Searching the web with ChatGPT
  3. OpenAI - Publishers and Developers FAQ
  4. Google Search Central - AI optimization guide
  5. Princeton / KDD - GEO: Generative Engine Optimization

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team