LIFEWOOD
Finalizing099
AI search

How ChatGPT Decides Which Sources to Cite

Short answer. ChatGPT answers from two different places: a live retrieval index when search is on, and trained memory when it is off. Only the first can cite a URL. When it does cite, the…

Lifewood Data Technology · August 2026 · 8 min read

Download PDF

Short answer. ChatGPT answers from two different places: a live retrieval index when search is on, and trained memory when it is off. Only the first can cite a URL. When it does cite, the Semrush AI Visibility Index 2026 — built on 126 million US AI search prompts — puts it at about 15 sources per answer, against roughly 3 for Gemini. And it repeats very few of them: asked the same question twice, Parse's study of 693,509 answers found ChatGPT reused only 21.2% of its cited domains. Visibility here is a rate, never a position.

Every citation, on every engine, is the end of the same two-stage process: a document has to be retrieved before it can be selected. Wording a page for citability changes the second stage and does nothing for the first, which is why so much AI visibility work produces no measurable movement — it is aimed at a stage the page never reached.

ChatGPT makes that distinction unusually visible, because it runs two answering modes with completely different retrieval behaviour. Confusing them is the single most common source of nonsense in AI visibility reporting.


Two modes, two different questions

Search on (retrieval) Search off (memory)
What happens Runs searches, reads pages, then answers Answers from training data
Can it cite a URL? Yes No — it has no document to point at
What moves it What is published and crawlable now What the wider web said before the training cut-off
Response time Days to weeks Only when a new model ships
Who can influence it You, fairly directly Mostly other people writing about you

If a report says your ChatGPT visibility is 0% and does not say which mode it measured, it has told you nothing. Being absent from retrieval is a publishing and crawlability problem you can fix this quarter. Being absent from memory is a reputation problem measured in model generations.


The retrieval pipeline, step by step

With search on, ChatGPT is a retrieval-augmented system, and every stage is a place you can be excluded.

  1. A crawler builds the index. OpenAI runs separate bots for separate jobs: OAI-SearchBot builds the search index ChatGPT answers from, GPTBot collects training data, and ChatGPT-User fetches a page live when a request requires it. Blocking the wrong one costs you the wrong thing — see AI crawlers and AI search visibility.
  2. The question becomes searches. The model rewrites the user's question into one or more queries. Small changes in phrasing produce different queries, and therefore different sources.
  3. Documents are retrieved and reranked. Candidates come back and are scored for relevance to the rewritten queries. This is the stage where passage-level writing decisions pay off or do not.
  4. The answer is composed with citations attached. The model writes from the documents in its context window and links the ones it used.

Two things follow from the shape of that pipeline. First, position in the retrieved context matters as well as content quality — the model sees a subset, not the web. Second, being used and being cited are not the same outcome. Research on citation absorption, notably Zhang, He and Yao's work distinguishing citation selection from citation absorption, finds that a retrieved source can contribute language, evidence or structure to an answer while a different source receives the visible link. Citation count is therefore an incomplete measure of influence, though it remains the only one that can be observed from outside.


Fifteen slots is a structural advantage

The Semrush AI Visibility Index 2026, drawn from 126 million US AI search prompts recorded between January and April 2026, found ChatGPT cites an average of about 15 sources per response while Gemini cites about 3.

That five-fold difference in citation slots is worth planning around. Fifteen slots means a mid-sized brand can realistically appear alongside the incumbents in the same answer. Three slots means winning is close to winner-take-all. The same content programme has very different odds on the two engines, which is one reason a blended "AI visibility score" across engines is not an actionable number.


Four in five sources change overnight

This is the most important and least-reported fact about ChatGPT visibility, and two independent studies agree closely on it.

Parse, analysing 693,509 answers across 16,143 ChatGPT prompts and 15,805 Google prompts between 26 March and 25 April 2026, found that asking the same question twice returned only 21.2% of the same cited domains on ChatGPT and 31.5% on Google AI Overviews. Widening the window to a week raised overlap only to 26.7% and 36.8%.

GetMentions, working from 530,875 citations across 181,225 distinct URLs and 2,398 real queries over seven consecutive days in June 2026, measured the same instability from the other direction:

Engine Day-over-day source churn Sources cited on all seven days
Gemini 88.3% 0.4%
ChatGPT 79.2% 1.1%
Google AI Mode 75.9% 2.6%
Perplexity 44.4% 11.1%

The same study found 84% of the sources for a question were cited by only one of the four engines.

One screenshot is not a measurement. It is one sample of a noisy distribution. Three consequences follow, all commercial rather than technical:

  • A single check proves nothing in either direction. Not appearing once is not absence; appearing once is not visibility.
  • A week-on-week change from a handful of prompts is noise. At roughly 79% daily churn you need repeated runs across a fixed question set before a movement means anything.
  • Engines must be measured separately. With 84% of sources unique to one engine, an averaged score hides the only thing you could act on. Perplexity in particular behaves nothing like ChatGPT.

What actually raises the odds

Given that volatility, the goal is not to hold a slot. It is to be in the candidate pool often enough that the dice land your way at a useful rate.

  • Be fetchable by the right bot. OAI-SearchBot decides whether ChatGPT can cite you at all, and many sites block it without knowing, because a platform default did it for them.
  • Be quotable at passage level. Self-contained paragraphs, question-shaped headings, one claim per paragraph, the answer in the first sentence.
  • Carry evidence, not adjectives. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande's GEO benchmark (ACM SIGKDD 2024), covering roughly 10,000 queries across nine datasets, found targeted content changes raised visibility in generative engine responses by up to 40%, with authority-style edits — adding citations, statistics and quotations — outperforming cosmetic edits such as rewriting, simplification and keyword work. Be careful with second-hand versions of that paper: per-tactic percentages circulate that are not in it. The safe claim is the direction.
  • Be present off your own domain. Omnibound's 2026 AEO statistics compilation puts roughly 85% of AI references on third-party sites. The AI Platform Citation Source Index 2026, synthesising six studies covering more than 680 million citations recorded between August 2024 and April 2026, found Reddit the single most-cited domain across generative engines at roughly 40% of aggregate citation frequency, with Wikipedia second, appearing in 26–48% of ChatGPT top-10 answers.
  • Stay current. Retrieval favours recently updated pages heavily on commercial questions.

What reduces the odds is the mirror image: content that exists only after client-side hydration, thin or duplicated pages, unclear page purpose, stale facts, and unsupported claims that are hard to stand behind as a reference.


How much of the market ChatGPT still is

ChatGPT is the default assumption in most AI visibility conversations, and it is becoming less true each quarter. Similarweb's generative AI traffic figures put ChatGPT's share of the category at around 53%, down from about 76% a year earlier, as Gemini passed a quarter of traffic and Claude grew fastest.

A programme scoped only to ChatGPT was defensible in 2024. In 2026 it leaves roughly half the category unmeasured, on engines that cite differently, refresh differently, and — per GetMentions — share only a sixth of their sources with each other. Google's own two surfaces do not agree with each other either, and AI Overviews select by fan-out rather than by rank.


What this cannot do

  • No submission endpoint exists. You cannot request a citation. You can only be crawlable, quotable and present in the sources retrieval already favours.
  • Memory mode is not addressable this quarter. If ChatGPT does not know who you are with search off, publishing more of your own pages will not change it before the next model.
  • Volatility caps how good the news can get. At roughly 1% seven-day persistence, "we hold position for query X" is not a claim anyone can honestly make about ChatGPT.
  • Blocking decisions are not reversible retroactively. A page excluded from the index while a bot was blocked was not cited during that period, and re-crawling runs on OpenAI's schedule.

How Lifewood approaches this

Lifewood measures the two modes separately as a matter of method, because blending them makes correct retrieval work look like failure for months. The instrument is a fixed set of buyer questions per market, run repeatedly rather than checked, reported as a rate with the raw answers retained — at roughly 79% daily churn, anything less is reporting noise with a decimal point on it.

Content work is aimed at the retrieval stage first: served HTML that carries the answer without JavaScript, self-contained passages, and evidence density in the passage rather than adjectives in the page. Producing that across markets is a delivery problem as much as a writing one, which is where 50+ languages and 40+ delivery centres across 30+ countries matter, since the question a buyer types in Portuguese is not the English question translated. See AEO services, GEO services and what gets you cited by AI answer engines.


Sources and further reading

  • Semrush, 2026 AI Visibility Index, 126 million US AI search prompts, January–April 2026.
  • Parse, AI citation volatility by industry, 693,509 answers, March–April 2026.
  • GetMentions, AI citation volatility: a 530,875-citation study, June 2026.
  • Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, GEO: Generative Engine Optimization, ACM SIGKDD 2024.
  • Zhang, He and Yao, From Citation Selection to Citation Absorption.
  • AI Platform Citation Source Index 2026, synthesis of six studies covering 680 million citations.
  • Similarweb, generative AI traffic share statistics, 2026.

Frequently asked questions

With search enabled it rewrites the question into search queries, retrieves candidate documents from OpenAI's own search index, reranks them, and composes an answer from the passages it selects, linking the ones it used. Selection happens at passage level, so a page is chosen for the specific paragraph that answers the rewritten query rather than for its overall quality.

About 15 on average, measured by the Semrush AI Visibility Index across 126 million US prompts between January and April 2026. Gemini cites about 3. That difference means ChatGPT has far more room for non-incumbent brands to appear in a given answer.

Because retrieval is probabilistic and the index moves. Parse found ChatGPT repeats only about 21% of its cited domains when the same question is asked again, and GetMentions found only about 1% of sources are cited on seven consecutive days. Any single check is one noisy sample.

OAI-SearchBot builds the search index ChatGPT cites from, so blocking it makes citation impossible. GPTBot collects training data and ChatGPT-User fetches pages during a live user request. They are separate directives in robots.txt and should be decided separately.

Not directly and not quickly. With search off the model answers from training data, so what moves it is what third parties published about you before the training cut-off. That is a reputation and earned-coverage problem, and it resolves only when a new model is trained.

No. ChatGPT search runs on OpenAI's own retrieval layer, crawled by OAI-SearchBot. Ranking well in Google does not by itself make you retrievable in ChatGPT, which is why the two need measuring separately.

Enough that daily churn of roughly 79% averages out. In practice that means a fixed set of buyer questions run repeatedly over weeks, reporting a rate rather than a position — not a handful of prompts checked once a month.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team