Skip to main content
AEO/GEO

How ChatGPT Decides Which Sources to Cite

August 2026 · 9 min read · Updated September 2026

Short answer. ChatGPT answers from two different places: a live retrieval index when search is on, and trained memory when it is off. Only the first can cite a URL. When it does cite, the Semrush AI Visibility Index 2026 — built on 126 million US AI search prompts — puts it at about 15 sources per answer, against roughly 3 for Gemini. And it repeats very few of them: asked the same question twice, Parse's study of 693,509 answers found ChatGPT reused only 21.2% of its cited domains. Visibility here is a rate, never a position.

Key takeaways

  • Every citation on every engine requires two stages: a document has to be retrieved before it can be selected, and wording a page for citability changes only the second stage.
  • ChatGPT runs two answering modes — search on (retrieval, can cite URLs) and search off (trained memory, cannot cite URLs) — and they respond to completely different levers.
  • The Semrush AI Visibility Index 2026 found ChatGPT cites about 15 sources per answer versus about 3 for Gemini, a five-fold difference in available slots.
  • Parse's study of 693,509 answers found ChatGPT reuses only 21.2% of cited domains when the same question is asked twice; GetMentions found just 1.1% of sources are cited on all seven days of a week.
  • Similarweb's traffic data put ChatGPT's share of generative AI traffic at around 53%, down from about 76% a year earlier, as Gemini and Claude both gained share.

How does ChatGPT decide which sources to cite?

With search enabled, ChatGPT rewrites the question into search queries, retrieves candidate documents from OpenAI's own index, reranks them, and composes an answer from the passages it selects, linking the ones it used. With search off, it answers from training data and has no document to point at, so it cannot cite anything.

Retrieval is the stage where a crawler-built index returns candidate documents for a rewritten query; selection is the later stage where the model chooses which retrieved passages to quote and link. A page that never reaches the first stage cannot benefit from work aimed at the second.

Search on (retrieval) Search off (memory)
What happens Runs searches, reads pages, then answers Answers from training data
Can it cite a URL? Yes No — it has no document to point at
What moves it What is published and crawlable now What the wider web said before the training cut-off
Response time Days to weeks Only when a new model ships
Who can influence it You, fairly directly Mostly other people writing about you

If a report says a brand's ChatGPT visibility is 0% and does not say which mode it measured, it has told you nothing. Being absent from retrieval is a publishing and crawlability problem that can be fixed this quarter. Being absent from memory is a reputation problem measured in model generations.

What happens at each stage of the retrieval pipeline?

With search on, ChatGPT behaves as a retrieval-augmented system, and every stage in that pipeline is a place a page can be excluded.

  1. A crawler builds the index. OpenAI runs separate bots for separate jobs: OAI-SearchBot builds the search index ChatGPT answers from, GPTBot collects training data, and ChatGPT-User fetches a page live when a request requires it. Blocking the wrong one costs the wrong thing — see AI crawlers and AI search visibility for which directive controls which bot.
  2. The question becomes searches. The model rewrites the user's question into one or more queries; small changes in phrasing produce different queries, and therefore different sources.
  3. Documents are retrieved and reranked. Candidates are scored for relevance to the rewritten queries. This is the stage where passage-level writing decisions pay off or do not.
  4. The answer is composed with citations attached. The model writes from the documents in its context window and links the ones it used.

Position in the retrieved context matters as well as content quality, since the model sees only a subset of the web. Being used and being cited are also not the same outcome: research on citation absorption, notably Zhang, He and Yao's work distinguishing citation selection from citation absorption, finds that a retrieved source can contribute language, evidence or structure to an answer while a different source receives the visible link. Citation count is an incomplete measure of influence, though it remains the only one observable from outside.

Why does ChatGPT have room for more brands than Gemini?

ChatGPT simply attaches more citation slots to each answer, so more distinct sources can appear in the same response.

The Semrush AI Visibility Index 2026, drawn from 126 million US AI search prompts recorded between January and April 2026, found ChatGPT cites an average of about 15 sources per response while Gemini cites about 3. That five-fold difference is worth planning around: fifteen slots means a mid-sized brand can realistically appear alongside incumbents in the same answer, while three slots makes winning close to winner-take-all. The same content programme has very different odds on the two engines, which is one reason a single blended "AI visibility score" across engines is not an actionable number.

Why do the cited sources change every time the same question is asked?

Because retrieval is probabilistic and the underlying index moves, so the set of documents ChatGPT retrieves for a given query is not fixed from one run to the next.

Parse, analysing 693,509 answers across 16,143 ChatGPT prompts and 15,805 Google prompts between 26 March and 25 April 2026, found that asking the same question twice returned only 21.2% of the same cited domains on ChatGPT and 31.5% on Google AI Overviews. Widening the window to a week raised overlap only to 26.7% and 36.8%.

GetMentions, working from 530,875 citations across 181,225 distinct URLs and 2,398 queries over seven consecutive days in June 2026, measured the same instability from the other direction:

Engine Day-over-day source churn Sources cited on all seven days
Gemini 88.3% 0.4%
ChatGPT 79.2% 1.1%
Google AI Mode 75.9% 2.6%
Perplexity 44.4% 11.1%

The same study found 84% of the sources for a question were cited by only one of the four engines. Three practical consequences follow:

  • A single check proves nothing in either direction — not appearing once is not absence, and appearing once is not visibility.
  • A week-on-week change measured from a handful of prompts is noise; at roughly 79% daily churn, a claim needs repeated runs across a fixed question set before it means anything.
  • Engines need measuring separately, since an averaged score across engines that share only a sixth of their sources hides the only thing a team could act on. Perplexity behaves nothing like ChatGPT in this respect.

What actually raises the odds of being cited?

Given that volatility, the goal is not to hold a fixed slot but to be in the candidate pool often enough that the odds land favourably at a useful rate.

  • Be fetchable by the right bot. OAI-SearchBot decides whether ChatGPT can cite a page at all, and many sites block it without knowing, because a platform default did it for them.
  • Be quotable at passage level. Self-contained paragraphs, question-shaped headings, one claim per paragraph, and the answer in the first sentence all make a passage easier to lift and cite intact — a practice covered in structuring a website so AI engines cite you.
  • Carry evidence, not adjectives. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande's GEO benchmark (ACM SIGKDD 2024), covering roughly 10,000 queries across nine datasets, found targeted content changes raised visibility in generative engine responses by up to 40% through Quotation Addition, with Statistics Addition adding a further roughly 30% uplift — authority-style edits that outperformed cosmetic changes such as rewriting, simplification and keyword work. Second-hand versions of that paper circulate with different per-tactic percentages; the direction, not the exact figure, is the safe claim.
  • Be present off-domain, on the sites that already get cited. Third-party mentions on forums, review sites and reference sites feed generative engines independently of a brand's own pages, a dynamic covered in why third-party brand mentions matter for GEO.
  • Stay current. Retrieval favours recently updated pages heavily on commercial questions.

What reduces the odds is the mirror image: content that exists only after client-side hydration, thin or duplicated pages, unclear page purpose, stale facts, and unsupported claims that are hard to stand behind as a reference.

How much of the AI search market does ChatGPT still represent?

Less than it did a year ago, and the share keeps shrinking each quarter as competing engines take a growing slice of generative AI traffic.

Similarweb's generative AI traffic figures put ChatGPT's share of the category at around 53%, down from about 76% a year earlier, as Gemini passed a quarter of traffic and Claude grew fastest. A programme scoped only to ChatGPT was defensible in 2024; today it leaves roughly half the category unmeasured, on engines that cite differently, refresh differently, and — per GetMentions — share only a sixth of their sources with each other. Google's own two surfaces do not agree with each other either, and AI Overviews select by fan-out rather than by rank.

What can't be fixed with content alone?

Some parts of ChatGPT visibility sit outside what publishing and technical work can move, at least on a useful timescale.

  • No submission endpoint exists — a brand cannot request a citation, only be crawlable, quotable and present in the sources retrieval already favours.
  • Memory mode is not addressable this quarter: if ChatGPT does not recognise a brand with search off, publishing more of that brand's own pages will not change it before the next model trains.
  • Volatility caps how good the news can get — at roughly 1% seven-day persistence, "we hold position for query X" is not a claim anyone can honestly make about ChatGPT.
  • Blocking decisions are not reversible retroactively: a page excluded from the index while a bot was blocked was not cited during that period, and re-crawling runs on OpenAI's own schedule.

How does Lifewood approach measuring and improving ChatGPT visibility?

Lifewood measures the two modes separately as a matter of method, because blending retrieval and memory results makes correct retrieval work look like failure for months. The instrument is a fixed set of buyer questions per market, run repeatedly rather than checked once, and reported as a rate with the raw answers retained — at roughly 79% daily churn, anything less is reporting noise with a decimal point attached.

Content work is aimed at the retrieval stage first: served HTML that carries the answer without JavaScript, self-contained passages, and evidence density in the passage rather than adjectives on the page. Producing that consistently across markets is a delivery problem as much as a writing one, which is where Lifewood's 100+ languages and 40+ delivery centres across 30+ countries matter, since the question a buyer types in Portuguese is not the English question translated. See what actually gets you cited by AI answer engines for the broader set of practices this applies across engines, not only ChatGPT.

Frequently asked questions

With search enabled it rewrites the question into search queries, retrieves candidate documents from OpenAI's own search index, reranks them, and composes an answer from the passages it selects, linking the ones it used. Selection happens at passage level, so a page is chosen for the specific paragraph that answers the rewritten query rather than for its overall quality.

About 15 on average, measured by the Semrush AI Visibility Index across 126 million US prompts between January and April 2026. Gemini cites about 3. That difference means ChatGPT has far more room for non-incumbent brands to appear in a given answer.

Because retrieval is probabilistic and the index moves. Parse found ChatGPT repeats only about 21% of its cited domains when the same question is asked again, and GetMentions found only about 1% of sources are cited on seven consecutive days. Any single check is one noisy sample.

OAI-SearchBot builds the search index ChatGPT cites from, so blocking it makes citation impossible. GPTBot collects training data and ChatGPT-User fetches pages during a live user request. They are separate directives in robots.txt and should be decided separately.

Not directly and not quickly. With search off the model answers from training data, so what moves it is what third parties published about a brand before the training cut-off. That is a reputation and earned-coverage problem, and it resolves only when a new model is trained.

No. ChatGPT search runs on OpenAI's own retrieval layer, crawled by OAI-SearchBot. Ranking well in Google does not by itself make a page retrievable in ChatGPT, which is why the two need measuring separately.

Sources and further reading

  1. Semrush, 2026 AI Visibility Index — 126 million US AI search prompts, January–April 2026.
  2. Parse, AI citation volatility by industry — 693,509 answers, March–April 2026.
  3. GetMentions, AI Citation Volatility: A 530,875-Citation Study — June 2026.
  4. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, GEO: Generative Engine Optimization, ACM SIGKDD 2024
  5. Similarweb, AI Search Stats in 2026 — generative AI traffic share.
  6. AEO services
  7. GEO services

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team