Skip to main content
AEO/GEO

How to Get Your Company Cited by Perplexity and Gemini

Short answer. Both answer from live retrieval, so on-site changes can earn citations within days — neither is answering from training memory. They read different indexes: PerplexityBot…

Mumu D. · July 2026 · 13 min read

Download PDF

Short answer. Both answer from live retrieval, so on-site changes can earn citations within days — neither is answering from training memory. They read different indexes: PerplexityBot builds Perplexity's own, while Googlebot's Search index feeds AI Overviews, AI Mode and the Gemini app's grounding. That distinction has one practical consequence worth knowing before you touch robots.txt: Google-Extended controls Gemini app training and grounding only, and has no effect whatsoever on Google Search's AI features, which use Googlebot.

Only 11% of the domains cited by ChatGPT are also cited by Perplexity, according to one 2026 citation index. Nobody has published the equivalent figure for Perplexity against Gemini, but the reasons for the gap are structural, and they apply here too.

Both engines answer from a live search rather than from memory. That is the important thing they have in common, and it is why a page can earn a citation on either one within days of being changed. But they search different indexes, take instructions from different crawlers, count a different number of sources per answer, and reach for different kinds of pages when the question turns commercial.

So the work splits into three parts: what serves both engines at once, what has to be done separately for each, and a single list you can hand to whoever owns the website.

What the two engines share Both are retrieval engines, not memory engines. Perplexity's own documentation describes the product as searching the internet in real time and attaching numbered citations to every answer. Google's guidance for its generative features describes retrieval-augmented generation, which it also calls grounding, as pulling relevant, up-to-date pages from the Search index and generating a response with clickable links to the pages that support it. In both cases the answer is assembled from pages fetched at query time. That is different from asking a model with no tools, which answers from training weights that on-site changes cannot touch.

Both reward the same shape of page. An engine building an answer under a time budget needs a passage it can lift cleanly.

Question-shaped headings, a direct answer in the first two sentences under each, one idea per section, and a stated date all raise the odds that the passage is extractable. The largest controlled study of this, Aggarwal and colleagues' GEO benchmark across 10,000 queries, found that adding authoritative quotations raised citation visibility by up to 40% and adding statistics by around 30%, while keyword stuffing scored minus 10%. Neither engine has published anything to contradict that.

Both lean on third parties for judgement questions. When the question is "what is X", your own page is a plausible source.

When the question is "which X is best", both engines reach for pages that already rank the options. On Gemini with Google Search grounding, one July 2026 study of 100 "best [vertical] in [city]" queries found directory and ranking sites cited in 78% of answers. On Perplexity, Profound's citation data puts G2, Gartner, NerdWallet, PCMag, TripAdvisor and Yelp among the most-cited domains for commercial intent. The lesson is the same: for recommendation queries, being described accurately on the pages the engine already trusts matters more than your own homepage.

Both are volatile. Profound's own research shows 40% to 60% of cited domains change month to month across the major platforms.

A citation is a state, not a possession, which is why the checklist at the end has a refresh line in it.

Where they differ, and why the differences are not cosmetic The index. Perplexity runs its own crawler, PerplexityBot, which its documentation says is designed to surface and link websites in Perplexity search results and is not used to train foundation models. Gemini in Google Search (AI Overviews and AI Mode) uses the ordinary Google Search index fetched by Googlebot. The Gemini app, when grounding is on, uses Google Search as its retrieval tool and returns the sources as structured grounding metadata. Practically: a page that ranks well in Google has a head start on Gemini surfaces, and no start at all on Perplexity unless PerplexityBot can reach it.

The crawler rules. This is the point most teams get wrong, and it can silently remove you from one engine while you optimise for the other. Google's Search Central documentation is explicit that its AI features use Googlebot, and the separate Google-Extended token controls whether content is used for Gemini model training and grounding but has no effect on Google Search, including AI Overviews and AI Mode. Perplexity documents two agents: PerplexityBot for indexing, and Perplexity-User, which fetches a page when a user asks a question and, because a user requested the fetch, generally ignores robots.txt. Perplexity also publishes IP ranges and a WAF whitelisting guide, because a security rule that blocks unknown bots is enough to remove a site from its results.

How many sources per answer. Semrush's 2026 AI Visibility Index, built on 126 million US prompts, reports that Gemini cites an average of three sources per response, drawing on a small pool that includes Wikipedia, Reddit and YouTube. Perplexity typically shows four to eight. Three slots is a much tighter door than eight, so on Gemini the competition is for a very short list and the question of whether you are the best single source for a sub-question matters more.

Which third parties. Perplexity's citation profile is unusually concentrated on Reddit: one 2026 index puts Reddit at 20% to 24% of Perplexity citations, with 99% of those citations pointing to specific threads rather than subreddit pages. Google's AI surfaces show a self-referential tilt, with roughly 43% of AI Overview citations going to Google-owned properties including YouTube by one analysis. So the "be present where the engine already looks" advice resolves differently: on Perplexity, that is community threads and review aggregators; on Gemini, that is YouTube, Google Business Profile and the pages Google already ranks.

What the vendor tells you to do. Google says plainly that there are no additional requirements to appear in AI Overviews or AI Mode, that llms.txt files and special markup are ignored, and that chunking content or rewriting it for AI is unnecessary. Perplexity publishes crawler and WAF guidance but no ranking guidance. Neither has published a source-selection specification, so everything beyond access is inferred from measurement.

Two retrieval engines, two doors PERPLEXITY GEMINI (SEARCH SURFACES AND APP)

INDEX INDEX Own crawl (PerplexityBot) plus live fetch (Perplexity-User)

Google Search index via Googlebot; app grounds through Google Search ACCESS RULE ACCESS RULE robots.txt for the bot; the user fetch generally ignores it. WAF must whitelist published IPs Googlebot for Search AI features; Google-Extended only affects app training and grounding SOURCES / ANSWER SOURCES / ANSWER Typically 4 to 8 Average 3 (Semrush, 2026)

LEANS ON LEANS ON Reddit threads, review aggregators (G2, Gartner, Yelp), recent pages Wikipedia, Reddit, YouTube and Google-owned properties; directory sites on "best in city" queries Same page shape works on both. The access layer, the number of slots and the third parties to be present on do not.

The question type changes which engine you can win from your own site Yext's 6.8-million-citation study found that for objective, unbranded factual questions, first-party websites accounted for more than 40% of citations across the engines tested, and its Q1 2026 follow-up reported website citation shares of 52% on Gemini and 51% on Perplexity in a location-grounded dataset. That is the good news: for "what does X cost", "how does X work" and "what is X's policy on Y", your own page is a strong candidate on both engines.

The bad news is the other half. Aleyda Solis's June 2026 analysis of 15 SaaS brands found 84% to 93% of AI citation weight on thirdparty sites for their categories, and Lily Ray's study of 100 B2B "best software" queries found that when Google cited a brand's own self-ranked listicle, it recommended a competitor 69% of the time. Publishing "best [category]" with yourself at the top gets you cited and not recommended.

So the sorting rule is: own-site work for factual and procedural questions on both engines; third-party presence for evaluative questions on both engines; and accept that the third parties differ by engine.

Where this connects to our own work Declaring the interest: Lifewood runs managed AEO and GEO programmes, and the access audit described below is the first thing we do on every one of them.

Two observations from that work that hold regardless of who does it.

The first is that most "we are invisible on Perplexity" cases are access cases, not content cases. A WAF rule inherited from a previous security review, a CDN bot-fight setting, or a robots.txt written when "AI crawler" meant "training crawler" is enough to remove a site from one engine while it performs perfectly on the other. Because the two engines have different crawlers and different robots behaviour, the same site can be fully visible on Gemini and absent from Perplexity without anyone having decided that. Server logs settle it in an afternoon; content changes take weeks and would not have helped.

The second is about multilingual programmes. Retrieval is language-scoped on both engines. A brand whose Thai-market facts exist only in English is competing for Thai answers with whatever Thai-language pages the engine can find, including inaccurate ones.

Google's guidance on managing multilingual sites is unchanged by AI features, and Perplexity searches in the language of the question. Producing native-language pages that a native reviewer has checked is therefore not a localisation nicety; it is the difference between being a candidate and not.

One checklist for both engines Do these in order. The first three are free and decide whether anything after them can work.

  1. Confirm access, per engine, from logs. Grep for Googlebot, PerplexityBot and Perplexity-User in server logs for the last 30 days. If any is absent from a public page you care about, that engine cannot cite it. Check robots.txt, CDN bot settings and WAF rules separately; Perplexity publishes IP ranges and a whitelisting guide.

  2. Do not confuse Google-Extended with Googlebot. Blocking Google-Extended removes you from Gemini app grounding and training only. It does nothing to AI Overviews or AI Mode. Decide each deliberately.

  3. Serve the main content in HTML. Google says it can process JavaScript but calls it more complex; Perplexity's fetcher has a shorter time budget than a full render. If the answer is behind a click, a tab or a script, treat it as absent.

  4. Put the answer in the first two sentences under a question-shaped heading. Both engines extract passages. Make the extractable span obvious and self-contained.

  5. Add evidence, not adjectives. Quotations from named sources, dated statistics, and a visible methodology. This is the change the 10,000-query GEO benchmark measured as effective; keyword density was measured as harmful.

  6. Date everything and refresh on a schedule. Both engines prefer recent pages for commercial and evaluation questions. Set a refresh cadence for the pages that matter, and change the content, not just the date stamp.

  7. Earn presence on the third parties each engine trusts. For Perplexity: genuine participation in the Reddit threads where buyers already discuss your category, and complete profiles on the review aggregators it cites. For Gemini: a complete Google Business Profile, YouTube content with transcripts, and inclusion in independent "best of" lists. Google's guidance warns against seeking inauthentic mentions; the evidence says authentic ones move both engines.

  8. Reconcile your facts across the web. Semrush found that on Gemini, the overlap between brands mentioned and domains cited can be as low as 30%. The engine can name you from third-party evidence without citing you at all. Make sure the name, category, pricing and claims on those third-party pages agree with your own, or the engine will be choosing which version to believe.

  9. Measure engine by engine, not as one number. A visibility score averaged across engines hides an access failure on one of them. Track a fixed prompt set on each engine separately and read the sources, not just the mention.

  10. Publish comparisons that name competitors honestly. Comparison and listicle formats took 40% of commercial-intent citations in Wix Studio's analysis of one million citations. A comparison that concedes where a rival wins is citable. A comparison that does not is a brief for the competitor the engine likes better.

Where the effort goes, by question type Factual and procedural questions: first-party site share of citations >40% (Yext, 6.8M citations)

Evaluative "best X" questions: third-party share of citation weight 84% to 93% (Solis, 2026)

Gemini "best in city" answers citing a directory or ranking site 78% (Acromatico, 100 queries)

Own self-ranked listicle cited but competitor recommended (AI Overviews)

69% (Ray, 2026)

Own-site work wins the factual questions on both engines. Third-party presence wins the evaluative ones, and the third parties differ by engine.


Key takeaways

  • Perplexity and Gemini both answer from live retrieval, so on-site changes can earn citations on either within days. Neither answers from training memory when search is on.
  • They use different indexes: PerplexityBot builds Perplexity's; Googlebot's Search index feeds AI Overviews, AI Mode and the Gemini app's grounding.
  • Google-Extended controls Gemini app training and grounding only. It has no effect on Google Search AI features, which use Googlebot.
  • Perplexity-User fetches pages at a user's request and generally ignores robots.txt; Perplexity publishes IP ranges and asks WAFs to whitelist them.
  • Gemini cites an average of three sources per answer; Perplexity typically four to eight. Fewer slots means tighter competition.
  • Perplexity leans on Reddit threads (20% to 24% of its citations by one index) and review aggregators. Google surfaces lean on YouTube, Wikipedia and Google-owned properties.
  • Google says no special markup, llms.txt, chunking or AI-specific rewriting is needed. Perplexity publishes access guidance only.
  • Both reward evidence: authoritative quotations lifted citation visibility up to 40% and statistics around 30% in the 10,000-query GEO benchmark; keyword stuffing scored minus 10%.
  • First-party sites earn over 40% of citations for objective factual questions; third parties hold 84% to 93% of citation weight for evaluative SaaS queries.
  • A brand's own self-ranked listicle was cited but the competitor recommended 69% of the time in Google AI Overviews.
  • On Gemini, mentioned brands and cited domains overlap as little as 30%, so facts on third-party pages must agree with your own.
  • 40% to 60% of cited domains change monthly; citation requires a refresh cadence, not a one-off project.
  • Access failures are the most common cause of single-engine invisibility and are diagnosed from server logs, not content audits.

Sources and further reading

Frequently asked questions

No. They retrieve from different indexes and show measurably different citation profiles. Perplexity concentrates on Reddit and review aggregators; Gemini and Google's Search AI features draw on a smaller pool including Wikipedia, Reddit and YouTube. Optimise for each, not for "AI".

No. Google states that Google-Extended affects Gemini app training and grounding, not Google Search. AI Overviews and AI Mode use Googlebot and the Search index.

It does not, for indexing. PerplexityBot respects robots.txt. Perplexity-User, which fetches a page because a user asked a question about it, is documented as generally ignoring robots.txt because the request came from a person.

Google's guidance says no. It ignores llms.txt and requires no special schema to appear in AI features; existing structured data remains useful for ordinary rich results.

Start with the pages that answer the factual questions buyers ask about you: pricing, process, specifications, policies. Those are the questions where first-party pages win on both engines. Evaluative questions are won elsewhere.

Both engines retrieve live, so a corrected page can be cited within days of being recrawled.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team