Skip to main content
AEO/GEO

AI Crawlers: Which Bots to Allow, Which to Block, and Why

August 2026 · 8 min read · Updated September 2026

Short answer. Training crawlers and retrieval crawlers are different bots doing different jobs, and most blocking decisions treat them as one. Blocking a training crawler costs you nothing you can currently measure. Blocking a retrieval crawler makes citation impossible. The trap is that common defaults do not distinguish them: since 1 July 2025, new domains on Cloudflare have blocked GPTBot, ClaudeBot and PerplexityBot by default, under a single toggle labelled for training.

Key takeaways

  • A training crawler collects data to improve a future model; a retrieval crawler fetches pages to answer a live query or build a citation index — blocking one does not have the same cost as blocking the other.
  • Since 1 July 2025, new domains on Cloudflare block GPTBot, ClaudeBot and PerplexityBot by default, and that control does not separate training crawlers from retrieval crawlers.
  • Google-Extended governs Gemini training use only; blocking it does not remove a site from Google AI Overviews, because those are built on the ordinary Googlebot Search crawl.
  • AI crawlers fetch far more pages than they send back in referral traffic, which is a real argument for blocking training bots but not for blocking the retrieval bots that can actually produce a citation or a visit.
  • llms.txt has no confirmed effect on AI Overviews or ChatGPT citations and should not substitute for checking whether the crawlers that matter can actually reach your pages.

What do training crawlers and retrieval crawlers actually do?

They are separate user agents doing separate jobs, and a site can allow one while blocking the other in robots.txt. A training crawler gathers content used to improve a future model; a retrieval crawler fetches a page to build a search index or answer a live user request, and blocking it removes the possibility of a citation immediately.

User agent Operator Job Blocking it costs you
OAI-SearchBot OpenAI Builds the index ChatGPT search answers from Any possibility of a ChatGPT citation
GPTBot OpenAI Collects training data Presence in a future model's memory
ChatGPT-User OpenAI Fetches a page live during a user request The page loading when a user asks about it
ClaudeBot Anthropic Collects training data Presence in a future model's memory
Claude-SearchBot / Claude-User Anthropic Retrieval and live user fetches Citation in Claude answers
PerplexityBot Perplexity Builds its retrieval index Citation in Perplexity answers
Google-Extended Google Controls Gemini training use only Nothing in Search or AI Overviews
Googlebot Google Search index, and the basis of AI Overviews Search and AI Overviews together

Google-Extended is the one people get wrong in the expensive direction: it governs training use, not Search, so blocking it does not remove a site from AI Overviews, because AI Overviews are built on the ordinary Googlebot crawl — and the same is true of AI Mode. There is no separate opt-out for AI Overviews that keeps a site in Search. The OpenAI split runs the other way: ChatGPT cites from OAI-SearchBot's index, not from GPTBot's training crawl, so blocking "OpenAI" as one thing loses the citable half along with the trainable half.

Could your site be blocking AI crawlers without anyone deciding to?

Quite possibly, because the most common default now blocks retrieval crawlers as a side effect of blocking training crawlers. Since 1 July 2025, new domains on Cloudflare have blocked GPTBot, ClaudeBot and PerplexityBot by default, and that control does not separate a purely retrieval bot like PerplexityBot from a training bot.

A site can be well written, fully structured and schema-marked, and still be completely uncitable because of an infrastructure default set on the day the domain was onboarded. Four checks catch it, in the order they tend to fail:

  • Fetch /robots.txt and read it. Look for Disallow under each AI user agent by name, not just under *.
  • Check the CDN or WAF layer separately. A permissive robots.txt means nothing if the edge itself returns a 403 to the bot.
  • Check server logs or CDN analytics for real hits from OAI-SearchBot, PerplexityBot and Googlebot. Permission is a claim; a 200 in the log is evidence.
  • Check for JavaScript-dependent content. Retrieval crawlers are not guaranteed to execute it, so a page whose text arrives only after hydration can be an empty page to them.

Who is blocking AI crawlers, and how much traffic do they send back?

Blocking training crawlers is now common practice, and it costs the blocking site little it can measure, because those crawlers were never going to send a referral. Independent robots.txt analyses in 2026 consistently find GPTBot, CCBot, ClaudeBot, Google-Extended and Bytespider among the most frequently disallowed user agents, with a large share of surveyed domains blocking at least one of them.

The traffic economics explain why. Multiple 2026 analyses of Cloudflare-scale data find that AI crawlers fetch orders of magnitude more pages than they send back in referral clicks, and that training and mixed-purpose crawling account for the large majority of AI crawler traffic, with search and user-initiated fetches making up a small minority. The exact ratios vary by dataset and panel — one report's multiple can differ from another's by an order of magnitude for the same bot — but the direction is consistent across every source: AI crawlers take far more than they return, and Google's ordinary Search crawl returns far more referral traffic per page crawled than any AI-specific bot.

That is a real argument for blocking training crawlers, and publishers with genuine bandwidth costs or licensing leverage are right to make it. It is not an argument for blocking retrieval crawlers, which is the decision the common CDN defaults actually make on a site's behalf unless someone corrects it.

Should you block AI crawlers?

For most organisations selling something rather than licensing content, the answer splits by crawler type: allow the ones that can cite or refer you, and decide the training ones on principle rather than traffic. A defensible policy has five parts.

  1. Allow every retrieval and user-action crawler — OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User, Claude-User, Googlebot. These are the only bots that can produce a citation or a visit, so there is no upside to blocking them unless content is being deliberately withheld.
  2. Decide training crawlers on principle, not on traffic — GPTBot, ClaudeBot, CCBot, Google-Extended. Allowing them is a bet on being in a future model's memory: real but unmeasurable and slow. Blocking them costs nothing observable today. Either answer is defensible; pick one and write down why.
  3. Block scrapers with no answer surface, such as Bytespider and similar bots that neither cite nor refer — nothing is given up by excluding them.
  4. Verify at the edge, not just in robots.txt. Confirm the CDN, WAF and bot-management rules actually agree with the file; this is where a policy is usually contradicted in practice.
  5. Re-check quarterly, since user agents get added, platform defaults change, and a routine security review can revert the whole policy in one commit.

The asymmetry in step 2 is worth stating plainly: the cost of blocking a training crawler is unobservable, while the cost of blocking a retrieval crawler is immediate and total for that engine. When a decision has one measurable side and one unmeasurable side, the measurable side should decide it.

Does llms.txt help AI crawlers find your content?

No confirmed evidence supports that, and it should not be treated as a substitute for a crawler-access check. Google's own guidance states that llms.txt has no effect on Search rankings or AI Overviews, and Google's John Mueller has said no Google Search system reads or acts on the file, comparing it to the old keywords meta tag that Google stopped using because it was self-declared and easy to game.

Adoption has not translated into use. Independent studies in 2026 found llms.txt published on a meaningful minority of surveyed domains, yet an Ahrefs analysis of 137,000 sites found that roughly 97% of published llms.txt files received zero requests, and of the small remainder that were fetched, actual AI retrieval bots accounted for only a sliver of those requests. Publishing the file is harmless. Counting it as AI visibility work is not, because it can displace the crawler-access check that actually decides whether a page can be cited — a file no crawler reads cannot substitute for a Disallow line that every crawler obeys. A related question — what llms.txt is and whether a site needs one — covers the format itself in more detail.

What does allowing crawlers actually get you?

Being fetchable is a precondition for citation, not a guarantee of it. Crawler access makes citation possible; it does not make a citation likely, and no legitimate provider can promise a placement in any specific AI answer.

The engines also differ in what they do with the pages they do fetch: Perplexity's source list is far steadier than ChatGPT's, and Google's Search and AI Overviews surfaces can diverge from each other even though they share one crawl. Blocking decisions are not reversible after the fact, either — a page excluded from an index while a bot was blocked was not cited during that period, and re-crawling runs on the operator's schedule, not the publisher's. What ultimately earns the citation once a page is reachable is a separate question, covered in what actually gets a brand cited by AI answer engines.

How does Lifewood approach crawler access for AEO and GEO?

Lifewood treats crawler access as the first gate on any AEO or GEO engagement, checked before content is scoped, because content aimed at a page an engine cannot fetch has no possible return. The check is evidence-based rather than declarative: robots.txt is read by user agent, edge and WAF rules are read separately from the file, and server logs are inspected for actual 200 responses to OAI-SearchBot, PerplexityBot and Googlebot, since permission stated in a file is a claim and a log line is proof.

The training-crawler question is handled as a documented client policy decision rather than a silent technical default, so a later security review does not quietly reverse it. Rendering is checked in the same pass, since a retrieval crawler that executes no JavaScript reads an empty page regardless of what robots.txt permits. Lifewood's broader AEO and GEO work builds on that access check rather than skipping past it.

Frequently asked questions

GPTBot collects data used to train OpenAI's models. OAI-SearchBot builds the search index ChatGPT cites from when web search is enabled. They are separate user agents with separate robots.txt directives, and blocking only OAI-SearchBot removes the ability to be cited in ChatGPT answers while leaving training access untouched.

No. Google-Extended controls whether content is used for Gemini model training. AI Overviews are generated from the ordinary Googlebot crawl of the Search index, so the only way to leave AI Overviews is to leave Search itself — there is no separate opt-out.

Quite possibly. Since 1 July 2025, new domains on Cloudflare block GPTBot, ClaudeBot and PerplexityBot by default, and the control does not separate training crawlers from retrieval crawlers. Check robots.txt, the CDN bot rules and server logs before assuming access.

Block training crawlers if you object to unpaid training use; the measurable cost of doing so is close to zero. Avoid blocking retrieval crawlers unless content is deliberately being withheld from AI answers, since they are the only bots that can cite a page or send a visit.

There is no confirmed evidence that it does. Google has said the file has no effect on Search rankings or AI Overviews, and an Ahrefs study of 137,000 sites found that around 97% of published llms.txt files received zero requests in the period measured.

Not reliably, and not all of them do. Content that exists only after client-side hydration can be invisible to a retrieval crawler even when the bot is fully allowed in robots.txt. Server-rendering the text meant to be quoted is the safer default.

Sources and further reading

  1. We Analyzed robots.txt Across Cloudflare's Network: Publishers Now Block Training Bots and Allow Answering Bots
  2. AI Crawler & Bot Traffic Statistics 2026
  3. AI Crawler Access Control: The 2026 Decision Matrix
  4. Google says normal SEO works for ranking in AI Overviews and llms.txt won't be used
  5. We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team