← AI search trends·Blog · AI search trends

AI crawlers and robots.txt: a practical guide for site owners

robots.txt shapes the behaviour of specific crawlers; it is not a single switch. Google-Extended governs the use of content for training and grounding in Gemini Apps and Vertex AI, but it does NOT govern inclusion in Google Search AI Overviews — there, what matters is Googlebot and page eligibility (indexed, and able to show a snippet). To manage visibility properly you need the exact bot names and the places where the rules do not apply.

May 2026·8 min read

Context / Problem

Visibility versus training

You want to appear in AI search without handing your content to someone else's model training for free.

Dozens of user agents

The logs show dozens of user agents and it is unclear which to block and which to let through.

robots.txt is not a switch

robots.txt looks like a universal switch — but some requests legitimately ignore it.

The key distinction: training, search, and a user request

Claim: the large AI companies split their crawlers by purpose, and blocking them wholesale is a mistake. Argument: the bot that gathers training data and the bot that indexes for AI search are different user agents under different policies. Evidence: OpenAI explicitly suggests you can "allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training OpenAI's generative AI foundation models" [70].

Claim: user-initiated requests are their own category. Argument: when a person asks an assistant to open a page, that is a user action rather than a bulk crawl. Evidence: OpenAI states of ChatGPT-User that "because these actions are initiated by a user, robots.txt rules may not apply" [70]; Perplexity documents the same for Perplexity-User, which generally ignores robots.txt because the fetch was requested by a person [74].

User agents of the main AI crawlers

GPTBot

Operator: OpenAI. Token: GPTBot/1.4. Purpose: crawling content that may be used to train generative AI foundation models. Obeys robots.txt: yes — Disallow to stay out of training.

OAI-SearchBot

Operator: OpenAI. Token: OAI-SearchBot/1.4. Purpose: surfacing websites in ChatGPT's search features. Obeys robots.txt: yes — Allow for visibility in search.

ChatGPT-User

Operator: OpenAI. Token: ChatGPT-User/1.0. Purpose: user-triggered actions in ChatGPT and Custom GPTs. Obeys robots.txt: "rules may not apply" (user-initiated).

OAI-AdsBot

Operator: OpenAI. Token: OAI-AdsBot/1.0. Purpose: validating the safety of pages submitted as ads on ChatGPT. Obeys robots.txt: applies to submitted ad landing pages.

PerplexityBot

Operator: Perplexity. Token: +https://perplexity.ai/perplexitybot. Purpose: indexing and linking in Perplexity search (not model training). Obeys robots.txt: yes — Allow is recommended.

Perplexity-User

Operator: Perplexity. Token: +https://perplexity.ai/perplexity-user. Purpose: visiting a page in response to a user's question. Obeys robots.txt: generally ignores it (user-initiated).

ClaudeBot

Operator: Anthropic. Identifier: claude.com/crawling/bots.json. Purpose: gathering web content for AI model training and development. Obeys robots.txt: yes — honours "do not crawl" directives.

Claude-User

Operator: Anthropic. Identifier: claude.com/crawling/bots.json. Purpose: user-initiated web access while Claude answers a question. Obeys robots.txt: a user action.

Claude-SearchBot

Operator: Anthropic. Identifier: claude.com/crawling/bots.json. Purpose: indexing content to improve search result quality. Obeys robots.txt: yes.

Googlebot

Operator: Google. Identifier: the general search crawler. Purpose: Search, Images, Video, News, Discover — including AI Overviews and AI Mode. Obeys robots.txt: yes.

Google-Extended

Operator: Google. Identifier: a training control only. Purpose: training future Gemini models and grounding in Gemini Apps / Vertex AI. Obeys robots.txt: yes — no effect on ranking or AI Overviews.

GoogleOther

Operator: Google. Identifier: a generic crawler. Purpose: one-off research crawls. Obeys robots.txt: yes — not tied to a specific product.

CCBot

Operator: Common Crawl. Token: CCBot/2.0 (https://commoncrawl.org/faq/). Purpose: the open web dataset, used among other things for research. Obeys robots.txt: yes — Disallow to block.

User agent sources: OpenAI [70], Perplexity [74], Anthropic [75], Google [76], Common Crawl [77].

Ready-made robots.txt directives

What follows is an example policy, not a universal recommendation: choose directives that match your goals and your jurisdiction.

Example policy "block training, stay in AI search" (directives listed inline): "User-agent: GPTBot · Disallow: /" · "User-agent: Google-Extended · Disallow: /" · "User-agent: CCBot · Disallow: /" · "User-agent: OAI-SearchBot · Allow: /" · "User-agent: PerplexityBot · Allow: /".

Example policy "limit ClaudeBot load instead of blocking it outright" (per Anthropic's documentation): "User-agent: ClaudeBot · Crawl-delay: 1".

Every directive is a trade-off, not a free improvement.

Disallow for GPTBot / CCBot

What you gain: privacy and a training opt-out — the content stays out of OpenAI's training data and out of the open Common Crawl dataset. What you pay: less AI discoverability outside Google Search — fewer chances of being cited in ChatGPT and in products built on Common Crawl.

Disallow for Google-Extended

What you gain: an opt-out from training future Gemini models and from grounding in Gemini Apps / Vertex AI. What you pay: nothing changes in Google Search AI Overviews or ranking — so it does not solve "take my site out of Google's AI answers".

Allow for OAI-SearchBot / PerplexityBot

What you gain: visibility and citation in ChatGPT and Perplexity search. What you pay: those platforms keep indexing the content, and your control over how it is reworded in an answer stays limited.

An important detail about Google

Claim: Google-Extended is not an off switch for AI Overviews. Argument: Google states plainly that "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". Evidence: blocking Google-Extended removes content only from Gemini training and from grounding in Gemini Apps and Vertex AI, while AI Overviews are built on the main Googlebot index [76]. This is the platform's own position, not an independent assessment.

The limits of robots.txt: what the file does not do

Claim: robots.txt is a request, not a technical block. Argument: compliance depends on the operator acting in good faith, and user-initiated fetches legitimately bypass it. Evidence: OpenAI and Perplexity both document that ChatGPT-User and Perplexity-User may not obey robots.txt because a human initiated the action [70, 74].

Claim: user agent spoofing exists, so blocking by string is unreliable. Argument: bots can masquerade as legitimate ones. Evidence: Common Crawl warns that "we are aware of crawlers falsely identifying themselves as CCBot" and publishes its IP ranges at index.commoncrawl.org/ccbot.json for verification; Anthropic likewise recommends checking requests against its published IP list and cautions that blocking IP addresses "may not work correctly or persistently" [77, 75].

E-E-A-T: the official sources

OpenAI

Official bot documentation — names and purposes of GPTBot / OAI-SearchBot / ChatGPT-User [70].

Perplexity

Official bot guide — robots.txt policy for PerplexityBot and Perplexity-User [74].

Anthropic

Support article — ClaudeBot / Claude-User / Claude-SearchBot, Crawl-delay [75].

Google

Developer documentation — Google-Extended does not affect ranking [76].

Common Crawl

CCBot page — the CCBot user agent, blocking, IP verification [77].

Every rule and phrasing above is the official position of the platform in question; policies change, so check the primary sources before you deploy anything.

Frequently asked questions

How do I stay in AI search without ending up in model training?

Disallow the training bots (GPTBot, Google-Extended, CCBot) and allow the search ones (OAI-SearchBot, PerplexityBot). They are different user agents under different policies.

Will blocking Google-Extended remove me from AI Overviews?

No. Google states that Google-Extended does not affect inclusion in Search or ranking. AI Overviews are built on the Googlebot index, not on Google-Extended.

Why did an AI assistant open a page I disallowed in robots.txt?

If a user initiated the fetch (ChatGPT-User, Perplexity-User), robots.txt may not apply — this is documented behaviour at both OpenAI and Perplexity.

Can I block AI bots by IP address?

Not reliably. Anthropic and Common Crawl both warn about user agent spoofing and recommend verifying requests against their official IP lists rather than blocking ranges by hand.

Sources (E-E-A-T)

70 · OpenAI, current version

OpenAI crawlers and user agents. https://developers.openai.com/api/docs/bots — official platform documentation.

74 · Perplexity Help Center, current version

How does Perplexity follow robots.txt? https://www.perplexity.ai/help-center/en/articles/10354969-how-does-perplexity-follow-robots-txt — official platform documentation.

75 · Anthropic Support, current version

Does Anthropic crawl data from the web and how can site owners block the crawler? https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler — official platform documentation.

76 · Google Search Central, 2026

Overview of Google crawlers and fetchers. https://developers.google.com/crawling/docs/crawlers-fetchers/overview-google-crawlers — official platform documentation.

77 · Common Crawl, current version

CCBot documentation. https://commoncrawl.org/ccbot — official data-infrastructure documentation.

How we checked this material

Author

The Enigma editorial team.

Published

18 May 2026.

Updated

18 May 2026.

Sources

Official platform documentation — the operator is stated next to each claim.

Verification

Bot names, purposes and robots.txt behaviour re-read from the OpenAI, Anthropic, Google and Common Crawl documentation; quotations are their own wording.

Caveat

Perplexity's help-centre page returned HTTP 403 to our fetch, so its behaviour is reported without a verbatim quotation — check that page directly before relying on it.

The full list of sources and their reliability levels lives in the research catalogue.