robots.txt shapes the behaviour of specific crawlers; it is not a single switch. Google-Extended governs the use of content for training and grounding in Gemini Apps and Vertex AI, but it does NOT govern inclusion in Google Search AI Overviews — there, what matters is Googlebot and page eligibility (indexed, and able to show a snippet). To manage visibility properly you need the exact bot names and the places where the rules do not apply.
You want to appear in AI search without handing your content to someone else's model training for free.
The logs show dozens of user agents and it is unclear which to block and which to let through.
robots.txt looks like a universal switch — but some requests legitimately ignore it.
Claim: the large AI companies split their crawlers by purpose, and blocking them wholesale is a mistake. Argument: the bot that gathers training data and the bot that indexes for AI search are different user agents under different policies. Evidence: OpenAI explicitly suggests you can "allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training OpenAI's generative AI foundation models" [70].
Claim: user-initiated requests are their own category. Argument: when a person asks an assistant to open a page, that is a user action rather than a bulk crawl. Evidence: OpenAI states of ChatGPT-User that "because these actions are initiated by a user, robots.txt rules may not apply" [70]; Perplexity documents the same for Perplexity-User, which generally ignores robots.txt because the fetch was requested by a person [74].
Operator: OpenAI. Token: GPTBot/1.4. Purpose: crawling content that may be used to train generative AI foundation models. Obeys robots.txt: yes — Disallow to stay out of training.
Operator: OpenAI. Token: OAI-SearchBot/1.4. Purpose: surfacing websites in ChatGPT's search features. Obeys robots.txt: yes — Allow for visibility in search.
Operator: OpenAI. Token: ChatGPT-User/1.0. Purpose: user-triggered actions in ChatGPT and Custom GPTs. Obeys robots.txt: "rules may not apply" (user-initiated).
Operator: OpenAI. Token: OAI-AdsBot/1.0. Purpose: validating the safety of pages submitted as ads on ChatGPT. Obeys robots.txt: applies to submitted ad landing pages.
Operator: Perplexity. Token: +https://perplexity.ai/perplexitybot. Purpose: indexing and linking in Perplexity search (not model training). Obeys robots.txt: yes — Allow is recommended.
Operator: Perplexity. Token: +https://perplexity.ai/perplexity-user. Purpose: visiting a page in response to a user's question. Obeys robots.txt: generally ignores it (user-initiated).
Operator: Anthropic. Identifier: claude.com/crawling/bots.json. Purpose: gathering web content for AI model training and development. Obeys robots.txt: yes — honours "do not crawl" directives.
Operator: Anthropic. Identifier: claude.com/crawling/bots.json. Purpose: user-initiated web access while Claude answers a question. Obeys robots.txt: a user action.
Operator: Anthropic. Identifier: claude.com/crawling/bots.json. Purpose: indexing content to improve search result quality. Obeys robots.txt: yes.
Operator: Google. Identifier: the general search crawler. Purpose: Search, Images, Video, News, Discover — including AI Overviews and AI Mode. Obeys robots.txt: yes.
Operator: Google. Identifier: a training control only. Purpose: training future Gemini models and grounding in Gemini Apps / Vertex AI. Obeys robots.txt: yes — no effect on ranking or AI Overviews.
Operator: Google. Identifier: a generic crawler. Purpose: one-off research crawls. Obeys robots.txt: yes — not tied to a specific product.
Operator: Common Crawl. Token: CCBot/2.0 (https://commoncrawl.org/faq/). Purpose: the open web dataset, used among other things for research. Obeys robots.txt: yes — Disallow to block.
User agent sources: OpenAI [70], Perplexity [74], Anthropic [75], Google [76], Common Crawl [77].
What follows is an example policy, not a universal recommendation: choose directives that match your goals and your jurisdiction.
Example policy "block training, stay in AI search" (directives listed inline): "User-agent: GPTBot · Disallow: /" · "User-agent: Google-Extended · Disallow: /" · "User-agent: CCBot · Disallow: /" · "User-agent: OAI-SearchBot · Allow: /" · "User-agent: PerplexityBot · Allow: /".
Example policy "limit ClaudeBot load instead of blocking it outright" (per Anthropic's documentation): "User-agent: ClaudeBot · Crawl-delay: 1".
Every directive is a trade-off, not a free improvement.
What you gain: privacy and a training opt-out — the content stays out of OpenAI's training data and out of the open Common Crawl dataset. What you pay: less AI discoverability outside Google Search — fewer chances of being cited in ChatGPT and in products built on Common Crawl.
What you gain: an opt-out from training future Gemini models and from grounding in Gemini Apps / Vertex AI. What you pay: nothing changes in Google Search AI Overviews or ranking — so it does not solve "take my site out of Google's AI answers".
What you gain: visibility and citation in ChatGPT and Perplexity search. What you pay: those platforms keep indexing the content, and your control over how it is reworded in an answer stays limited.
Claim: Google-Extended is not an off switch for AI Overviews. Argument: Google states plainly that "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". Evidence: blocking Google-Extended removes content only from Gemini training and from grounding in Gemini Apps and Vertex AI, while AI Overviews are built on the main Googlebot index [76]. This is the platform's own position, not an independent assessment.
Claim: robots.txt is a request, not a technical block. Argument: compliance depends on the operator acting in good faith, and user-initiated fetches legitimately bypass it. Evidence: OpenAI and Perplexity both document that ChatGPT-User and Perplexity-User may not obey robots.txt because a human initiated the action [70, 74].
Claim: user agent spoofing exists, so blocking by string is unreliable. Argument: bots can masquerade as legitimate ones. Evidence: Common Crawl warns that "we are aware of crawlers falsely identifying themselves as CCBot" and publishes its IP ranges at index.commoncrawl.org/ccbot.json for verification; Anthropic likewise recommends checking requests against its published IP list and cautions that blocking IP addresses "may not work correctly or persistently" [77, 75].
Official bot documentation — names and purposes of GPTBot / OAI-SearchBot / ChatGPT-User [70].
Official bot guide — robots.txt policy for PerplexityBot and Perplexity-User [74].
Support article — ClaudeBot / Claude-User / Claude-SearchBot, Crawl-delay [75].
Developer documentation — Google-Extended does not affect ranking [76].
CCBot page — the CCBot user agent, blocking, IP verification [77].
Every rule and phrasing above is the official position of the platform in question; policies change, so check the primary sources before you deploy anything.
Disallow the training bots (GPTBot, Google-Extended, CCBot) and allow the search ones (OAI-SearchBot, PerplexityBot). They are different user agents under different policies.
No. Google states that Google-Extended does not affect inclusion in Search or ranking. AI Overviews are built on the Googlebot index, not on Google-Extended.
If a user initiated the fetch (ChatGPT-User, Perplexity-User), robots.txt may not apply — this is documented behaviour at both OpenAI and Perplexity.
Not reliably. Anthropic and Common Crawl both warn about user agent spoofing and recommend verifying requests against their official IP lists rather than blocking ranges by hand.
OpenAI crawlers and user agents. https://developers.openai.com/api/docs/bots — official platform documentation.
How does Perplexity follow robots.txt? https://www.perplexity.ai/help-center/en/articles/10354969-how-does-perplexity-follow-robots-txt — official platform documentation.
Does Anthropic crawl data from the web and how can site owners block the crawler? https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler — official platform documentation.
Overview of Google crawlers and fetchers. https://developers.google.com/crawling/docs/crawlers-fetchers/overview-google-crawlers — official platform documentation.
CCBot documentation. https://commoncrawl.org/ccbot — official data-infrastructure documentation.
The Enigma editorial team.
18 May 2026.
18 May 2026.
Official platform documentation — the operator is stated next to each claim.
Bot names, purposes and robots.txt behaviour re-read from the OpenAI, Anthropic, Google and Common Crawl documentation; quotations are their own wording.
Perplexity's help-centre page returned HTTP 403 to our fetch, so its behaviour is reported without a verbatim quotation — check that page directly before relying on it.
The full list of sources and their reliability levels lives in the research catalogue.