robots.txt decides which agents may read your content. It is the cheapest check in a GEO audit — and a frequent place where visibility leaks unnoticed.
Different AI platforms use different user agents with different crawling purposes. A blocked bot will not extract your content; but robots.txt is access control, not removal from Search. A robots.txt that contradicts your visibility goals is a common error, and a cheap one to fix.
There is no single "AI bot". OpenAI documents its own crawlers: GPTBot for training and OAI-SearchBot for search scenarios [70]. Anthropic describes ClaudeBot, Claude-User and Claude-SearchBot — the last a separate search crawler whose blocking can reduce visibility in Claude search-like scenarios [75]. Perplexity distinguishes PerplexityBot (surfacing and linking, governed by robots.txt) from Perplexity-User (a user-requested fetch that typically ignores robots.txt) [74]. Policy is set per agent, not "for AI in general".
robots.txt governs crawler access to a URL; it does not keep the page out of the index. Google's overview of crawlers and fetchers explains which agents it uses for which tasks [76]. A blocked URL can still appear without a description if other pages link to it. Full exclusion from Search needs noindex or password protection — robots.txt is not the tool for that.
OpenAI's training crawler and search agent [70].
Anthropic crawling, with a separate search crawler [75].
Surfacing and linking (robots.txt) versus user-requested fetch, which typically ignores robots.txt [74].
Google's various tasks, AI features included [76].
The base for Microsoft Copilot and AI search [71].
The open web corpus feeding many systems [77].
Blocking live AI bots does not mean absence from their data. Common Crawl documents CCBot as open web-data infrastructure that many AI systems and research projects rely on [77]. What you decide about CCBot is a separate strategic decision from access for GPTBot or ClaudeBot.
Visibility in the Microsoft ecosystem starts with classic rules. The Bing Webmaster Guidelines set out requirements for quality, crawling, indexing and prohibited practices [71]. Through Bing's link to Copilot, those rules remain the baseline access layer for the corresponding AI products.
Crawlers and user agents (GPTBot, OAI-SearchBot). developers.openai.com/api/docs/bots
Webmaster Guidelines (quality, crawling).
PerplexityBot / Perplexity-User and robots.txt.
ClaudeBot / Claude-User / Claude-SearchBot and how to block them.
Overview of crawlers and fetchers.
CCBot documentation. commoncrawl.org/ccbot
No. robots.txt governs crawler access but does not remove a URL from Search: a blocked URL can appear without a description. Exclusion needs noindex or password protection [76].
No. Blocking one live agent does not remove content from open corpora such as Common Crawl, which other systems use [77].