← Research Lab·Research Lab · Technical guide

AI bot access: robots.txt for GPTBot, ClaudeBot, PerplexityBot, CCBot

robots.txt decides which agents may read your content. It is the cheapest check in a GEO audit — and a frequent place where visibility leaks unnoticed.

May 2026·6 min read

Different AI platforms use different user agents with different crawling purposes. A blocked bot will not extract your content; but robots.txt is access control, not removal from Search. A robots.txt that contradicts your visibility goals is a common error, and a cheap one to fix.

Every platform has its own agent

There is no single "AI bot". OpenAI documents its own crawlers: GPTBot for training and OAI-SearchBot for search scenarios [70]. Anthropic describes ClaudeBot, Claude-User and Claude-SearchBot — the last a separate search crawler whose blocking can reduce visibility in Claude search-like scenarios [75]. Perplexity distinguishes PerplexityBot (surfacing and linking, governed by robots.txt) from Perplexity-User (a user-requested fetch that typically ignores robots.txt) [74]. Policy is set per agent, not "for AI in general".

What robots.txt actually does

robots.txt governs crawler access to a URL; it does not keep the page out of the index. Google's overview of crawlers and fetchers explains which agents it uses for which tasks [76]. A blocked URL can still appear without a description if other pages link to it. Full exclusion from Search needs noindex or password protection — robots.txt is not the tool for that.

A map of the main agents

GPTBot / OAI-SearchBot

OpenAI's training crawler and search agent [70].

ClaudeBot / Claude-User / Claude-SearchBot

Anthropic crawling, with a separate search crawler [75].

PerplexityBot / Perplexity-User

Surfacing and linking (robots.txt) versus user-requested fetch, which typically ignores robots.txt [74].

Googlebot and fetchers

Google's various tasks, AI features included [76].

Bingbot

The base for Microsoft Copilot and AI search [71].

CCBot (Common Crawl)

The open web corpus feeding many systems [77].

Why Common Crawl deserves its own decision

Blocking live AI bots does not mean absence from their data. Common Crawl documents CCBot as open web-data infrastructure that many AI systems and research projects rely on [77]. What you decide about CCBot is a separate strategic decision from access for GPTBot or ClaudeBot.

Bing as the base for Copilot

Visibility in the Microsoft ecosystem starts with classic rules. The Bing Webmaster Guidelines set out requirements for quality, crawling, indexing and prohibited practices [71]. Through Bing's link to Copilot, those rules remain the baseline access layer for the corresponding AI products.

Sources (E-E-A-T)

70 · OpenAI

Crawlers and user agents (GPTBot, OAI-SearchBot). developers.openai.com/api/docs/bots

71 · Microsoft Bing

Webmaster Guidelines (quality, crawling).

74 · Perplexity

PerplexityBot / Perplexity-User and robots.txt.

75 · Anthropic

ClaudeBot / Claude-User / Claude-SearchBot and how to block them.

76 · Google Search Central

Overview of crawlers and fetchers.

77 · Common Crawl

CCBot documentation. commoncrawl.org/ccbot

Frequently asked questions

Can I hide a page from Google with robots.txt?

No. robots.txt governs crawler access but does not remove a URL from Search: a blocked URL can appear without a description. Exclusion needs noindex or password protection [76].

Is one rule "for AI bots" enough?

No. GPTBot, ClaudeBot, PerplexityBot, Googlebot and CCBot are different agents with different purposes; policy is set for each [70, 74, 75].

If I block ClaudeBot, does my content disappear from all AI?

No. Blocking one live agent does not remove content from open corpora such as Common Crawl, which other systems use [77].

Where should an access-focused GEO audit start?

By checking robots.txt and the firewall for Google, Bing, OpenAI, Perplexity, Anthropic and CCBot — one of the first and cheapest checks [76, 71].