Agent Traffic

See and govern which
AI agents read your site.

GPTBot/OAI-SearchBot, ClaudeBot/Claude-SearchBot, PerplexityBot/Perplexity-User, Googlebot and CCBot are different agents with different purposes and different attitudes to robots.txt. Agent Traffic brings them into one access and policy map.

"AI bot" is not one agent

Every platform has its own user agent and its own crawling purpose; a single rule "for AI" is not enough.

robots.txt ≠ removal from Search

A blocked bot will not extract your content, but robots.txt is access control — not a way to take a page out of Search.

Common Crawl is its own decision

What you decide about open corpora differs from access for live bots, and the decision is often made implicitly.

Agent Traffic tells you exactly who is allowed to read the page

A page about AI agents owes the reader more than a list of bots. It has to surface the decision: who to open, who to close, where training differs from retrieval, and how to catch policy drift.

Short AI answer

Agent Traffic is an access map for AI, search and user-requested agents: GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Googlebot and CCBot are each checked separately. The decision separates training, search retrieval and user-triggered fetches [70, 74, 75, 76, 77].

A real choice

Default-open and split policy carry different risks

A team can open search and retrieval agents while restricting the training crawler. That is not "correct for everyone" — it is a business decision about visibility, data and legal comfort.

Evidence
The OpenAI, Anthropic, Perplexity, Google and Common Crawl docs describe different user agents with different purposes.
Owner action
Build a policy table: agent, purpose, allowed, blocked, owner, review date.
Limit of the conclusion
A legal position cannot be derived from an SEO audit; the business owner has to confirm it.
Technical story

robots.txt can look right while the CDN still cuts access

What usually breaks inside a project is not the idea of the policy but its execution: robots.txt allows the bot, and a WAF rule, CDN rule or geo restriction returns an unexpected status.

Evidence
Compare robots.txt, a live fetch, headers, status codes and the behaviour of different user agents.
Owner action
Add a recurring accessibility check and record drift as a technical finding.
Limit of the conclusion
Without server logs you cannot claim real bot traffic volume; that figure is marked N/A.
Ownership

Bot access deserves an owner, like a security rule

A policy kept "somewhere in SEO" goes stale with every release. It holds up better as an infrastructure rule with a responsible owner, a review date and an expected behaviour.

Evidence
An open loop appears wherever the intended policy and the actual fetch disagree.
Owner action
Put the policy checklist into the release process and link it to the GEO audit.
Limit of the conclusion
Granting crawler access does not oblige any model to cite the page.
Agent Traffic · agent access mapSchematic
1Input: your content and your bot access policy
2Processing: preparation / extraction / measurement
3Signal: citations, absorption, gaps
4Output: prioritized actions (with no guarantee of inclusion)
Illustrative process schematic, not real data.
01

The agent map

GPTBot/OAI-SearchBot, ClaudeBot/Claude-User/Claude-SearchBot, PerplexityBot/Perplexity-User — a policy for each one separately [70, 74, 75].

02

What robots.txt actually does

It controls crawler access to a URL; it does not keep the page out of the index [76].

03

Common Crawl on its own

CCBot is open web-data infrastructure; the decision belongs in its own explicit loop [77].

04

Policy drift

Track the gaps between intent and the actual state of robots.txt and the firewall [70, 74, 75].

What the product includes.

🧭

Agent map

One access map, per agent.

Read more →
🚦

Access vs Search

Access control and Search exclusion are not the same thing.

Read more →
🌐

The Common Crawl loop

A separate strategic decision about CCBot.

Read more →
📉

Policy drift

Intent versus the actual state of access.

Read more →
70 · OpenAI
Crawlers and user agents (GPTBot, OAI-SearchBot).
↗
74 · Perplexity
PerplexityBot / Perplexity-User and robots.txt.
↗
75 · Anthropic
ClaudeBot / Claude-User / Claude-SearchBot.
↗
76 · Google Search Central
Overview of crawlers and fetchers.
↗
77 · Common Crawl
CCBot documentation.
↗

Stories from audits, not just theory

What we put into the product is not abstract SEO advice but the patterns that keep recurring in GEO audits: what the model saw, which source it cited, where the brand went missing, and which page has to be rewritten.

Cited definition

An Enigma field note is an anonymized audit story containing a checkable signal, the action the page owner took, and the limits of the data. The format helps people and AI understand not only what to do, but why the recommendation appeared at all.

B2B SaaS · brand visibility

The brand is mentioned, but the source belongs to someone else

In a typical B2B SaaS audit the brand shows up for a direct branded query, while comparison and buying-intent prompts lean on review sites or competitor pages. The problem is not the brand name — it is the absence of a cited methodology and a comparison of its own.

Evidence
The prompt matrix shows mentions without an owned citation; citation rows point to a third-party source as the basis of the answer.
Action
Publish a methodology page, comparison answer blocks, and an FAQ covering selection criteria.
Limitation
Client names and uplift figures are not shown without confirmed permission; impact stays N/A.
E-commerce · category intent

AI picks directories over your category pages

In e-commerce the model often cites marketplace directories, because the category page carries no short block on selection criteria, availability, returns and alternatives. The page exists — it just offers no answer-ready evidence.

Evidence
Citation context resolves to an external directory; the brand category is in the sitemap but does not cover source-authority intent.
Action
Add a buyer guide, a criteria table, a shipping and returns block, and an FAQ that matches the visible content.
Limitation
Demand estimates made without a live SEO export are labelled Estimated, never Measured.
Agency · multi-client governance

The content team and the technical team pull different levers

A recurring agency problem: an editor adds FAQs and comparisons while robots.txt, the CDN or a WAF quietly block some AI and user-requested agents. The content is citation-ready on paper and not always reachable in practice.

Evidence
The crawler checklist records the gap between the intended access policy and the rules bots actually meet.
Action
Split training crawler policy, retrieval/search access and user-triggered fetch agents into a separate decision table.
Limitation
The final policy follows the brand's legal position and should not be imposed by a template.
Research-led content · proof depth

The evidence is there, buried too deep

Research-led articles usually do have sources, but the claim, the date, the study limits and the action for the reader sit in different places. AI can lift a fragment out of context and lose the point of the recommendation.

Evidence
The content-gap audit flags long paragraphs with no opening claim and a weak link from claim to evidence.
Action
Rewrite sections as claim, argument, proof, limitation — and put a summary box ahead of the deep material.
Limitation
Where primary sources cannot be reached, the block is marked needs review rather than published as fact.
Ownership workflowHow a story becomes a backlog item
Input

URL, intent cluster, model, run date, and the visible fragment of the answer.

Evidence

Mention, citation, absorption, the cited source, a crawler finding or a content gap.

Decision

A specific page, answer block, schema, technical rule or source layer.

Guardrail

Unverified metrics stay N/A; client names are never published without permission.

Can I hide a page from Search using Agent Traffic?

No. Agent Traffic governs crawler access, but robots.txt does not remove a URL from Search — that needs noindex or password protection [76].

Why govern CCBot separately?

Because blocking live bots does not take your content out of the open Common Crawl corpora that other systems use; it is a separate decision [77].

Is one rule "for AI bots" enough?

No. GPTBot, ClaudeBot, PerplexityBot and Googlebot are different agents with different purposes; set the policy for each one [70, 74, 75].

Where should I start?

By inventorying the actual robots.txt and firewall against the OpenAI, Perplexity, Anthropic, Google and Common Crawl documentation — the cheapest check in a GEO audit [76, 74].

Know which agents are reading you.

One access map instead of scattered rules.

Start free