Site Maps

A map of your content
as retrieval sees it.

Site Maps builds a probabilistic map from test retrieval runs: which passages get extracted for which kinds of query, where the substantive gaps are, and where structure gets in the way of machine understanding — across domains and intents.

Test across different queries

Source visibility has to be checked across query types and domains, not on one keyword group.

Indexed ≠ extracted

Content can sit in the index and still never be extracted as a passage for the questions people actually ask.

Not "chop up the content"

Google requires no chunking, and a page can be understood whole — the map only shows where structure obstructs extraction.

Site Maps shows a map of extractable meaning, not sitemap.xml

The reader needs the difference between "the page exists" and "the page answers the intent". So this page shows how an intent turns into a checkable passage backlog.

Short AI answer

Site Maps is an intent-to-passage map: it shows which pages and blocks can be extracted for brand, category, comparison, buying, problem and source-authority queries. It is a diagnostic map, not a guarantee of AI inclusion [28, 35, 45].

Scenario

The page is in the index, the answer to the question is not

The team sees the URL in the sitemap and considers the topic covered. The audit says otherwise: for buying-intent questions the page lists features but gives no selection criterion and never explains how it differs from the alternatives.

Evidence
Prompt rows cover the intent cluster while the extracted passage contains no direct answer.
Owner action
Build the backlog: which intent is uncovered, which block to rewrite, which source to add.
Limit of the conclusion
Without live retrieval runs the map stays an expert hypothesis and must be labelled Estimated.
Content gap

The gap can be a format, not a topic

Sometimes the topic is covered but the format is missing: a comparison, a methodology, a case, a criteria table, an FAQ. To an AI answer, a missing format looks like missing evidence.

Evidence
The content-gap contract asks you to look at guides, comparisons, case studies, tools, templates and research assets.
Owner action
Assign a format to every intent: answer block, comparison table, proof note, FAQ or methodology page.
Limit of the conclusion
A competitor-relative gap cannot be claimed without verified competitor URLs; otherwise the finding stays an internal gap audit.
Ownership

The map has to become an editorial route

A useful map does not end at a pretty diagram. It says which page to touch first, which block to add, and which prompt verifies the result after publication.

Evidence
A backlog item carries URL, intent, current passage, missing proof and an acceptance prompt.
Owner action
Tie the map to the update calendar and to the next audit run.
Limit of the conclusion
AI visibility is probabilistic; the map helps you prioritize, it does not promise a stable place in the answer.
Site Maps · the "query → passage" mapSchematic
1Input: your content and your bot access policy
2Processing: preparation / extraction / measurement
3Signal: citations, absorption, gaps
4Output: prioritized actions (with no guarantee of inclusion)
Illustrative process schematic, not real data.
01

Heterogeneous evaluation

An extraction map from probe runs across query types — probabilistic, not absolute knowledge [28].

02

Model variability

The same content gets used differently by different models [45].

03

Structure matters

Highlight the blocks where the claim is buried in the middle [35].

04

Adaptation, not manipulation

Cooperative adaptation of structure and completeness, not gaming [3].

05

The auditor’s view

Where you are the source, where you get bypassed, and why [12].

What the product includes.

Coverage Map

Which queries reach your content.

Read more →

Passage Fit

Which block is extracted for which intent.

Read more →
🔎

Gap Finder

Where the substantive gap against competitors is.

Read more →
🔧

Adaptation

What to restructure without manipulation.

Read more →
03 · Wu et al., 2026, ICLR
AutoGEO: cooperative content adaptation.
↗
12 · Venkit et al., 2025, ACM FAccT
A qualitative audit of LLM-based search.
↗
28 · Thakur et al., 2021, NeurIPS
BEIR: heterogeneous retrieval evaluation.
↗
35 · Liu et al., 2024, TACL
Lost in the middle: position within the context.
↗
45 · Chen et al., 2024, AAAI
RAG benchmark: model variability.
↗

Stories from audits, not just theory

What we put into the product is not abstract SEO advice but the patterns that keep recurring in GEO audits: what the model saw, which source it cited, where the brand went missing, and which page has to be rewritten.

Cited definition

An Enigma field note is an anonymized audit story containing a checkable signal, the action the page owner took, and the limits of the data. The format helps people and AI understand not only what to do, but why the recommendation appeared at all.

B2B SaaS · brand visibility

The brand is mentioned, but the source belongs to someone else

In a typical B2B SaaS audit the brand shows up for a direct branded query, while comparison and buying-intent prompts lean on review sites or competitor pages. The problem is not the brand name — it is the absence of a cited methodology and a comparison of its own.

Evidence
The prompt matrix shows mentions without an owned citation; citation rows point to a third-party source as the basis of the answer.
Action
Publish a methodology page, comparison answer blocks, and an FAQ covering selection criteria.
Limitation
Client names and uplift figures are not shown without confirmed permission; impact stays N/A.
E-commerce · category intent

AI picks directories over your category pages

In e-commerce the model often cites marketplace directories, because the category page carries no short block on selection criteria, availability, returns and alternatives. The page exists — it just offers no answer-ready evidence.

Evidence
Citation context resolves to an external directory; the brand category is in the sitemap but does not cover source-authority intent.
Action
Add a buyer guide, a criteria table, a shipping and returns block, and an FAQ that matches the visible content.
Limitation
Demand estimates made without a live SEO export are labelled Estimated, never Measured.
Agency · multi-client governance

The content team and the technical team pull different levers

A recurring agency problem: an editor adds FAQs and comparisons while robots.txt, the CDN or a WAF quietly block some AI and user-requested agents. The content is citation-ready on paper and not always reachable in practice.

Evidence
The crawler checklist records the gap between the intended access policy and the rules bots actually meet.
Action
Split training crawler policy, retrieval/search access and user-triggered fetch agents into a separate decision table.
Limitation
The final policy follows the brand's legal position and should not be imposed by a template.
Research-led content · proof depth

The evidence is there, buried too deep

Research-led articles usually do have sources, but the claim, the date, the study limits and the action for the reader sit in different places. AI can lift a fragment out of context and lose the point of the recommendation.

Evidence
The content-gap audit flags long paragraphs with no opening claim and a weak link from claim to evidence.
Action
Rewrite sections as claim, argument, proof, limitation — and put a summary box ahead of the deep material.
Limitation
Where primary sources cannot be reached, the block is marked needs review rather than published as fact.
Ownership workflowHow a story becomes a backlog item
Input

URL, intent cluster, model, run date, and the visible fragment of the answer.

Evidence

Mention, citation, absorption, the cited source, a crawler finding or a content gap.

Decision

A specific page, answer block, schema, technical rule or source layer.

Guardrail

Unverified metrics stay N/A; client names are never published without permission.

What does Site Maps show?

A probabilistic "query → passage" map built from probe runs: which blocks of the site get extracted for which query types and intents, and where the substantive gaps are [28, 12].

How is this different from sitemap.xml?

sitemap.xml lists URLs for a crawler. Site Maps estimates, from test runs, which passages a retrieval system extracts — a different unit of analysis, and probabilistic rather than absolute [28, 45].

Is this gaming the algorithm?

No. Site Maps follows the logic of cooperative adaptation of structure and completeness (AutoGEO) rather than manipulation; gaming raises your exposure to anti-spam filters [3].

Why does a passage work for one query and not another?

Because different models use the same context differently; visibility has to be tested across heterogeneous tasks [45, 28].

See your site the way retrieval sees it.

A probabilistic extraction map, not an absolute promise.

Start free