Insights

From "where are we cited"
to "what do we do about it".

Insights adapts the evaluation metrics of RAG systems — faithfulness, relevance, robustness — into a diagnosis of your content, and turns unstable visibility measurements into a prioritized list of fixes.

Raw measurements do not give you "why"

Metrics show what is happening, not why, and not what to change first.

Feedback only on the final answer

In generative search the signal often arrives only for the final answer, not for the retrieval and generation stages.

Blind fixes

Without a diagnosis, content gets edited on instinct rather than at the bottleneck.

Insights explains why a recommendation matters right now

The Insights page has to show more than "we found an insight": the path from evidence to decision — which prompt failed, which source won, which content is missing, and how to verify the fix.

Short AI answer

Insights is the interpretation layer of a GEO audit: it connects prompt evidence, citation context, technical findings and content gaps to a prioritized backlog. The goal is to explain the next best action, not to replace an editorial or engineering decision [20, 36, 44].

Scenario

A recommendation without evidence reads as an opinion

"Add an FAQ" is useless if nobody can see which prompt failed and why an FAQ would fix it. Insights has to put the evidence next to every action.

Evidence
Each action item carries a prompt id, source context, the affected intent and the technical finding, if any.
Owner action
Make every recommendation evidential: evidence refs, confidence, impact, effort, owner.
Limit of the conclusion
With no evidence refs, a recommendation belongs in "needs review", not in high priority.
Priority

Not every gap is worth closing first

A gap in brand facts and a gap in methodology can look identical and land very differently. Priority only appears once intent, model, frequency and effort are considered together.

Evidence
Ranking uses severity, impact, confidence, effort and affected_models.
Owner action
Show why one task outranks another: which rows support it and what outcome is expected.
Limit of the conclusion
Without a business goal the scoring stays technical rather than revenue-based.
Ownership

An insight has to be verifiable after publication

A good insight ends in an acceptance prompt: what to ask the models again, which source should appear, and which answer counts as an improvement.

Evidence
The open loop closes only after a repeat run, not when the text is merged.
Owner action
Store a validation prompt and a re-check date for every task.
Limit of the conclusion
AI answers are unstable, so verification watches the direction of the signal rather than one perfect answer.
Insights · from signal to prioritySchematic
1Input: your content and your bot access policy
2Processing: preparation / extraction / measurement
3Signal: citations, absorption, gaps
4Output: prioritized actions (with no guarantee of inclusion)
Illustrative process schematic, not real data.
01

Break it down by stage

Visibility is decomposed across decomposition/retrieval/generation: where the influence is lost [20].

02

RAG quality axes

Faithfulness, relevance, robustness — adapted for GEO diagnosis [36].

03

Automated evaluation

The RAGAS logic: how well the content is used in the answer [44].

04

Variance as a signal

High instability is a proxy worth investigating, not a proven cause [6].

05

Filter diagnostics

Hypotheses about weak passage fit, to be verified [46].

What the product includes.

🪜

Stage Breakdown

At which stage the influence is lost.

Read more →
📐

Quality Axes

Faithfulness / relevance / robustness.

Read more →

Variance Signal

What is unstable, and why that is a priority.

Read more →
🧪

Filter Diagnostics

Hypotheses about weak passage fit.

Read more →
06 · Schulte, 2026
Instability as a diagnostic signal.
↗
20 · Dai et al., 2025, ACM SIGIR
Stage-level feedback in AI search.
↗
36 · Yu et al., 2024
A survey of RAG evaluation metrics.
↗
44 · Es et al., 2024, EACL Demo
RAGAS: automated RAG metrics.
↗
46 · Yan et al., 2024
CRAG: correcting retrieval errors.
↗

Stories from audits, not just theory

What we put into the product is not abstract SEO advice but the patterns that keep recurring in GEO audits: what the model saw, which source it cited, where the brand went missing, and which page has to be rewritten.

Cited definition

An Enigma field note is an anonymized audit story containing a checkable signal, the action the page owner took, and the limits of the data. The format helps people and AI understand not only what to do, but why the recommendation appeared at all.

B2B SaaS · brand visibility

The brand is mentioned, but the source belongs to someone else

In a typical B2B SaaS audit the brand shows up for a direct branded query, while comparison and buying-intent prompts lean on review sites or competitor pages. The problem is not the brand name — it is the absence of a cited methodology and a comparison of its own.

Evidence
The prompt matrix shows mentions without an owned citation; citation rows point to a third-party source as the basis of the answer.
Action
Publish a methodology page, comparison answer blocks, and an FAQ covering selection criteria.
Limitation
Client names and uplift figures are not shown without confirmed permission; impact stays N/A.
E-commerce · category intent

AI picks directories over your category pages

In e-commerce the model often cites marketplace directories, because the category page carries no short block on selection criteria, availability, returns and alternatives. The page exists — it just offers no answer-ready evidence.

Evidence
Citation context resolves to an external directory; the brand category is in the sitemap but does not cover source-authority intent.
Action
Add a buyer guide, a criteria table, a shipping and returns block, and an FAQ that matches the visible content.
Limitation
Demand estimates made without a live SEO export are labelled Estimated, never Measured.
Agency · multi-client governance

The content team and the technical team pull different levers

A recurring agency problem: an editor adds FAQs and comparisons while robots.txt, the CDN or a WAF quietly block some AI and user-requested agents. The content is citation-ready on paper and not always reachable in practice.

Evidence
The crawler checklist records the gap between the intended access policy and the rules bots actually meet.
Action
Split training crawler policy, retrieval/search access and user-triggered fetch agents into a separate decision table.
Limitation
The final policy follows the brand's legal position and should not be imposed by a template.
Research-led content · proof depth

The evidence is there, buried too deep

Research-led articles usually do have sources, but the claim, the date, the study limits and the action for the reader sit in different places. AI can lift a fragment out of context and lose the point of the recommendation.

Evidence
The content-gap audit flags long paragraphs with no opening claim and a weak link from claim to evidence.
Action
Rewrite sections as claim, argument, proof, limitation — and put a summary box ahead of the deep material.
Limitation
Where primary sources cannot be reached, the block is marked needs review rather than published as fact.
Ownership workflowHow a story becomes a backlog item
Input

URL, intent cluster, model, run date, and the visible fragment of the answer.

Evidence

Mention, citation, absorption, the cited source, a crawler finding or a content gap.

Decision

A specific page, answer block, schema, technical rule or source layer.

Guardrail

Unverified metrics stay N/A; client names are never published without permission.

How does Insights differ from Monitoring?

Monitoring measures visibility — what and where. Insights explains cause and priority: at which stage the influence is lost and which passage to rework first [20, 36].

Where do the quality metrics come from?

They adapt the RAG evaluation axes — faithfulness, relevance, robustness (RAGAS and the RAG evaluation surveys) — rather than an opaque in-house score [44, 36].

Why look at instability separately?

Because high variance can indicate unstable relevance or weak passage fit — a proxy signal and a candidate for investigation, not a proven cause [6].

Does Insights guarantee more citations?

No. Insights diagnoses bottlenecks; the fixes raise the odds, but inclusion stays probabilistic and depends on the system [46].

Turn the signal into priorities.

Bottleneck diagnosis — with no promises of inclusion.

Start free