← GEO Playbook·GEO Playbook · Chapter 05

Why AI visibility is unstable — and how to measure it properly

AI answers and the set of cited sources change between runs. A single measurement paints a false picture. This chapter sets the logic of repeated measurement.

6 min read

Checking one prompt once is not your position in AI search. The same query on different days returns different answers and different sources. Without repeated measurement there is no way to tell signal from noise.

One measurement misleads

Visibility in AI search is probabilistic, not fixed. Schulte shows that the answers of AI search engines and their sets of cited sources are unstable between runs, so a single measurement can misrepresent a brand's position [6]. The conclusion is direct: monitoring is repeated measurement, not a one-off prompt check.

Why the direct feedback loop disappeared

In classic web search, clicks and dwell time can serve as fine-grained feedback for improving ranking models; in generative search the feedback is often attached to the final answer, which makes it harder to map onto the retrieval and generation stages. Dai et al. propose restoring fine-grained feedback at the decomposition, retrieval and generation stages [20]. So visibility depends not only on the document but on behaviour at the stages of AI search.

What exactly to measure

If the result is unstable, the metric has to be a distribution rather than a point.

Several runs of the same prompt

Answers change between runs [6].

Several days and time windows

Visibility drifts over time [6].

Several platforms

The source sets differ between systems.

A set of prompts, not one

A single query does not represent an intent cluster.

Metrics borrowed from RAG evaluation

What deserves evaluation is not only the fact of a mention but the quality of source use. A survey of RAG evaluation gives the dimensions: relevance, accuracy, claim support, robustness [36]. RAGAS formalizes faithfulness, answer relevancy and context relevance for RAG pipelines [44]. They can be adapted for a GEO dashboard, but they are not a ready-made standard: the dashboard has to record platform, prompt set, run count and citation/absorption separately.

One piece of content, different systems

The same page behaves differently across products. The Chen et al. benchmark shows that models use identical context unevenly: accuracy and failure modes differ between them [45]. So share of voice is computed per platform first; an aggregate figure is acceptable only as a secondary layer on top of that breakdown, never instead of it.

Sources (E-E-A-T)

06 · Schulte, 2026, arXiv

Answers and citations are unstable between runs. arxiv.org/abs/2604.07585

20 · Dai et al., 2025, SIGIR

Feedback in generative search has weakened. arxiv.org/abs/2505.14680

36 · Yu et al., 2024, arXiv

A set of dimensions for evaluating RAG. arxiv.org/abs/2405.07437

44 · Es et al., 2024, EACL Demo

RAGAS: faithfulness, relevancy, context relevance. aclanthology.org/2024.eacl-demo.16/

45 · Chen et al., 2024, AAAI

Models use the same context differently. arxiv.org/abs/2309.01431

Frequently asked questions

Why is the AI answer about my brand different every time?

Because answers and the set of cited sources are unstable between runs — a property of generative search, not a bug [6].

How many times should visibility be checked?

Many: several runs of one prompt, across several days and several platforms. A single measurement is methodologically weak and can mislead; statistical reliability comes from the repetition protocol — run count, windows, variance [6].

Why can visibility not be averaged "across all AI"?

Because different systems use the same content differently; accuracy and source sets vary between models [45]. Share of voice is computed per platform.

Which metrics belong on a GEO dashboard?

Faithfulness, answer relevancy and context relevance from RAG evaluation — adapted, not treated as a finished standard; computed over repeated measurements, recording platform, prompt set, run count and citation/absorption [36, 44].