AI answers and the set of cited sources change between runs. A single measurement paints a false picture. This chapter sets the logic of repeated measurement.
Checking one prompt once is not your position in AI search. The same query on different days returns different answers and different sources. Without repeated measurement there is no way to tell signal from noise.
Visibility in AI search is probabilistic, not fixed. Schulte shows that the answers of AI search engines and their sets of cited sources are unstable between runs, so a single measurement can misrepresent a brand's position [6]. The conclusion is direct: monitoring is repeated measurement, not a one-off prompt check.
In classic web search, clicks and dwell time can serve as fine-grained feedback for improving ranking models; in generative search the feedback is often attached to the final answer, which makes it harder to map onto the retrieval and generation stages. Dai et al. propose restoring fine-grained feedback at the decomposition, retrieval and generation stages [20]. So visibility depends not only on the document but on behaviour at the stages of AI search.
If the result is unstable, the metric has to be a distribution rather than a point.
What deserves evaluation is not only the fact of a mention but the quality of source use. A survey of RAG evaluation gives the dimensions: relevance, accuracy, claim support, robustness [36]. RAGAS formalizes faithfulness, answer relevancy and context relevance for RAG pipelines [44]. They can be adapted for a GEO dashboard, but they are not a ready-made standard: the dashboard has to record platform, prompt set, run count and citation/absorption separately.
The same page behaves differently across products. The Chen et al. benchmark shows that models use identical context unevenly: accuracy and failure modes differ between them [45]. So share of voice is computed per platform first; an aggregate figure is acceptable only as a secondary layer on top of that breakdown, never instead of it.
Answers and citations are unstable between runs. arxiv.org/abs/2604.07585
Feedback in generative search has weakened. arxiv.org/abs/2505.14680
A set of dimensions for evaluating RAG. arxiv.org/abs/2405.07437
RAGAS: faithfulness, relevancy, context relevance. aclanthology.org/2024.eacl-demo.16/
Models use the same context differently. arxiv.org/abs/2309.01431
Because answers and the set of cited sources are unstable between runs — a property of generative search, not a bug [6].
Many: several runs of one prompt, across several days and several platforms. A single measurement is methodologically weak and can mislead; statistical reliability comes from the repetition protocol — run count, windows, variance [6].
Because different systems use the same content differently; accuracy and source sets vary between models [45]. Share of voice is computed per platform.