The answers of AI search engines and their sets of cited sources change between runs. Two papers examined for why GEO monitoring has to be a repeatable process.
"I checked one prompt, the brand is not in AI" is common and mistaken logic. If the answer changes between runs, a one-off measurement does not describe the real position. Without understanding the sources of instability there is no way to build a correct visibility metric.
Schulte shows directly that the answers of AI search engines and their sets of cited sources are unstable between runs, so a single visibility measurement can misrepresent a brand's position [6]. That is the central argument against a one-off "prompt check".
From source 06 follows a methodological conclusion: a reliable assessment of visibility needs a repeatable measurement design — prompt sets, several runs, time windows, probabilistic metrics, rather than a single observation. That is a well-founded methodological conclusion from the research, not the one mandatory standard for every product.
Dai et al. (NExT-Search) describe how, in classic search, user behaviour improves ranking, while in generative search the feedback often attaches only to the final answer [20]. A less direct feedback loop is one reason visibility is harder to capture in a single snapshot.
Instability of answers and citations between runs. arxiv.org/abs/2604.07585
A less direct feedback loop in generative search. arxiv.org/abs/2505.14680
Because answers and cited sources are unstable between runs; one measurement can falsely show the brand as absent [6].
Repeated measurement: prompt sets, several runs, time windows, probabilistic metrics — a methodological conclusion from the GEO measurement research, not the one mandatory standard [6].
Because feedback often attaches only to the final answer rather than to the retrieval and generation stages — the loop is less direct [20].
It is a property of generative systems, to be built into the measurement methodology rather than ignored [6].