← Research Lab·Research Lab · Instability

AI answer instability: why one measurement misleads

The answers of AI search engines and their sets of cited sources change between runs. Two papers examined for why GEO monitoring has to be a repeatable process.

May 2026·5 min read

"I checked one prompt, the brand is not in AI" is common and mistaken logic. If the answer changes between runs, a one-off measurement does not describe the real position. Without understanding the sources of instability there is no way to build a correct visibility metric.

Answers and citations are unstable between runs

Schulte shows directly that the answers of AI search engines and their sets of cited sources are unstable between runs, so a single visibility measurement can misrepresent a brand's position [6]. That is the central argument against a one-off "prompt check".

The methodological consequence

From source 06 follows a methodological conclusion: a reliable assessment of visibility needs a repeatable measurement design — prompt sets, several runs, time windows, probabilistic metrics, rather than a single observation. That is a well-founded methodological conclusion from the research, not the one mandatory standard for every product.

Why feedback complicates the picture

Dai et al. (NExT-Search) describe how, in classic search, user behaviour improves ranking, while in generative search the feedback often attaches only to the final answer [20]. A less direct feedback loop is one reason visibility is harder to capture in a single snapshot.

How this affects metrics

One prompt, one run

An instantaneous snapshot — the risk of a false "brand absent" [6].

Repeated measurement

A distribution of presence — more expensive, and correct.

Only the final answer

An end result with no signal from the retrieval and generation stages [20].

Sources (E-E-A-T)

06 · Schulte, 2026

Instability of answers and citations between runs. arxiv.org/abs/2604.07585

20 · Dai et al., 2025, ACM SIGIR

A less direct feedback loop in generative search. arxiv.org/abs/2505.14680

Frequently asked questions

Why can visibility not be checked with one prompt?

Because answers and cited sources are unstable between runs; one measurement can falsely show the brand as absent [6].

How should it be measured instead?

Repeated measurement: prompt sets, several runs, time windows, probabilistic metrics — a methodological conclusion from the GEO measurement research, not the one mandatory standard [6].

Why is generative search harder to pin down?

Because feedback often attaches only to the final answer rather than to the retrieval and generation stages — the loop is less direct [20].

Is instability a platform bug?

It is a property of generative systems, to be built into the measurement methodology rather than ignored [6].