What to target is a group of questions and an intent, not an individual key. This is a practical strategy derived from SEO and IR research, not a proven formula.
One keyword is not one intent: people phrase queries differently. Checking a single phrase does not represent the whole cluster of questions. Without clusters there is no systematic way to measure intent coverage in AI answers.
Query strategy is a long game. Erdmann et al. study the long-term strategy of keyword choice in SEO and its effect on competitiveness and traffic stability [61]. Carried across to AEO and GEO: optimize for groups of questions and intents — prompt and query clusters — rather than individual keys. A practical strategy, not a guarantee of outcome.
AEO grew out of QA logic. Natural Questions is built on real user questions paired with Wikipedia documents [30]. Phrase cluster members as natural questions — "how do I…", "which is better…", "X versus Y" — because these systems look for the fragment that answers directly.
Retrieval quality cannot be measured on one keyword group. BEIR is a benchmark for evaluating retrieval models across tasks and domains, and it shows that retrieval quality does not reduce to a single dataset [28]. A cluster has to cover different intent types rather than one narrow group.
"best tool for…" — high intent, high stakes.
"how do I solve…" — catches solution-seekers before they know the category name.
"X versus Y", "alternatives to Y" — shows positioning against the market.
Direct questions about the brand — accuracy and tone control, not just coverage.
ORCAS describes a large set of aggregated query-document click pairs and shows the role of clicks in training and evaluating classic systems [64]. In generative search, clicks on documents fall and feedback becomes less direct — so a cluster is measured by presence in the answer, not by clicks alone.
LLM answers are not equally useful everywhere. Caramancion studies when users prefer an LLM answer and when they prefer classic search [15]. AEO has the greatest effect on complex, comparative and explanatory questions, where the user wants a finished answer.
The long-term strategy of keyword choice.
BEIR: retrieval quality across tasks and domains. arxiv.org/abs/2104.08663
When users choose an LLM over search. arxiv.org/abs/2401.05761
Natural Questions: real user questions.
ORCAS: the role of clicks in training and evaluating search. arxiv.org/abs/2006.05324
A group of queries and phrasings that express one intent ("best X", "X versus Y", "how to do X"). You target the cluster, not the individual key [61].
Because one phrase does not represent an intent cluster, and retrieval quality depends on query type and domain [28].
As natural user questions — answer engines look for the fragment that answers directly, and QA datasets are built on real questions [30].
On complex, comparative and explanatory queries, where the user wants a finished answer rather than a list of links [15].