A link under a ChatGPT or Perplexity answer is not proof that your text shaped it. This piece draws the line between "we were cited" and "we were actually used".
Marketers and content teams measure success in AI search with a single number: how many times the domain landed in the citation list. The trouble is that the number says almost nothing about what the page contributed. Concretely:
The page sits in the source list while the answer text carries none of your phrasing and none of your facts.
The content budget gets optimized for "landing in the citations" rather than for shaping the answer.
You cannot explain to leadership why growing citation counts fail to move brand awareness.
A citation in an AI answer is not one process but a chain of three independent events. Failure or success at one stage does not predict the others.
Claim: choosing a source is a retriever decision, not evidence of its usefulness. Argument: the model first runs a search and picks pages by semantic proximity to the query, and only then decides what to do with them. Evidence: in From Citation Selection to Citation Absorption (Zhang, He, Yao, 2026, preprint) selection is defined as the stage where "a platform triggers search and chooses sources", separately from whether the source reaches the answer [7].
Claim: a link pinned to a sentence is a claim by the model, not a verified fact. Argument: the LLM generates the marker as part of the text, and that marker may not support the statement it sits beside. Evidence: the ALCE benchmark (Gao, Yen, Yu, Chen, 2023, peer-reviewed, EMNLP) found that "on the ELI5 dataset, even the best models lack complete citation support 50% of the time" — half the links do not back what is claimed [32].
On the ELI5 dataset even the best models lack complete citation support 50% of the time — half the links do not back the claim [32].
Claim: influence is when the page gives the answer its language, facts, structure or evidence. Argument: probable absorption is judged by the source's contribution to the final text, not by the presence of a link. Evidence: the same preprint defines absorption as the case where "a cited page contributes language, evidence, structure, or factual support to the final answer" — a metric distinct from the fact of citation [7].
One important limit: measuring absorption directly requires access to the model's internal signals (as in the MIRAGE approach), which closed platforms do not offer. External tools estimate probable absorption indirectly — through overlap of language, facts and structure between answer and source — not as certain knowledge of the internal mechanism.
Claim: a platform can cite broadly and absorb narrowly. Argument: engines balance citation breadth and depth differently. Evidence: across 602 controlled prompts and 21,143 valid search-layer citations on three platforms, Zhang et al. found that Perplexity and Google AI Overview cite more sources on average, while "ChatGPT cites fewer sources but shows substantially higher average citation influence among fetched pages" [7].
Claim: a correct citation is not the same as a faithful one. Argument: a link can point to a source that genuinely contains the fact while the model arrived at the answer by another route. That is the correctness versus faithfulness distinction argued in Correctness is not Faithfulness in RAG Attributions (Wallat et al., 2024/2025, peer-reviewed, ACM Digital Library): a correct attribution does not establish that the source was actually used during generation [39]. Note: the full text of this source was unavailable at the time of writing; only the conceptual distinction stated in its title is reproduced here, with no figures.
Claim: to learn the real influence you have to look inside the model rather than at its links. Argument: attribution from internal states is decoupled from what the model "decided" to cite. Evidence: the MIRAGE method (Qi et al., 2024, peer-reviewed, EMNLP) "detects context-sensitive answer tokens and pairs them with retrieved documents contributing to their prediction via saliency methods" — precisely because self-citing LLMs "fail to faithfully reflect LLMs' context usage throughout the generation" [40].
Claim: what gets absorbed is not the "optimized" page but the extractable one. Argument: models pull specific units of meaning into an answer, not general relevance. Evidence: Zhang et al. (preprint) identify a cluster of traits in high-influence pages — greater length, clear structure, semantic fit to the query, and density of extractable evidence: definitions, numeric facts, comparisons and step-by-step procedures [7].
The retriever picked the page — presence in the source list — Source 7 (preprint).
The model pinned a marker to a sentence — citation precision/recall, support for the claim — Source 32 (EMNLP), Source 39 (ACM).
The answer text was shaped by the page — contribution of language, facts, structure; internal saliency — Source 7 (preprint), Source 40 (EMNLP).
From Citation Selection to Citation Absorption — preprint (arXiv) — introduces the selection/absorption distinction; 602 prompts, 21,143 citations, 3 platforms [7].
Enabling LLMs to Generate Text with Citations — peer-reviewed (EMNLP) — the ALCE benchmark; on ELI5 no complete citation support 50% of the time [32].
Correctness is not Faithfulness in RAG Attributions — peer-reviewed (ACM DL) — the correctness versus faithfulness distinction (full text unavailable) [39].
Model Internals-based Answer Attribution / MIRAGE — peer-reviewed (EMNLP) — attribution from internal model states instead of self-citation [40].
No. That is citation selection — the retriever picked the page. Influence (absorption) is a separate event in which the page's text actually reached the answer. Without access to the model's internal signals, probable absorption can only be estimated indirectly.
The citation marker is part of the generated text, not a verified fact. On the ALCE benchmark, even the best models lack complete citation support 50% of the time on ELI5.
Through attribution from the model's internal states. MIRAGE pairs answer tokens with the documents that actually contributed to their prediction, bypassing the model's unreliable self-citation.
Length, clear structure, semantic fit to the query, and a high density of extractable evidence: definitions, numbers, comparisons and step-by-step instructions.
The Enigma editorial team.
18 May 2026.
18 May 2026.
Peer-reviewed work, preprints and industry research — the type is stated next to each claim.
Definitions and figures reproduce the sources' own abstracts; where a full text was unavailable, that is stated at the claim.
Preprints and industry reports are cited with their methodological limits; verify the conclusions on your own project.
The full list of sources and their reliability levels lives in the research catalogue.
From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms. https://arxiv.org/abs/2604.25707 — preprint.
Enabling Large Language Models to Generate Text with Citations. https://aclanthology.org/2023.emnlp-main.398/ — peer-reviewed (EMNLP).
Correctness is not Faithfulness in RAG Attributions. https://dl.acm.org/doi/10.1145/3731120.3744592 — peer-reviewed (ACM DL).
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation. https://aclanthology.org/2024.emnlp-main.347/ — peer-reviewed (EMNLP).