How retrieval actually decides citation
An engine does not read your page and judge it: it retrieves passages from an index, fuses spans that corroborate one another, and attributes the result, which means citation is decided at the passage level before the model composes a sentence.
Key takeaways
- Retrieval happens before generation. A page that cannot be pulled apart into coherent passages is not evaluated and rejected, it is never considered.
- The competitive unit is the section, not the page. Each of your sections competes against a different set of rival sections, and can lose while the rest of the page wins.
- Answers are usually fused from several spans rather than lifted from one sentence, so the density of independently liftable passages matters more than one perfect line.
- Schema is an entity and rich-result lever, not a citation lever. Google states this in its own documentation, and it is worth reading before anyone bills you for markup as an AI play.
Retrieval: what does the engine actually read?
The model does not open your website. A retrieval layer sits in front of it, takes the question, expands it into a set of related sub-questions, and fetches candidate material from an index. What reaches the model is that material, and nothing else about you exists as far as that answer is concerned.
Which index matters, because the engines do not share one. Google's AI Overviews and Gemini ground on Google's own Search index; Google's documentation says the system relies on its core Search ranking systems to retrieve relevant, up to date pages from the Search index. ChatGPT grounds substantially on Bing and its own crawl. Perplexity runs its own retrieval. That is why per-engine results diverge for the same brand, and why the fix for one engine is per-index search work rather than a generic AEO treatment.
There is also a hard eligibility gate before any of this applies. Google states that a page must be indexed and eligible to be shown in Google Search with a snippet, and separately that a site must be included in Search generative AI features in Search Console to be eligible for display. A noindex, a nosnippet directive, a zero-length snippet limit or that setting will remove you regardless of how good the content is. It is the cheapest check in the discipline and it is skipped constantly.
Typical timelines run from two to six weeks depending on scope, with Provider A citing a shorter turnaround for standard work and Provider B noting that permitting can extend it.
illustrative, not a capture Two sources, fused into one sentence, each contributing a span rather than a quotable line. That fusion is why passage density beats one perfect sentence, and why the section is the unit that competes. Written to show the shape of an engine answer. It has no session, no capture date and no sample size, so it is not evidence about any category, including yours.
Extraction: why is the passage the unit, not the page?
What the retrieval layer hands over is fragments: snippets, passages, extracted units. That extraction is chunking, and it happens before the model sees anything. The practical consequence is stark. If a section cannot be pulled out as a coherent unit, it does not lose the comparison. It never enters it.
Three habits make sections unretrievable. Vague headings, which give the layer nothing to match a sub-question against, so a heading like our approach is invisible while a heading shaped as the question a buyer asks is a target. Cross-section dependency, where a paragraph relies on a pronoun or a phrase resolved three sections earlier, so nothing coherent can be lifted without the rest of the page. And the answer arriving after two hundred words of preamble, so the liftable part sits below whatever the extractor took.
The corrective is not to fragment your site into stub pages. Google says directly that there is no requirement to break content into tiny pieces for AI to understand it, and in the same guidance concedes the mechanism, saying its systems understand the nuance of multiple topics on a page and show the relevant piece to users. Both statements are true. The synthesis we work to is normal length pages whose sections are each independently extractable: the claim first, the subject restated rather than referenced, no dependency on what came before.
Corroboration: what settles a tie between sources?
Your content is not assessed as a finished document. It is assessed as a set of claims against what the model already believes from training, what retrieval just handed it, and the source-quality rules it operates under. Winning means the same claim arriving from several places the system already trusts.
That is why agreement is worth more than novelty at the point of citation. A passage that contradicts every other retrieved source is a risk the system routes around; a passage that matches the consensus and adds one specific, sourced detail is safe to attribute and worth attributing. Consensus first, distinctive detail second, in that order.
It is also why the single quotable sentence is the wrong target. Synthesis usually stitches spans together, often two or three passages from different parts of one page and sometimes from different sources entirely. You are not writing a hero line for a machine to lift. You are raising the density of independently liftable spans, close together, saying compatible things, so that any two adjacent spans can fuse into a defensible attribution.
Selection: why is being retrieved not the same as being cited?
Entering the pool of eligible sources and being chosen from it are separate events with separate causes, which is the subject of the essay on composite scores. Selection is stochastic across runs, so the honest unit is a rate over many runs rather than a position. There is no first place in an AI answer, only a probability of being named, and a single capture is one sample from it.
Two things widely sold as citation levers do not decide selection. Google's own guidance says structured data is not required for generative AI search, that there is no special markup you need to add, and that it remains a good idea for the rest of your search work. That settles it in both directions: keep shipping schema for entity clarity and rich results, stop billing it as a citation lever. Our own house rule goes one step further, and it is Google's guideline too: every answer in your markup must exist in the visible copy, so if you cannot point at the sentence on the page it does not go in the markup.
The same applies to llms.txt. Google has said plainly that Search does not use these files and that publishing one will neither harm nor help visibility in Search. It costs almost nothing and other consumers may read it, so ship it if you like. It is a cheap ticket, not a mechanism, and anybody quoting it as a lever is quoting something Google has already denied in writing.
Questions, answered plainly.
Does schema markup get me cited by AI?
Not on the evidence Google publishes. Its guidance says structured data is not required for generative AI search and that there is no special markup to add, while recommending it for the rest of your search work. Ship schema for entity clarity and rich results, and treat any provider billing it as an AI citation lever as quoting something the source has denied.
Should I break my content into small pages so it chunks better?
No. Google says explicitly that there is no requirement to break content into tiny pieces, and separately warns that spawning a page per query variation to influence generative answers falls under its scaled content abuse policy. Keep normal length pages and make each section independently extractable.
Does llms.txt help?
Google has stated that Search does not use these files and that publishing one neither helps nor harms visibility in Search. It is cheap and other consumers may read it, so there is no argument against having one. There is a strong argument against paying for it as a visibility lever.
Why do different engines cite completely different sources?
Because they retrieve from different indexes. Google's generative answers ground on Google's index, ChatGPT grounds substantially on Bing and its own crawl, Perplexity runs its own retrieval. The mechanism is the same everywhere; the pool of candidates is not, which is why engine coverage has to be stated on any report rather than implied.
What is measured, and what this page is not.
This is an explainer. It carries no figures, and it is not a reading of your category. The disclosure below states the instrument that produces the numbers the essay refers to, so the distinction is on the page rather than assumed.
- instrument
- Caul
- what was measured
- Nothing on this page. Where the essay refers to citation share, that figure is produced separately, per account.
- how
- A prompt set written once for a category and then frozen, run against every engine in clean sessions, with each answer stored unmodified.
- over what window
- Reviewed on 2026-08-15. The engines change, so read the essay against the date on the byline.
- what this cannot tell you
- An explainer is not evidence about your category. Being named is not being recommended, and it is not traffic or revenue. Any figure about your own visibility has to come from a capture of your own category, carrying its sample size and its window.
The rest of the cluster.
See where you stand.
The audit is a real sweep of your category, benchmarked against competitors you name, delivered on a call so the findings get explained rather than emailed. You keep the report and the underlying data whatever you decide afterwards.