AI engines retrieve candidate pages, re-rank them, and pick a few to cite based on relevance, authority, and freshness. Understanding that pipeline guides AEO.
When an AI engine answers a question, it does not consult the entire web — it retrieves a set of candidate pages, re-ranks them by relevance and trust, and selects a few to draw on and cite. Understanding that pipeline — retrieval, re-ranking, selection — tells you where AEO effort actually pays off, because each stage rewards different things about your content.
Most AI answers that reference live sources follow a similar path. First the engine retrieves candidate documents related to the query. Then it re-ranks those candidates to find the most relevant and trustworthy. Finally it selects a handful to synthesize into the answer and, often, to cite.
You have to pass all three stages to be cited. Being retrievable but low-ranked, or high-ranked but not selected, both leave you out of the answer.
Thinking in stages clarifies why a single tactic is never enough — different stages care about different things.
Retrieval gathers candidate pages that might answer the query. To be in this set, your page has to be crawlable, indexed, and relevant to the question's topic and language.
If your content is blocked, client-rendered into invisibility, or off-topic for the query, you never enter the candidate pool — and the later stages cannot save you.
From the candidates, the engine ranks by how well each answers the query and how much it can be trusted. Relevance, authority, and freshness all factor in.
A page that answers the specific question directly outranks one that merely mentions the topic. A source the engine considers authoritative outranks an unknown one. And current content often outranks stale content on topics where recency matters.
This is where answer-first writing, entity clarity, off-site authority, and freshness compound — they push you up the ranking among the candidates.
The engine then chooses a few top sources to actually use and cite, aiming for a set that answers the query well without redundancy. Being ranked highly makes selection likely but not guaranteed.
Selection tends to favor sources that add distinct, quotable value — a clear answer, a specific fact, a useful comparison. Content that is easy to extract cleanly is easier to select.
The upshot: it is not enough to be relevant and trusted; you also need to be the clearest, most quotable answer among the finalists.
Each stage maps to work you can do. Fetchability and relevance get you retrieved; answer-first content, entity clarity, authority, and freshness get you re-ranked; clean, extractable, distinctive passages get you selected.
TrueCite tracks how its nine tracked engines answer your prompts and which sources they cite, so you can see where in this pipeline you are winning or dropping out. That tells you whether to fix reachability, authority, or the clarity of the answer itself.
The pipeline is also a diagnostic. When you are not cited, the stage where you drop out tells you what to fix — and fixing the wrong stage wastes effort.
If your content is never retrieved, the problem is reachability or relevance: fetchability, crawler access, and topic match. If it is retrieved but not selected, the problem is ranking or quotability: authority, freshness, and how clearly you state the answer.
So when an answer omits you, ask where in the three stages you fell out. Are you even in the candidate pool? Are you ranked but passed over? Matching your fix to the stage you are losing at is how you improve deliberately instead of applying random tactics and hoping one lands.
AI engines cite through a pipeline: retrieve candidates, re-rank by relevance and trust, select a quotable few. To be cited you must clear all three — be reachable and relevant, be authoritative and fresh, and be the clearest answer worth quoting. Mapping your AEO work to those stages is how you improve deliberately rather than by guesswork.
Using TrueCite? See the How Scans Work docs →