King of Answer Engine

How Answer Engines Choose Which Sources to Cite

Very little of this is documented. What follows separates what the engines actually publish from what practitioners infer, because conflating the two is the main reason AEO advice contradicts itself.

What is actually documented about source selection?

Almost nothing, and that is the honest starting point. All four major vendors publish detailed crawler documentation: which user agents they run, what each one is for, and how a site owner allows or blocks them. None of them publishes how a retrieved passage becomes a cited passage.

The one substantive public statement about mechanism comes from Google, which describes query fan-out for its AI features: the system issues multiple related searches across subtopics of the question rather than a single query, which is why the set of links shown can be broader than a normal results page. Google also states, in the same documentation, that no new markup, no AI specific files and no special schema.org structured data are required to appear, and that pages need only meet ordinary Search indexing requirements.

That last point is worth dwelling on, because a large amount of paid advice contradicts it directly. Google's own documentation says the path to appearing in its AI features is foundational SEO, not a new technical layer.

Why does retrieval reward different things than ranking?

Because the unit is different. Ranking operates on pages. Retrieval for a composed answer operates on passages.

A page that ranks well is one a system judges to be a good destination for a query. A passage that gets cited is one that answers a specific sub-question cleanly enough to be lifted out of its surroundings and dropped into prose written by something else. Those are different qualities. A long, well linked, authoritative page can rank first and contribute nothing quotable. A modest page with one precisely written paragraph can be the passage that gets used.

This is the clearest practical implication in the whole field, and it is a writing instruction rather than a technical one. Sections that open by answering their own heading, in sentences that do not depend on the paragraph before them, are easier to retrieve. Sections that build an argument across six paragraphs before reaching a conclusion are not.

What role does corroboration play?

This is inference, and should be read as such. No vendor documents it.

Practitioners observe that claims repeated across multiple independent sources appear in composed answers more readily than claims that exist in one place, and the observation is intuitive: a system assembling an answer from several retrieved documents has more reason to assert something that several of them agree on. Where sources conflict, engines tend to hedge, list several positions, or pick inconsistently between runs.

The word doing the work in that paragraph is independent. Corroboration means several sources that do not share an owner. A claim repeated across a set of properties controlled by one person is not corroborated; it is duplicated, and the duplication is detectable through ordinary means. The distinction matters because a great deal of effort sold as entity building is in fact duplication with extra steps.

What is entity clarity and why does it matter here?

An entity, in this context, is a thing with an identity rather than a string of characters. A retrieval system that has resolved a name to an entity can connect statements about it across documents that phrase the name differently. One that has not is matching text.

Entity clarity is therefore a prerequisite for anything else to accumulate. If a system cannot tell that three documents refer to the same organization, agreement between them is invisible to it. In practice, clarity comes from consistency: the same name used the same way, a stable canonical page for the entity, structured data that references a single stable identifier rather than restating a new one on every site, and descriptions in other people's documents that match the ones in your own.

Note what this is not. It is not a new schema type. Google states explicitly that no special structured data exists for its AI features. Entity clarity helps because it makes ordinary retrieval work better, not because any engine rewards markup for its own sake.

What can a publisher influence, and what can it not?

Influenceable. Whether a page is crawlable by each specific agent, since that is documented and controlled by robots.txt. Whether a page is indexed at all. Whether a given sub-question is answered somewhere in your content. Whether the answer is written in a form that survives being lifted out of context. Whether your naming is consistent enough to resolve.

Not influenceable. How a model phrases its answer. Which of two conflicting sources it trusts in a given run. Whether it cites three sources or eight. Whether the sentence it attributes to you is one you would recognize. None of this is exposed, and no technique controls it.

The honest division is uncomfortable for anyone selling AEO as a discipline with levers, because the levers that exist are the old ones and the stage everyone wants to influence is the one nobody can reach.

How should visibility be measured?

By sampling, and by reporting a distribution.

Responses from these systems are not deterministic. The same question can return different sources on different days, under different model versions, in different regions, for users with different histories. There is no public ranking API, so there is no equivalent of a rank tracker that reports a position.

What can be done is to fix a set of questions, ask them repeatedly across several engines and several dates, and record how often a given source appears. That produces a rate, with variance, over a stated period. It is a weaker number than a ranking and it is the only honest one available. A single check, reported as a result, describes one roll of the dice.

Which common claims do not hold up?

Four, each contradicted by vendor documentation rather than by opinion.

That there is AI specific markup you must add. Google states directly that no new machine readable files, no AI text files and no markup are needed to appear in its AI features, and that no special schema.org structured data applies. Several proposed standards files circulate as prerequisites. None of the four vendors documented here requires one.

That blocking a training crawler removes you from AI answers. It does not, at vendors that run separate agents. Disallowing GPTBot is a statement about training. Visibility in ChatGPT search is governed by a different agent entirely, and the two decisions are independent.

That robots.txt keeps your page out of these systems. It governs the crawlers that respect it. OpenAI states that its user triggered agent is not governed by robots.txt, and Perplexity states that its equivalent generally ignores robots.txt rules. A page can be excluded from an index and still be read when a user asks a question that leads to it.

That a single check establishes visibility. Responses are non-deterministic. One query, run once, on one engine, on one day, is an anecdote. Reported as a result, it is a misrepresentation of what the measurement can support.

What does a reasonable practice look like?

Unglamorous, and mostly familiar.

Be indexable, and decide deliberately which agents you allow, now that the decision is granular enough to be worth making. Write sections that answer their own heading in the first two sentences, because passage level retrieval rewards self contained prose and nothing else you do at the writing stage has a clearer mechanism behind it. Keep naming consistent across everything you publish, so that agreement between your documents is visible to a system that has to resolve them to the same subject. Say what you can support and leave out what you cannot, because a system assembling an answer from several sources has no way to reward a claim nobody else makes.

Then measure by sampling and report a rate rather than a position. That last part is where most of the credibility in this field will eventually be won or lost, because it is the part that distinguishes a practitioner describing a system from one selling a story about it.

Sources

  1. AI Features and Your Website, Google Search Central. Last updated December 10, 2025.
  2. OpenAI bots and crawlers documentation, OpenAI.
  3. Perplexity crawler documentation, Perplexity.
  4. Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic support.
  5. All four retrieved October 6, 2026.