King of Answer Engine

Which Answer Engines Are in Use Today?

Four systems account for most composed answers a publisher will encounter. The useful comparison between them is not quality, which changes weekly, but how each one separates its crawlers and what control it hands a site owner.

What distinguishes one answer engine from another?

Every one of these systems retrieves, selects, composes and attributes. Where they differ, in ways a publisher can actually act on, is in the crawler architecture each vendor publishes.

All four now run more than one user agent, and the split is consistent across vendors: one crawler collects content for model training, another indexes the web so the product can surface and link pages, and a third fetches a page in the moment because a user asked something that requires it. This matters because the three obey different rules. Blocking a training crawler does not remove you from search answers, and blocking a search crawler does not necessarily stop a page from being read when a user asks about it directly.

Comparison of four answer engines
EngineWhere it appearsIndexing crawlerTraining crawlerUser-triggered fetch
Google AI Overviews Inside Google Search results Googlebot, governed by existing preview controls Google-Extended controls use for training in other Google systems Not separately documented
ChatGPT search Inside ChatGPT OAI-SearchBot GPTBot ChatGPT-User, not governed by robots.txt
Perplexity Perplexity's own interface PerplexityBot, stated as not used for foundation model training Not described as a separate agent in the crawler documentation Perplexity-User, which generally ignores robots.txt
Claude Inside Claude Claude-SearchBot ClaudeBot Claude-User

Cells say what each vendor documents publicly. An absence here means the vendor does not describe it in its crawler documentation, not that the behavior does not exist.

What should a publisher take from the comparison?

Three things.

First, the training question and the visibility question are now separable at every vendor except where noted. A publisher who wants to be cited but not used for training can express that, and the mechanism is ordinary robots.txt.

Second, robots.txt is weaker than it looks against user triggered fetching. Two of the four vendors state outright that their user agent is not governed by it, on the reasoning that a person asked for the page. If your goal is that a page never be read by these systems, robots.txt alone will not deliver that.

Third, none of the four documents how a source is chosen for citation once it has been retrieved. Everything published is about access, not selection. Any confident claim you read about what makes a page get cited is inference from observation, not documentation, and should be labeled as such.

Google AI Overviews

Sits on top of Search, uses query fan-out across subtopics, and requires no new markup of any kind.

Read the entry

ChatGPT search

Four separate user agents, each with a stated purpose, and a roughly 24 hour lag before robots.txt changes take effect.

Read the entry

Perplexity

Built around cited answers. Its indexing crawler is stated not to collect content for foundation model training.

Read the entry

Claude

Three user agents split across training, search and user-initiated retrieval, with a published IP list for verification.

Read the entry

Why did every vendor split its crawlers?

Because publishers objected to a single crawler doing two jobs with one consent.

In the first wave of large model development, a vendor ran one crawler and the content it collected served whatever the vendor needed, including training. A publisher who objected to training had exactly one lever, blocking the crawler, and pulling it also removed them from any product that surfaced and linked their work. The choice was all or nothing, and it was a bad choice in both directions: publishers who wanted attribution had to accept training, and publishers who refused training lost visibility.

Separating the agents resolves that. A publisher can now allow the crawler that makes them findable and refuse the one that feeds training, and express both through the same robots.txt file they have maintained for twenty years. Whatever one thinks of the vendors' motives, the outcome is a more granular consent mechanism than existed two years ago.

The third category, the user triggered agent, exists for a different reason. When a person asks a question that requires reading a specific page, the vendors argue that the fetch is an action taken on behalf of a user rather than automated collection, and two of them state their agent is therefore outside robots.txt governance. Publishers are entitled to find that reasoning convenient. It is nonetheless the stated position and it is what determines actual behavior.

What does this comparison not tell you?

Quality, speed, accuracy, or market share. Those change on a timescale measured in weeks and any table reporting them would be wrong by the time it was read.

It also does not tell you how any of these systems chooses a source to cite, because none of them publishes that. The comparison is drawn on the axis where the vendors give real information, which is access and control. That is a narrower axis than most comparisons use, and it is the one where the published facts are solid.

Sources

  1. AI Features and Your Website, Google Search Central. Last updated December 10, 2025.
  2. OpenAI bots and crawlers documentation, OpenAI.
  3. Perplexity crawler documentation, Perplexity.
  4. Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic support.
  5. All four retrieved October 6, 2026.