King of Answer Engine

ChatGPT Search

OpenAI runs four separate user agents with four stated purposes. The split lets a publisher be findable without being training data, and it is the clearest crawler documentation of the four vendors here.

What is ChatGPT search?

ChatGPT search is the retrieval layer inside ChatGPT: when a question benefits from current web content, the system searches, reads candidate pages and composes an answer with links to sources. It lives inside a conversation rather than on a results page, which means a single thread can accumulate several searches and the attribution for any one claim can be several turns back.

Which crawlers does OpenAI run?

Four, each documented with a distinct purpose.

OAI-SearchBot exists to surface websites in ChatGPT search results. This is the agent that determines whether a site can appear in a composed answer. A site that disallows it in robots.txt will not be shown in ChatGPT search answers, though OpenAI notes such a site can still appear as a navigational link.

GPTBot crawls content for training generative AI foundation models. Disallowing GPTBot signals that a site's content should not be used in that training. It is used for automatic crawling, not user initiated requests.

OAI-AdsBot validates the safety of landing pages submitted as ChatGPT ads and assesses ad relevance. It visits only pages explicitly submitted as ads, and OpenAI states the data it collects is not used for AI model training.

ChatGPT-User handles user triggered actions, visiting a page or interacting with an external application because a person asked for it. OpenAI states this agent is not governed by robots.txt, on the reasoning that the action originates from a user request rather than automatic crawling, and that it does not determine search eligibility.

What does the four-way split mean in practice?

It means the training question and the visibility question are genuinely separable here. A publisher can allow OAI-SearchBot and disallow GPTBot, and the documented result is a site that can be cited in ChatGPT answers without contributing to foundation model training. For publishers who want attribution but object to training, this is the mechanism.

It also means robots.txt is a weaker instrument than it appears against user triggered reading. If your objective is that a page never be read by ChatGPT under any circumstance, blocking the two crawlers you can block does not achieve that, because ChatGPT-User is explicitly outside robots.txt governance.

How long do robots.txt changes take to apply?

OpenAI states that updates to robots.txt take approximately 24 hours to take effect. This is a small detail with practical consequences: a publisher who changes a directive and checks the same afternoon will draw the wrong conclusion. Allow a day before testing whether a change worked.

What is not publicly known?

OpenAI documents access thoroughly and selection not at all. There is nothing published about how a retrieved page becomes a cited one, how many sources a given answer will carry, or what causes one source to be named over another similar page.

Nor is there documentation of how the search layer decides to search in the first place. Some questions trigger retrieval and some are answered from the model's own weights, and the boundary is not described. For a publisher, that means there is no way to determine which topics are even in play.

One further gap worth naming: the distinction between appearing in a composed answer and appearing as a navigational link is drawn in the documentation but not explained. What determines which treatment a site receives is not stated.

What does the ads crawler imply?

OAI-AdsBot is the newest of the four agents and the only one tied to a commercial placement rather than to organic retrieval. Its documented behavior is narrow: it visits only pages explicitly submitted as ads, it validates that a landing page is safe and assesses relevance, and OpenAI states the data it collects is not used for AI model training.

For a publisher not running ChatGPT ads, it is simply not relevant, and that is worth saying because crawler lists circulate without that context and lead site owners to write rules for agents that will never visit them.

What it does signal is structural. A system that carries advertising alongside composed answers now has an organic retrieval path and a paid placement path operating on the same surface, which is the arrangement search engines settled into two decades ago. Publishers who lived through that transition on the web will recognize the shape, and the questions it eventually raises about how the two are distinguished for a reader are the same questions, arriving again.

How should a publisher configure robots.txt here?

By deciding two things separately and writing them separately.

Whether you want to be findable inside ChatGPT answers is a decision about OAI-SearchBot. Whether you want your content used to train foundation models is a decision about GPTBot. They are independent, and the most common configuration among publishers who have thought about it is to allow the first and disallow the second.

What neither decision controls is ChatGPT-User, which OpenAI documents as outside robots.txt governance. Write your rules understanding that the page can still be fetched when a person asks for it, and allow about a day before testing whether a change has taken effect.

Sources

  1. OpenAI bots and crawlers documentation, OpenAI developer documentation. Retrieved October 6, 2026. Source for the four user agents and their stated purposes, the robots.txt controls, the approximately 24 hour propagation delay, and the navigational link behavior for opted-out sites.