Skip to content
LumiRank
Answers
Answers · ChatGPT

Where does ChatGPT get its information from?

Two different systems produce a ChatGPT answer: a trained model with a fixed cutoff, and live search that crawls and cites pages in real time. What each one means for a business trying to be found.

ChatGPTAI crawlersGEOAI search
A single tall apartment tower shot from its base, with an airliner passing overhead in a clear blue sky

Short answer: ChatGPT answers come from two different systems. One is the underlying model, trained on a large body of text up to a fixed cutoff date, which it draws on when it answers from memory alone. The other is ChatGPT search, which sends out a live crawler, reads current pages, and cites them in the answer. Which one produced a given answer changes both how current it is and what you can do about it.

OpenAI documents the mechanics on its developer site: OAI-SearchBot fetches pages for live search results, GPTBot collects the separate corpus used to train future models, and ChatGPT-User fetches a single page on request, when someone asks ChatGPT to look at something specific or a Custom GPT calls out to a tool. Three different jobs, three different robots.txt tokens. Blocking one does not block the others.

This page covers the mechanism itself. Getting the model to recommend a business once it can see it, and correcting something it already has wrong, are separate questions, covered on their own pages.

At a glance

Training data

The base model learned from a large, fixed snapshot of text gathered before its training cutoff. It can't see anything published after that date unless a live search fills the gap, and it can't be corrected by asking it directly.

Live search (ChatGPT search)

When ChatGPT search is active, OAI-SearchBot fetches current pages and the answer cites them. This is the part of the system that responds to a page you publish or fix today.

Third-party pages it cites

A live search answer is only as good as the pages it finds. Review sites, directories, comparison articles and news coverage feed the same retrieval step your own site does, often with more weight because they're independent of you.

The specific request in front of it

ChatGPT-User fetches a page only when a user or a Custom GPT asks it to look at that exact page. It's a one-off fetch, not a standing index, and robots.txt rules may not even apply to it since a person requested it directly.

Training data: a snapshot, not a subscription

Every GPT model has a training cutoff, a date after which nothing new was in its training set. Ask a question that depends on the model's memory alone, no browsing, no plugins, and the answer reflects whatever the web said about the topic before that date, filtered through everything else the model learned at the same time.

This is why an old address, a discontinued product line or a former business name can persist in answers long after the change happens. The model isn't checking a database. It generated an answer from patterns learned once, and it stays that way until a newer model is trained on newer data.

There's no way to edit this directly. OpenAI provides feedback controls on individual answers and separate channels for personal-data requests, but neither functions as a correction service for a single fact. The lever that actually moves things is what live search finds when it looks for you today.

Live search: a crawler, not a static index

ChatGPT search works differently. When it's active, the system retrieves current web pages relevant to the question and builds the answer from what it finds, citing sources inline. OAI-SearchBot is the crawler behind that retrieval, and OpenAI's documentation is explicit that blocking it removes a site from those cited results specifically, separate from whatever GPTBot does with training data.

That means a page published this week can show up in a ChatGPT search answer this week, which isn't true of the trained model at all. It also means the answer depends on what the crawler can actually read. A page that renders its content with client-side JavaScript, sits behind a paywall, or sits behind a bot-blocking CDN setting gives the crawler nothing to cite, even with a perfectly configured robots.txt.

Exactly how much of that retrieval runs on OpenAI's own crawled index versus a licensed search partner isn't something OpenAI has broken down publicly in a way stable enough to state as fact here. That detail has shifted before. Check OpenAI's current documentation rather than any fixed description of the architecture, including this one.

The crawlers, and why blocking one doesn't block the others

OAI-SearchBot, GPTBot and ChatGPT-User are separate robots.txt tokens with separate jobs. OAI-SearchBot pulls pages into live search answers. GPTBot gathers material for training future models and has nothing to do with what ChatGPT cites today. ChatGPT-User fetches one page at a time when a person or a Custom GPT asks for it directly.

A business that wants to stay out of model training but still show up in live answers can disallow GPTBot and allow OAI-SearchBot in the same robots.txt file. A business that blocks all three, often by accident through a security plugin or CDN default, disappears from ChatGPT search entirely while potentially still existing in whatever an older model already learned.

Check which of the three a current robots.txt allows before assuming a site is invisible for the wrong reason, or visible for a reason nobody intended.

Why the same question gets different answers

Ask ChatGPT about the same business twice and the answers can differ, sometimes within the same day. Live search retrieves whatever pages rank or match best for that specific phrasing at that moment, and a slightly different question can pull in a different set of sources.

The model itself also introduces some variation in how it phrases and weighs what it retrieves. That's normal for a language model, not a sign anything is broken. It's also why judging visibility from one conversation is unreliable; a fixed set of questions tracked over weeks tells you far more than a single exchange.

None of that makes the system arbitrary. An answer reflects the sources available at the moment of the query, and those sources are the thing worth improving.

OpenAI's crawlers, and what blocking each one does

Four bots, three distinct jobs. Confusing them is the most common mistake in advice about this topic.

OpenAI's crawlers and the effect of blocking each one
CrawlerWhat it doesIf you block it
OAI-SearchBotFetches pages to answer live ChatGPT search queriesRemoved from ChatGPT search citations specifically
GPTBotCollects content for training future modelsExcluded from training data; no effect on live search citations
ChatGPT-UserFetches one page when a user or Custom GPT asks ChatGPT to look at itThat single request fails; this fetch may not follow robots.txt, since it's user-initiated
OAI-AdsBotChecks landing pages submitted for ChatGPT adsAffects ad approval only, not organic visibility

Source: OpenAI's crawler documentation at developers.openai.com. It's the most reliable place to check when any of this changes.

Frequently asked

Where does ChatGPT get its information from?: common questions

Does ChatGPT use Google or Bing to search the web?

OpenAI hasn't published a fixed answer to that, and it's the kind of infrastructure detail that has changed since ChatGPT search launched. What's documented is the crawler doing the fetching on OpenAI's side, OAI-SearchBot, and that it's kept separate from GPTBot. Whether the underlying retrieval also draws on a licensed search partner isn't something to state as settled fact; check OpenAI's current documentation.

Is every ChatGPT answer based on a live search?

No. ChatGPT search runs when the feature is triggered, either automatically for a query that looks time-sensitive or because the user turned it on. A question answered from the base model alone reflects only what was in its training data up to the cutoff, with no live retrieval at all.

What's the difference between GPTBot, OAI-SearchBot and ChatGPT-User?

GPTBot crawls content to train future models. OAI-SearchBot crawls pages to power live ChatGPT search answers. ChatGPT-User fetches a single page on request, when a user or a Custom GPT asks ChatGPT to look at something specific. They're controlled separately in robots.txt, and blocking one has no effect on the others.

Why does ChatGPT sometimes give different answers to the same question?

Live search pulls whatever sources best match the exact wording of a query at that moment, so a slightly different phrasing can retrieve a different set of pages. The model also varies its own output somewhat by design. Track a fixed set of questions over time rather than judging visibility from a single conversation.

Related
Answers · ChatGPT

Getting your product into ChatGPT

Answers · Brand accuracy

When AI gets your company wrong

Answers · AI crawlers

What is llms.txt?

Answers · Google

Getting into AI Overviews

Related services

+3

Further reading from the journal

+1

Want to be the brand the models name?