
GPTBot vs OAI-SearchBot vs ChatGPT-User: what each OpenAI bot does
GPTBot collects content that may be used to train OpenAI’s foundation models. OAI-SearchBot builds the index behind ChatGPT search results and citations. ChatGPT-User fetches a page when a person’s request in ChatGPT or a custom GPT needs it, and OpenAI says robots.txt rules may not apply to it. Each is controlled independently.
Key takeaways
- Block GPTBot to opt out of training; it does not affect ChatGPT search.
- Allow OAI-SearchBot if you want to appear in ChatGPT search answers.
- ChatGPT-User is user-triggered and may not follow robots.txt.
- OpenAI says robots.txt changes can take about 24 hours to apply.
OpenAI documents its crawlers on one page and states plainly that “each setting is independent of the others.” That independence is the whole point: you can keep your content out of training and still be cited in ChatGPT search.
What does each OpenAI bot do?
| Agent | Purpose (OpenAI’s words) | If you disallow it | robots.txt |
|---|---|---|---|
| GPTBot | “Used to make our generative AI foundation models more useful and safe” | “Indicates a site’s content should not be used in training generative AI foundation models” | Followed |
| OAI-SearchBot | “Used to surface websites in search results in ChatGPT’s search features” | The site “will not be shown in ChatGPT search answers” | Followed |
| ChatGPT-User | Used “for certain user actions in ChatGPT and Custom GPTs”; “not used for crawling the web in an automatic fashion” | Not a dependable control | “robots.txt rules may not apply” |
How do you block training but stay in ChatGPT search?
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
If your * group has disallows you rely on (for example /admin/), repeat them inside the OAI-SearchBot and ChatGPT-User groups — a named group replaces the * group for that crawler. OpenAI says it can take about 24 hours from a robots.txt update for its systems to adjust. More policies are in robots.txt for AI crawlers: copy-paste examples.
GPTBot is about training. OAI-SearchBot is about being found. Most sites that block one block both without meaning to.
Adnan Arodiya, Crawlwise
Can ChatGPT cite you if OAI-SearchBot is blocked?
Treat a block as removing you from ChatGPT search answers — that is how OpenAI describes it. OpenAI also says ChatGPT search draws on third-party search providers as well as its own crawler, so being well indexed in Bing remains part of being findable in ChatGPT; see how assistants find web pages. ChatGPT-User may still fetch a page a person explicitly asks about.
What do popular sites do with OpenAI’s bots?
From the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026), among the 1,837 domains with a readable robots.txt: 785 (42.7%) disallow GPTBot, 738 (40.2%) disallow OAI-SearchBot and 753 (41.0%) disallow ChatGPT-User at the homepage. Of the 785 that disallow GPTBot, 728 also disallow OAI-SearchBot; just 57 use the split policy above. At the live request, GPTBot was answered normally by 2,100 of all 10,000 domains and refused or given an error status by 3,433. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.
What about OAI-AdsBot?
OpenAI lists OAI-AdsBot as the agent that checks the safety of pages submitted as ads on ChatGPT, and says its data is not used to train foundation models. If you do not advertise on ChatGPT, you will not need to think about it.
How Crawlwise tests this
The AI crawler checker shows GPTBot, OAI-SearchBot and ChatGPT-User as separate rows, each with its robots.txt result and live response. A GPTBot disallow is labelled as a training opt-out; an OAI-SearchBot block is flagged because it affects ChatGPT search. Use the robots.txt tester to test a path for each agent and see the rules in the group that names it.
Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.
What a plan actually costs
The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.
| Plan | Monthly | Credits | A good fit when |
|---|---|---|---|
| Starter | $4.99 | 10 | You look after one site and check it now and then |
| Pro | $14.99 | 40 | A few sites, plus the API and a handful of watched URLs |
| Studio | $49.99 | 150 | Client work that would burn through Pro mid-month |
| Agency | $99.99 | 400 | Many locations, reported under your own name |
A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.
Frequently asked questions
Which OpenAI bot do I need to allow for ChatGPT search?
OAI-SearchBot. OpenAI says sites that opt out of it will not be shown in ChatGPT search answers.
Does blocking GPTBot remove content already used for training?
OpenAI describes the GPTBot disallow as a signal that content should not be used in training. Its crawler page does not describe removal from models already trained.
What is OAI-AdsBot?
OpenAI lists it as the agent used to validate the safety of pages submitted as ads on ChatGPT, and says its data is not used to train foundation models. It only matters if you advertise on ChatGPT.
Sources
Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.
Crawlwise crawler check vs checking by hand
Resolves robots.txt groups the way crawlers do (RFC 9309)
- Crawlwise
- Yes, per crawler token
- By hand
- By eye; group precedence is easy to misread
Fetches the page as each AI crawler
- Crawlwise
- Yes, one live request per agent
- By hand
- One curl at a time
Spots CDN and firewall blocks (403, challenge pages)
- Crawlwise
- Yes, with challenge fingerprints
- By hand
- Only if you read the response body
Separates a training opt-out from an accidental block
- Crawlwise
- Labelled separately
- By hand
- Up to you
Sees blocks that only happen in other regions or on real crawler IPs
- Crawlwise
- No — one location, one moment
- By hand
- Only from your own server logs
Meet the author

Adnan Arodiya
Crawlwise
Writes about what Crawlwise actually measures: on-page evidence, crawler access, and performance signals — with the limits stated up front.


