
Does robots.txt control AI crawlers? What it does and does not do
Partly. robots.txt is a request that documented crawlers from OpenAI, Anthropic, Google, Apple, Perplexity and Common Crawl say they follow. It does not reliably govern user-triggered fetchers, which several vendors say may ignore it. And it cannot stop your CDN or firewall from blocking a crawler you meant to allow.
Key takeaways
- A crawler obeys only the most specific group that names it; the * group stops applying.
- User-triggered fetchers (ChatGPT-User, Perplexity-User, Google-Agent) may ignore robots.txt by design.
- CDN and WAF rules act before robots.txt is ever consulted.
- robots.txt controls fetching, not indexing or removal.
robots.txt is a plain-text file at the root of a host that tells crawlers which paths they may fetch. Since 2022 it has a standard, RFC 9309. It is still a convention: a crawler follows it because its operator chose to. For AI crawlers that is mostly good news — the big vendors document that they comply — but there are three gaps worth knowing.
What does robots.txt actually control?
Fetching, for crawlers that choose to obey it. It does not control indexing: Google documents that a disallowed URL can still be indexed without its content if other pages link to it. It does not authenticate anything, and it is public, so it is not a place to hide private paths.
Which group does an AI crawler follow?
This is the rule that breaks most hand-written files. A crawler looks for the group whose User-agent line most specifically matches its token and follows only that group. Google describes it as “the group with the most specific user agent that matches”. The * group is a fallback, not a base layer.
User-agent: *
Disallow: /admin/
Disallow: /checkout/
User-agent: GPTBot
Disallow: /blog/
Here GPTBot is blocked from /blog/ but allowed into /admin/ and /checkout/, because once a named group matches, the * rules no longer apply to it. Repeat any shared rules inside each named group. Within a group, the longest matching path wins, and Google uses the least restrictive rule on a tie.
Do user-triggered fetchers obey robots.txt?
Not reliably, and the vendors say so. OpenAI: for ChatGPT-User, “robots.txt rules may not apply.” Perplexity: Perplexity-User “generally ignores robots.txt rules.” Google, for its user-triggered fetchers including Google-Agent: “these fetchers generally ignore robots.txt rules.” Meta: Meta-ExternalFetcher “may bypass robots.txt rules.” The reasoning is that a person asked for that page. Anthropic is the exception among the four AI labs here: it describes blocking Claude-User through robots.txt. If you need a hard block on a user-triggered fetcher, it has to happen at your server or CDN.
robots.txt says what you would like. Your CDN decides what actually happens. Test both.
Adnan Arodiya, Crawlwise
Can a CDN or firewall block AI crawlers regardless of robots.txt?
Yes — in both directions. A firewall rule that returns 403 to a bot blocks it even if robots.txt says Allow: /. Cloudflare’s Block AI bots setting “blocks verified bots that are classified as crawling for the purpose of AI training, as well as a number of unverified bots that behave similarly,” and excludes bots used for both training and search. Since July 1, 2025, Cloudflare asks every new domain at sign-up whether to allow AI crawlers. Generic bot-fight or managed-challenge modes can also serve a JavaScript challenge that no crawler solves.
How common is this? In the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026), 3,365 domains refused at least one AI crawler at the live request while their robots.txt did not disallow it, and 363 domains (3.6%) answered at least one crawler with a challenge page. “Refused” there means a 401, 403 or other error status; because the list contains many hosts that do not serve web pages, part of that count is not a deliberate AI policy. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.
What should you do about it?
- Write the policy you intend; start from these examples.
- Repeat shared disallows in every named group.
- Check your CDN’s bot and AI settings so they match the file.
- Test the live result, not the file — see how to check whether AI crawlers can access your site.
How Crawlwise tests this
The robots.txt tester fetches your live file and shows which group and which rule decided the result for a given crawler and path. The AI crawler checker adds the live request, so a 403 or a Cloudflare, Akamai or DataDome challenge shows up even when robots.txt allows the crawler. To write a file from scratch, use the robots.txt generator.
Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.
What a plan actually costs
The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.
| Plan | Monthly | Credits | A good fit when |
|---|---|---|---|
| Starter | $4.99 | 10 | You look after one site and check it now and then |
| Pro | $14.99 | 40 | A few sites, plus the API and a handful of watched URLs |
| Studio | $49.99 | 150 | Client work that would burn through Pro mid-month |
| Agency | $99.99 | 400 | Many locations, reported under your own name |
A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.
Frequently asked questions
How long until a robots.txt change takes effect?
It depends on the vendor. OpenAI says it can take about 24 hours for its systems to adjust after a robots.txt update. Other vendors do not all publish a number.
Does Disallow remove a page from an AI index?
Not necessarily. robots.txt controls fetching. A page already indexed, or known from links elsewhere, can still be referenced. Removal needs the vendor’s own mechanism or, for Google, noindex on a crawlable page.
Are user-agent names case-sensitive in robots.txt?
No. RFC 9309 says product tokens are matched case-insensitively, so GPTBot and gptbot are the same group.
Sources
Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.
- RFC 9309 — Robots Exclusion Protocol
- Google Search Central — How Google interprets the robots.txt specification
- OpenAI — Overview of OpenAI crawlers
- Anthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity — Perplexity crawlers
- Google Search Central — Google user-triggered fetchers
- Meta — Meta web crawlers
- Cloudflare Docs — Block AI bots
- Cloudflare — press release, July 1 2025: permission-based approach for AI crawlers
- Crawlwise — The AI Crawler Access Report (top-10000 run, 2026-Q4)
Crawlwise crawler check vs checking by hand
Resolves robots.txt groups the way crawlers do (RFC 9309)
- Crawlwise
- Yes, per crawler token
- By hand
- By eye; group precedence is easy to misread
Fetches the page as each AI crawler
- Crawlwise
- Yes, one live request per agent
- By hand
- One curl at a time
Spots CDN and firewall blocks (403, challenge pages)
- Crawlwise
- Yes, with challenge fingerprints
- By hand
- Only if you read the response body
Separates a training opt-out from an accidental block
- Crawlwise
- Labelled separately
- By hand
- Up to you
Sees blocks that only happen in other regions or on real crawler IPs
- Crawlwise
- No — one location, one moment
- By hand
- Only from your own server logs
Meet the author

Adnan Arodiya
Crawlwise
Writes about what Crawlwise actually measures: on-page evidence, crawler access, and performance signals — with the limits stated up front.


