Skip to content
Does robots.txt control AI crawlers? What it does and does not do — illustrated banner

Does robots.txt control AI crawlers? What it does and does not do

Partly. robots.txt is a request that documented crawlers from OpenAI, Anthropic, Google, Apple, Perplexity and Common Crawl say they follow. It does not reliably govern user-triggered fetchers, which several vendors say may ignore it. And it cannot stop your CDN or firewall from blocking a crawler you meant to allow.

Share article:

Key takeaways

  • A crawler obeys only the most specific group that names it; the * group stops applying.
  • User-triggered fetchers (ChatGPT-User, Perplexity-User, Google-Agent) may ignore robots.txt by design.
  • CDN and WAF rules act before robots.txt is ever consulted.
  • robots.txt controls fetching, not indexing or removal.

robots.txt is a plain-text file at the root of a host that tells crawlers which paths they may fetch. Since 2022 it has a standard, RFC 9309. It is still a convention: a crawler follows it because its operator chose to. For AI crawlers that is mostly good news — the big vendors document that they comply — but there are three gaps worth knowing.

What does robots.txt actually control?

Fetching, for crawlers that choose to obey it. It does not control indexing: Google documents that a disallowed URL can still be indexed without its content if other pages link to it. It does not authenticate anything, and it is public, so it is not a place to hide private paths.

Which group does an AI crawler follow?

This is the rule that breaks most hand-written files. A crawler looks for the group whose User-agent line most specifically matches its token and follows only that group. Google describes it as “the group with the most specific user agent that matches”. The * group is a fallback, not a base layer.

User-agent: *
Disallow: /admin/
Disallow: /checkout/

User-agent: GPTBot
Disallow: /blog/

Here GPTBot is blocked from /blog/ but allowed into /admin/ and /checkout/, because once a named group matches, the * rules no longer apply to it. Repeat any shared rules inside each named group. Within a group, the longest matching path wins, and Google uses the least restrictive rule on a tie.

Do user-triggered fetchers obey robots.txt?

Not reliably, and the vendors say so. OpenAI: for ChatGPT-User, “robots.txt rules may not apply.” Perplexity: Perplexity-User “generally ignores robots.txt rules.” Google, for its user-triggered fetchers including Google-Agent: “these fetchers generally ignore robots.txt rules.” Meta: Meta-ExternalFetcher “may bypass robots.txt rules.” The reasoning is that a person asked for that page. Anthropic is the exception among the four AI labs here: it describes blocking Claude-User through robots.txt. If you need a hard block on a user-triggered fetcher, it has to happen at your server or CDN.

robots.txt says what you would like. Your CDN decides what actually happens. Test both.

Adnan Arodiya, Crawlwise

Can a CDN or firewall block AI crawlers regardless of robots.txt?

Yes — in both directions. A firewall rule that returns 403 to a bot blocks it even if robots.txt says Allow: /. Cloudflare’s Block AI bots setting “blocks verified bots that are classified as crawling for the purpose of AI training, as well as a number of unverified bots that behave similarly,” and excludes bots used for both training and search. Since July 1, 2025, Cloudflare asks every new domain at sign-up whether to allow AI crawlers. Generic bot-fight or managed-challenge modes can also serve a JavaScript challenge that no crawler solves.

How common is this? In the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026), 3,365 domains refused at least one AI crawler at the live request while their robots.txt did not disallow it, and 363 domains (3.6%) answered at least one crawler with a challenge page. “Refused” there means a 401, 403 or other error status; because the list contains many hosts that do not serve web pages, part of that count is not a deliberate AI policy. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.

What should you do about it?

  1. Write the policy you intend; start from these examples.
  2. Repeat shared disallows in every named group.
  3. Check your CDN’s bot and AI settings so they match the file.
  4. Test the live result, not the file — see how to check whether AI crawlers can access your site.

How Crawlwise tests this

The robots.txt tester fetches your live file and shows which group and which rule decided the result for a given crawler and path. The AI crawler checker adds the live request, so a 403 or a Cloudflare, Akamai or DataDome challenge shows up even when robots.txt allows the crawler. To write a file from scratch, use the robots.txt generator.

Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.

What a plan actually costs

The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.

PlanMonthlyCreditsA good fit when
Starter$4.9910You look after one site and check it now and then
Pro$14.9940A few sites, plus the API and a handful of watched URLs
Studio$49.99150Client work that would burn through Pro mid-month
Agency$99.99400Many locations, reported under your own name

A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.

Frequently asked questions

How long until a robots.txt change takes effect?

It depends on the vendor. OpenAI says it can take about 24 hours for its systems to adjust after a robots.txt update. Other vendors do not all publish a number.

Does Disallow remove a page from an AI index?

Not necessarily. robots.txt controls fetching. A page already indexed, or known from links elsewhere, can still be referenced. Removal needs the vendor’s own mechanism or, for Google, noindex on a crawlable page.

Are user-agent names case-sensitive in robots.txt?

No. RFC 9309 says product tokens are matched case-insensitively, so GPTBot and gptbot are the same group.

Sources

Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.

Crawlwise crawler check vs checking by hand

  • Resolves robots.txt groups the way crawlers do (RFC 9309)

    Crawlwise
    Yes, per crawler token
    By hand
    By eye; group precedence is easy to misread
  • Fetches the page as each AI crawler

    Crawlwise
    Yes, one live request per agent
    By hand
    One curl at a time
  • Spots CDN and firewall blocks (403, challenge pages)

    Crawlwise
    Yes, with challenge fingerprints
    By hand
    Only if you read the response body
  • Separates a training opt-out from an accidental block

    Crawlwise
    Labelled separately
    By hand
    Up to you
  • Sees blocks that only happen in other regions or on real crawler IPs

    Crawlwise
    No — one location, one moment
    By hand
    Only from your own server logs

Meet the author

Photo of Adnan Arodiya