
robots.txt in the wild: what the top 10,000 domains actually serve
Of the first 10,000 domains on the Cisco Umbrella popularity list, probed on October 1, 2026, only 1,837 served a robots.txt a crawler could read as one. 2,503 answered 404 or 410, which means everything is allowed. The other 5,660 returned no file we could read, usually because the host serves no website at all.
Key takeaways
- 1,837 of 10,000 domains (18.4%) served a readable robots.txt; 2,503 (25.0%) returned 404 or 410; 5,660 (56.6%) could not be read.
- Of the 5,660 unreadable, 4,008 gave no HTTP response to any homepage request either: they are not websites.
- Among the 5,949 hosts that answered at all, 30.4% had a readable file and 41.8% had none.
- A 404 robots.txt allows everything; a 5xx one tells RFC 9309 crawlers to stay out.
Most robots.txt advice assumes there is a file. We wanted to know how often that is true for the domains people hit most, so we counted. This post uses the Crawlwise AI Crawler Access Report dataset (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026). The run requested https://domain/robots.txt for every domain and tested the homepage path against 17 AI crawler tokens. The robots.txt part of that run is a clean census of one question: did this host serve a robots file a crawler could actually read?
How many popular domains serve a readable robots.txt?
| What /robots.txt returned | Domains | Share of 10,000 |
|---|---|---|
| A readable robots.txt | 1,837 | 18.4% |
| 404 or 410 (no file, so everything is allowed) | 2,503 | 25.0% |
| Nothing readable as a robots file | 5,660 | 56.6% |
The split barely moves with rank. In the top 1,000 it was 204 readable, 276 missing and 520 unreadable; for ranks 5,001 to 10,000 it was 894, 1,255 and 2,851. The Umbrella list ranks hostnames by DNS popularity, so it includes API, CDN, telemetry and update hosts that never serve a web page, and only the homepage path was tested. Treat every share as a description of that list, not of the web.
How did we decide what counts as “readable”?
The rules are the ones the AI crawler checker uses, in lib/checks/ai-crawlers.ts. We requested the file up to three times — with a Googlebot user agent, a desktop Chrome user agent and our own CrawlwiseAudit/1.0 — following redirects, with an 8-second timeout per request. The first answer that qualified won:
- Readable: HTTP 200, a content type that is not HTML or JSON, no HTML or challenge-page markers in the first 8,000 characters, and either nothing but comments or at least one line starting with a known directive (
User-agent,Allow,Disallow,Sitemap,Crawl-delay,HostorClean-param). - No file: no attempt qualified as readable and at least one returned 404 or 410.
- Unread: everything else — timeouts, connection and DNS failures, 403s, 5xx errors, and 200 responses that were really an HTML page or a bot challenge.
That is stricter than Google, which says it will “try to parse the content and extract rules” even from a file that is not plain text. We are stricter on purpose: an HTML error page that happens to parse to no rules should not be reported as a policy.
A robots.txt you cannot read is not a policy. It is a guess, and every crawler resolves the guess its own way.
Adnan Arodiya, Crawlwise
Why could 5,660 domains not be read at all?
Mostly because they are not websites. For 4,008 of the 5,660, every one of the 17 homepage requests in the same run also ended without any HTTP response — a timeout, a refused connection, a TLS or DNS failure. Telemetry endpoints, update servers and API hosts rank high on a DNS list and answer nothing on https://host/. Counting only the 5,949 domains whose homepage answered at least one request with an HTTP status, the picture is closer to what a site owner would expect:
| Among 5,949 hosts that answered | Domains | Share |
|---|---|---|
| Readable robots.txt | 1,808 | 30.4% |
| 404 or 410 | 2,489 | 41.8% |
| Unread | 1,652 | 27.8% |
The 1,652 are the interesting ones: a host that serves something at its homepage but returns an HTML page, a challenge, a 403 or an error at /robots.txt. Our dataset stores the verdict, not the status code of the robots request, so we cannot split that group further without guessing — and we will not.
What do the readable files actually say?
We only tested the homepage path, and only for the 17 AI crawler tokens in the report, so this says nothing about Googlebot rules or deeper paths. Of the 1,837 readable files, 1,008 (54.9%) allow all 17 tokens at /, 689 (37.5%) disallow all 17, and 140 (7.6%) give different crawlers different answers. The most-disallowed token was CCBot at 797 (43.4%), the least was Applebot-Extended at 729 (39.7%). Many of the 689 full disallows belong to asset and API hosts — gstatic.com, ranked second, disallows all 17 tokens at its homepage — so a blanket Disallow: / there is housekeeping, not a stance on AI. The vendor-by-vendor picture is in the AI Crawler Access Report and in does robots.txt control AI crawlers?
Why does the status code of robots.txt matter so much?
Because crawlers treat the three outcomes in opposite ways. RFC 9309 says that when a server answers in the 400–499 range the file is “unavailable” and the crawler “MAY access any resources on the server”. When it is unreachable “due to server or network errors”, the crawler “MUST assume complete disallow”. Google follows the same split: it treats “all 4xx errors, except 429, as if a valid robots.txt file didn't exist”, and on a 5xx it stops crawling the site for the first 12 hours, then uses the last good version for the next 30 days while it keeps retrying.
So the worst robots.txt is not a missing one. It is one that returns 503 because a firewall rule or an origin outage only hits that path. Other details from the specification worth knowing:
- Size: RFC 9309 sets the parsing limit at no less than 500 KiB; Google enforces exactly 500 KiB and ignores anything after it.
- Redirects: crawlers “SHOULD follow at least five consecutive redirects”; Google follows at least five hops and then treats the file as a 404.
- Caching: crawlers “SHOULD NOT use the cached version for more than 24 hours”; Google says it generally caches for up to 24 hours. A fix can take a day to be seen.
- Precedence: the most specific (longest) matching path wins, and on a tie Google uses the least restrictive rule.
What should you check on your own site?
curl -sI https://example.com/robots.txt
# Expect: HTTP/2 200 and content-type: text/plain
# A 404 is fine if you have no rules.
# A 403, 5xx, or text/html is the case to fix.
Run it from outside your network, and once with a crawler user agent, because CDNs often answer bots differently. Then confirm the rules do what you meant; robots.txt for AI crawlers: copy-paste examples has working files, and noindex vs robots.txt disallow explains why robots.txt is the wrong tool for keeping a page out of search.
How Crawlwise checks this
The free robots.txt tester fetches your live /robots.txt, applies the same “readable” test described above, and then tells you whether a path is allowed or disallowed for the user agent you choose, and lists up to eight rules from the group that names that agent so you can see which one applies. It tests the path only, without the query string. When the file is not readable it says so plainly, and notes that a missing robots.txt normally means crawling is allowed. It reads the file as served to one request today; a CDN rule that answers real crawler IPs differently is outside what it can see, which is what the AI crawler checker probes per bot.
What each check reads, what counts as a failure, and what the checks cannot see are written up on the methodology page.
What a plan actually costs
The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.
| Plan | Monthly | Credits | A good fit when |
|---|---|---|---|
| Starter | $4.99 | 10 | You look after one site and check it now and then |
| Pro | $14.99 | 40 | A few sites, plus the API and a handful of watched URLs |
| Studio | $49.99 | 150 | Client work that would burn through Pro mid-month |
| Agency | $99.99 | 400 | Many locations, reported under your own name |
A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.
Frequently asked questions
Is a missing robots.txt a problem?
Not for crawling. RFC 9309 says that when the file is unavailable (a 4xx answer) a crawler may access any resource, and Google treats every 4xx except 429 as if no robots.txt existed. It only becomes a problem if you meant to keep crawlers out of something.
What happens if robots.txt returns a 500 error?
RFC 9309 tells crawlers to assume complete disallow when the file is unreachable because of server or network errors. Google says it stops crawling the site for the first 12 hours, then uses the last good copy for up to 30 days. A robots.txt that errors is far worse than one that is missing.
Why do the shares not add up to the whole web?
Because the list is not the web. Umbrella ranks hostnames by how often they are looked up in DNS, so it is full of API, CDN and telemetry hosts. 4,008 of the 10,000 gave no HTTP response at all to our homepage requests.
Sources
Documentation was read on 2026-10-05. Search engines change these pages without notice, so check the original before you act on a detail.
Crawlwise robots.txt check vs reading the file by hand
Tells a 404 robots.txt apart from an unreadable one
- Crawlwise
- Yes, labelled separately
- By hand
- Only if you check the status code yourself
Rejects an HTML page or challenge served at /robots.txt
- Crawlwise
- Yes, treated as unread
- By hand
- Easy to miss in a browser
Applies RFC 9309 group and longest-match rules
- Crawlwise
- Yes, per crawler token
- By hand
- By eye; precedence is easy to misread
Shows what Googlebot saw last week
- Crawlwise
- No — one live request, now
- By hand
- Search Console robots.txt report
Meet the author

Adnan Arodiya
Crawlwise
Writes about what Crawlwise actually measures: on-page evidence, crawler access, and performance signals — with the limits stated up front.


