The AI Crawler Access Report
39.2% of 102 domains block at least one AI crawler.
Sample of 102 domains, collected 2026-09-22.
Headline
| Sample | 102 domains |
|---|---|
| Blocking at least one AI crawler | 40 (39.2%) |
| Blocking a retrieval crawler | 26 (25.5%) |
| Challenged rather than blocked | 41 (40.2%) |
By crawler
| Crawler | Allowed | Blocked | Challenged | Not established |
|---|---|---|---|---|
| Bytespider | 28 | 28 | 30 | 16 |
| Meta-ExternalAgent | 33 | 25 | 30 | 14 |
| GPTBot | 35 | 24 | 29 | 14 |
| CCBot | 35 | 27 | 26 | 14 |
| ClaudeBot | 35 | 25 | 27 | 15 |
| Google-Extended | 36 | 24 | 28 | 14 |
| ChatGPT-User | 36 | 17 | 34 | 15 |
| Perplexity-User | 37 | 22 | 28 | 15 |
| Amazonbot | 38 | 22 | 28 | 14 |
| PerplexityBot | 40 | 22 | 27 | 13 |
| cohere-ai | 38 | 22 | 27 | 15 |
| Diffbot | 38 | 22 | 27 | 15 |
| OAI-SearchBot | 40 | 15 | 33 | 14 |
| Claude-User | 40 | 20 | 28 | 14 |
| Timpibot | 39 | 22 | 26 | 15 |
| Applebot-Extended | 40 | 21 | 26 | 15 |
| Claude-SearchBot | 41 | 18 | 28 | 15 |
By industry
| Industry | Domains | Blocking anything |
|---|---|---|
| reference | 1 | 1 (100%) |
| news | 5 | 5 (100%) |
| food | 4 | 3 (75%) |
| entertainment | 4 | 3 (75%) |
| property | 4 | 3 (75%) |
| health | 4 | 2 (50%) |
| employment | 4 | 2 (50%) |
| industrial | 4 | 2 (50%) |
| consumer | 5 | 2 (40%) |
| technology | 16 | 6 (37.5%) |
| ecommerce | 6 | 2 (33.3%) |
| retail | 6 | 2 (33.3%) |
| automotive | 6 | 2 (33.3%) |
| software | 10 | 3 (30%) |
| education | 5 | 1 (20%) |
| travel | 11 | 1 (9.1%) |
| finance | 3 | 0 (0%) |
| logistics | 4 | 0 (0%) |
What changed this quarter
No earlier run of a comparable sample exists to compare against, so nothing is claimed here.
Methodology, in full
Sample. This run probed 102 domains: a hand-assembled list of well-known sites, published at scripts/fixtures/crawler-report-sample.csv, with every industry label assigned by us. It is not a random sample of the web and it is not a top-sites ranking, so every share below describes these domains and not the web. A domain that fails to resolve is recorded as not established rather than dropped, because dropping failures is how a sample silently becomes favourable.
robots.txt. Fetched from the origin, verified to be a robots document rather than an HTML challenge page, then resolved with an RFC 9309 matcher against each crawler’s product token and the probed path. A file that could not be read as a robots file leaves that crawler not established.
Live fetch. The URL is fetched with that crawler’s user-agent. The response is classified by status, by challenge fingerprints (Cloudflare, Akamai, DataDome), and by body size against an ordinary fetch. A 403 or a challenge is a block. A 429 or a timeout is not established, not a block.
Verdicts. Allowed, blocked, challenged, or not established. A disallow in robots.txt for a training crawler is recorded as a deliberate opt-out, and it is counted separately from an accidental block everywhere on this page.
What this cannot tell you. One probe from one location at one moment: a CDN answering a crawler differently in another region is invisible here. A crawler that ignores robots.txt is not stopped by it, and this measures policy as much as enforcement. Rate limiting is only called a block when the response is a 403. And a domain blocking a training crawler while allowing the retrieval crawlers from the same owner is reported as exactly that, not as one number.
Crawlers probed
| Crawler | Owner | Category | User agent sent |
|---|---|---|---|
| GPTBot | OpenAI | training | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.3; +https://openai.com/gptbot |
| OAI-SearchBot | OpenAI | retrieval | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot |
| ChatGPT-User | OpenAI | retrieval | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot |
| ClaudeBot | Anthropic | training | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com) |
| Claude-SearchBot | Anthropic | retrieval | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-SearchBot/1.0; +Claude-SearchBot@anthropic.com) |
| Claude-User | Anthropic | retrieval | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +Claude-User@anthropic.com) |
| Google-Extended | training | Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html) | |
| PerplexityBot | Perplexity | retrieval | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) |
| Perplexity-User | Perplexity | retrieval | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user) |
| CCBot | Common Crawl | research | CCBot/2.0 (https://commoncrawl.org/faq/) |
| Bytespider | ByteDance | training | Mozilla/5.0 (compatible; Bytespider; +https://zhanzhang.toutiao.com/) |
| Applebot-Extended | Apple | training | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot) |
| Meta-ExternalAgent | Meta | training | Mozilla/5.0 (compatible; Meta-ExternalAgent/1.0; +https://developers.facebook.com/docs/sharing/webmasters/crawler) |
| Amazonbot | Amazon | training | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/600.2.5 (KHTML, like Gecko) Version/8.0.2 Safari/600.2.5 (Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) |
| cohere-ai | Cohere | retrieval | Mozilla/5.0 (compatible; cohere-ai/1.0; +https://cohere.com/bot) |
| Diffbot | Diffbot | research | Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com) |
| Timpibot | Timpi | research | Mozilla/5.0 (compatible; Timpibot/0.1; +https://timpi.io) |
Reproduce it
node scripts/crawler-report.mjs --list scripts/fixtures/crawler-report-sample.csv --limit 1000 --label "well-known"
A larger sample is the same command with a longer list, at roughly nineteen outbound requests per domain. 10,000 domains is about 190,000 requests and several hours; the page states whichever sample size it is reading, and nothing here is extrapolated from a smaller one.