
How to check whether AI crawlers can access your site (step by step)
Check five layers in order: read robots.txt and work out which group each crawler matches; review CDN and firewall bot rules; request the page with each crawler’s user agent and compare status codes and sizes; confirm the content is in the server HTML, not only JavaScript; then confirm real visits in server logs.
Key takeaways
- robots.txt is only the first of five layers that decide access.
- A 200 with a much smaller body than a browser gets is a hidden block.
- Spoofed user-agent tests can disagree with what the real crawler sees.
- Server logs, checked against vendor IP ranges, are the final word.
“Can ChatGPT read my site?” has a definite answer, but it lives in five places at once. Most people check the first one, robots.txt, and stop. These steps go in the order a crawler’s request actually travels.
Step 1: What does robots.txt say for each crawler?
Open https://your-site/robots.txt. For each crawler you care about, find the group whose User-agent most specifically matches its token — that group alone applies, and the * group is ignored for it. Then find the longest Allow or Disallow path that matches the URL. If the file returns a 404, everything is allowed. If it returns an HTML page or a challenge instead of a text file, crawlers cannot read your policy at all; in the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026), that was true for 5,660 of 10,000 domains. The rules are explained in does robots.txt control AI crawlers?
Step 2: What do your CDN and firewall do with bots?
Open your CDN or WAF dashboard and look for AI-bot settings, bot-fight or bot-management modes, managed challenges, country blocks and custom rules that match user agents. On Cloudflare, check “Block AI bots” and AI Crawl Control; new domains have been asked about AI crawlers at sign-up since July 2025. A rule here overrides whatever robots.txt says.
Step 3: What does a request with each user agent return?
Request the page as a browser and as each crawler, and compare status code and size:
URL=https://example.com/
curl -s -o /dev/null -w "browser %{http_code} %{size_download}\n" -A "Mozilla/5.0" "$URL"
curl -s -o /dev/null -w "GPTBot %{http_code} %{size_download}\n" \
-A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.3; +https://openai.com/gptbot)" "$URL"
curl -s -o /dev/null -w "ClaudeBot %{http_code} %{size_download}\n" \
-A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)" "$URL"
Read the result like this: a 403 or 401 is a block. A 200 whose body is a fraction of the browser’s, or contains “Just a moment” or a captcha, is a challenge page. A 429 is rate limiting, not a policy. Remember the limit of this test: CDNs that verify bots by IP may treat your spoofed request differently from the real crawler, in either direction.
A 200 is not the same as access. Compare the size of what the crawler got with what a browser got.
Adnan Arodiya, Crawlwise
Step 4: Is the content in the server HTML?
Most AI vendors do not document JavaScript rendering for their crawlers, so assume they read the HTML your server sends. View source (not the inspector) or run curl -s "$URL" | grep -i "a sentence from your page". If your main copy is missing, it is rendered client-side, and a non-rendering crawler gets an empty shell.
Step 5: Do your server logs show real visits?
Search access logs for the tokens (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User) and note the status codes they received. Anyone can send those strings, so verify the IP against the vendor’s published ranges where one exists — OpenAI publishes ranges for its bots on its crawler page. A run of 403s against a verified crawler IP is the clearest evidence of a block you can get.
What does a typical result look like?
For GPTBot across the same 10,000 domains, the live request was answered normally on 2,100, refused or answered with an error status on 3,433, met with a challenge page on 228, rate-limited on 58, and got no answer in time on 4,181. Many of the silent hosts are API and CDN hostnames rather than websites. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.
How Crawlwise tests this
The AI crawler checker runs steps 1 and 3 together for one URL: it reads robots.txt with an RFC 9309 matcher, sends one live request per crawler, fingerprints Cloudflare, Akamai and DataDome challenge pages, and flags a response that is materially smaller than an ordinary fetch. It cannot see other regions or your logs (steps 2 and 5 stay with you). The robots.txt tester shows the deciding rule for any path, and the AI Crawler Access Report shows how 10,000 popular domains compare.
Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.
What a plan actually costs
The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.
| Plan | Monthly | Credits | A good fit when |
|---|---|---|---|
| Starter | $4.99 | 10 | You look after one site and check it now and then |
| Pro | $14.99 | 40 | A few sites, plus the API and a handful of watched URLs |
| Studio | $49.99 | 150 | Client work that would burn through Pro mid-month |
| Agency | $99.99 | 400 | Many locations, reported under your own name |
A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.
Frequently asked questions
Is a curl with a GPTBot user agent a reliable test?
It is a good first test but not proof. CDNs that verify bots check the IP address too, so a spoofed user agent from your laptop can be treated differently from the real crawler. Server logs show what the real one received.
My logs show no AI crawlers at all. Am I blocked?
Not necessarily. Crawlers visit on their own schedule, and Google-Extended and Applebot-Extended never appear because they are tokens, not crawlers. Check robots.txt and your CDN first, then watch logs over a few weeks.
How often should I re-check?
After any CDN, firewall, hosting or robots.txt change, and after your CDN vendor changes its bot defaults. A monthly check is a reasonable baseline for important pages.
Sources
Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.
- RFC 9309 — Robots Exclusion Protocol
- Google Search Central — How Google interprets the robots.txt specification
- OpenAI — Overview of OpenAI crawlers
- Anthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity — Perplexity crawlers
- Cloudflare Docs — Block AI bots
- Cloudflare — press release, July 1 2025: permission-based approach for AI crawlers
- Crawlwise — The AI Crawler Access Report (top-10000 run, 2026-Q4)
Crawlwise crawler check vs checking by hand
Resolves robots.txt groups the way crawlers do (RFC 9309)
- Crawlwise
- Yes, per crawler token
- By hand
- By eye; group precedence is easy to misread
Fetches the page as each AI crawler
- Crawlwise
- Yes, one live request per agent
- By hand
- One curl at a time
Spots CDN and firewall blocks (403, challenge pages)
- Crawlwise
- Yes, with challenge fingerprints
- By hand
- Only if you read the response body
Separates a training opt-out from an accidental block
- Crawlwise
- Labelled separately
- By hand
- Up to you
Sees blocks that only happen in other regions or on real crawler IPs
- Crawlwise
- No — one location, one moment
- By hand
- Only from your own server logs
Try it on your URL
Run the checks on your live site
Free tools answer one question. A full audit scores the page, tests answer-engine crawlers, and saves to history on a plan.
- Live crawler probes
- Core Web Vitals
- One free full audit
Meet the author

Adnan Arodiya
Crawlwise
Writes about what Crawlwise actually measures: on-page evidence, crawler access, and performance signals — with the limits stated up front.


