Skip to content
How to check whether AI crawlers can access your site (step by step) — illustrated banner

How to check whether AI crawlers can access your site (step by step)

Check five layers in order: read robots.txt and work out which group each crawler matches; review CDN and firewall bot rules; request the page with each crawler’s user agent and compare status codes and sizes; confirm the content is in the server HTML, not only JavaScript; then confirm real visits in server logs.

Share article:

Key takeaways

  • robots.txt is only the first of five layers that decide access.
  • A 200 with a much smaller body than a browser gets is a hidden block.
  • Spoofed user-agent tests can disagree with what the real crawler sees.
  • Server logs, checked against vendor IP ranges, are the final word.

“Can ChatGPT read my site?” has a definite answer, but it lives in five places at once. Most people check the first one, robots.txt, and stop. These steps go in the order a crawler’s request actually travels.

Step 1: What does robots.txt say for each crawler?

Open https://your-site/robots.txt. For each crawler you care about, find the group whose User-agent most specifically matches its token — that group alone applies, and the * group is ignored for it. Then find the longest Allow or Disallow path that matches the URL. If the file returns a 404, everything is allowed. If it returns an HTML page or a challenge instead of a text file, crawlers cannot read your policy at all; in the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026), that was true for 5,660 of 10,000 domains. The rules are explained in does robots.txt control AI crawlers?

Step 2: What do your CDN and firewall do with bots?

Open your CDN or WAF dashboard and look for AI-bot settings, bot-fight or bot-management modes, managed challenges, country blocks and custom rules that match user agents. On Cloudflare, check “Block AI bots” and AI Crawl Control; new domains have been asked about AI crawlers at sign-up since July 2025. A rule here overrides whatever robots.txt says.

Step 3: What does a request with each user agent return?

Request the page as a browser and as each crawler, and compare status code and size:

URL=https://example.com/
curl -s -o /dev/null -w "browser   %{http_code} %{size_download}\n" -A "Mozilla/5.0" "$URL"
curl -s -o /dev/null -w "GPTBot    %{http_code} %{size_download}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.3; +https://openai.com/gptbot)" "$URL"
curl -s -o /dev/null -w "ClaudeBot %{http_code} %{size_download}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)" "$URL"

Read the result like this: a 403 or 401 is a block. A 200 whose body is a fraction of the browser’s, or contains “Just a moment” or a captcha, is a challenge page. A 429 is rate limiting, not a policy. Remember the limit of this test: CDNs that verify bots by IP may treat your spoofed request differently from the real crawler, in either direction.

A 200 is not the same as access. Compare the size of what the crawler got with what a browser got.

Adnan Arodiya, Crawlwise

Step 4: Is the content in the server HTML?

Most AI vendors do not document JavaScript rendering for their crawlers, so assume they read the HTML your server sends. View source (not the inspector) or run curl -s "$URL" | grep -i "a sentence from your page". If your main copy is missing, it is rendered client-side, and a non-rendering crawler gets an empty shell.

Step 5: Do your server logs show real visits?

Search access logs for the tokens (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User) and note the status codes they received. Anyone can send those strings, so verify the IP against the vendor’s published ranges where one exists — OpenAI publishes ranges for its bots on its crawler page. A run of 403s against a verified crawler IP is the clearest evidence of a block you can get.

What does a typical result look like?

For GPTBot across the same 10,000 domains, the live request was answered normally on 2,100, refused or answered with an error status on 3,433, met with a challenge page on 228, rate-limited on 58, and got no answer in time on 4,181. Many of the silent hosts are API and CDN hostnames rather than websites. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.

How Crawlwise tests this

The AI crawler checker runs steps 1 and 3 together for one URL: it reads robots.txt with an RFC 9309 matcher, sends one live request per crawler, fingerprints Cloudflare, Akamai and DataDome challenge pages, and flags a response that is materially smaller than an ordinary fetch. It cannot see other regions or your logs (steps 2 and 5 stay with you). The robots.txt tester shows the deciding rule for any path, and the AI Crawler Access Report shows how 10,000 popular domains compare.

Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.

What a plan actually costs

The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.

PlanMonthlyCreditsA good fit when
Starter$4.9910You look after one site and check it now and then
Pro$14.9940A few sites, plus the API and a handful of watched URLs
Studio$49.99150Client work that would burn through Pro mid-month
Agency$99.99400Many locations, reported under your own name

A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.

Frequently asked questions

Is a curl with a GPTBot user agent a reliable test?

It is a good first test but not proof. CDNs that verify bots check the IP address too, so a spoofed user agent from your laptop can be treated differently from the real crawler. Server logs show what the real one received.

My logs show no AI crawlers at all. Am I blocked?

Not necessarily. Crawlers visit on their own schedule, and Google-Extended and Applebot-Extended never appear because they are tokens, not crawlers. Check robots.txt and your CDN first, then watch logs over a few weeks.

How often should I re-check?

After any CDN, firewall, hosting or robots.txt change, and after your CDN vendor changes its bot defaults. A monthly check is a reasonable baseline for important pages.

Sources

Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.

Crawlwise crawler check vs checking by hand

  • Resolves robots.txt groups the way crawlers do (RFC 9309)

    Crawlwise
    Yes, per crawler token
    By hand
    By eye; group precedence is easy to misread
  • Fetches the page as each AI crawler

    Crawlwise
    Yes, one live request per agent
    By hand
    One curl at a time
  • Spots CDN and firewall blocks (403, challenge pages)

    Crawlwise
    Yes, with challenge fingerprints
    By hand
    Only if you read the response body
  • Separates a training opt-out from an accidental block

    Crawlwise
    Labelled separately
    By hand
    Up to you
  • Sees blocks that only happen in other regions or on real crawler IPs

    Crawlwise
    No — one location, one moment
    By hand
    Only from your own server logs

Try it on your URL

Run the checks on your live site

Free tools answer one question. A full audit scores the page, tests answer-engine crawlers, and saves to history on a plan.

  • Live crawler probes
  • Core Web Vitals
  • One free full audit

Meet the author

Photo of Adnan Arodiya