
What is an AI crawler? Training, search, and user-triggered fetchers
An AI crawler is an automated client run by an AI company that fetches web pages. There are three kinds: training crawlers that gather text for future models, search crawlers that build an index for cited answers, and user-triggered fetchers that load one page because a person asked. Each kind needs a separate decision.
Key takeaways
- Training, search, and user-triggered agents are separate bots with separate robots.txt tokens.
- Blocking a training crawler does not, by itself, remove you from AI answers.
- Google-Extended and Applebot-Extended are control tokens, not crawlers you will see in logs.
- User-triggered fetchers may not follow robots.txt; vendors say so in their own docs.
“AI crawler” gets used for any bot with an AI company’s name on it. That hides the distinction that matters to a site owner: what the fetched page is used for. A page fetched to train a model, a page fetched to build a search index, and a page fetched because a person pasted your URL into a chat are three different events with three different consequences.
What counts as an AI crawler?
An AI crawler is any automated HTTP client operated by an AI company that requests your pages and identifies itself with a documented user agent or robots.txt product token. The vendors that document theirs — OpenAI, Anthropic, Google, Apple, Perplexity, Meta and Common Crawl — each publish a page listing the names and what each one is for. Those pages are the only reliable source; third-party “bot directories” often copy each other and go stale.
What are the three kinds of AI crawler?
| Kind | What it does with your page | Documented examples | robots.txt |
|---|---|---|---|
| Training | Collects content that may be used to train future models | GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), Meta-ExternalAgent (Meta) | Followed, per each vendor |
| Search / index | Builds the index an assistant searches before it answers and cites | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Followed, per each vendor |
| User-triggered | Fetches one page at the moment a person’s request needs it | ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher, Google-Agent | OpenAI says rules “may not apply”; Perplexity and Google say these fetchers “generally ignore” them; Meta says it “may bypass” them |
Two names sit outside the table because they are not crawlers at all. Google-Extended “doesn’t have a separate HTTP request user agent string”, according to Google; it is a robots.txt token that controls whether content Google already crawls may be used for Gemini training and grounding. Applebot-Extended works the same way for Apple: Apple says it “does not crawl webpages”. You will never see either in a log file, and blocking them does not change what Googlebot or Applebot fetch.
Why does the distinction matter?
Because the trade-offs point in opposite directions. Opting out of training is a reasonable business decision with no direct effect on whether assistants can cite you. Blocking a search crawler removes you from that assistant’s index: OpenAI says sites that opt out of OAI-SearchBot “will not be shown in ChatGPT search answers”, and Anthropic says disabling Claude-SearchBot “prevents our system from indexing your content for search optimization”. Mixing the two up — usually with one broad rule — is how sites disappear from answers they wanted to be in. Which AI crawlers should you allow? walks through that decision vendor by vendor.
Ask what the fetched page is used for, not whose bot it is. That one question sorts almost every AI crawler decision.
Adnan Arodiya, Crawlwise
How is an AI crawler different from Googlebot or Bingbot?
Classic search crawlers are now AI-search crawlers too. Google’s AI Overviews and AI Mode are built on the regular Google index, and Google says a page must be indexed and eligible for a snippet to appear as a supporting link. Microsoft’s Copilot answers are grounded on the Bing index. So the most important “AI crawlers” for many sites are still Googlebot and Bingbot — see how ChatGPT, Claude, Gemini, Perplexity, Copilot and Grok find pages.
How many sites block AI crawlers today?
From the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026): 4,159 of 10,000 domains (41.6%) refused at least one of the 17 AI crawlers we test, either with a robots.txt disallow or with a refused or error response to the live fetch. Only 1,837 of the 10,000 served a robots.txt we could read as one. Among those, 785 (42.7%) disallow GPTBot at the homepage and 736 (40.1%) disallow Claude-SearchBot. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.
How Crawlwise tests this
The free AI crawler checker reads your robots.txt with an RFC 9309 matcher, then requests your URL once with each crawler’s user agent and classifies the answer: allowed, blocked, challenged, or not established. A training opt-out in robots.txt is labelled as a deliberate choice, not a defect. To test a single path against one token, use the robots.txt tester.
Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.
What a plan actually costs
The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.
| Plan | Monthly | Credits | A good fit when |
|---|---|---|---|
| Starter | $4.99 | 10 | You look after one site and check it now and then |
| Pro | $14.99 | 40 | A few sites, plus the API and a handful of watched URLs |
| Studio | $49.99 | 150 | Client work that would burn through Pro mid-month |
| Agency | $99.99 | 400 | Many locations, reported under your own name |
A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.
Frequently asked questions
Is Googlebot an AI crawler?
Not by name, but its index feeds Google’s AI Overviews and AI Mode. Google-Extended, the token people associate with Google AI, is not a separate crawler: it only governs whether content Googlebot already fetched may be used for Gemini training and grounding.
Do AI crawlers run JavaScript?
Most AI vendors do not document JavaScript rendering for their crawlers. The safe assumption is that the text an answer engine can use is the text in your server HTML.
Can I see AI crawlers in my analytics?
Usually not. Crawlers rarely run analytics scripts. Look in server or CDN logs for the user-agent strings, and verify them against the vendor’s published IP ranges where one exists.
Sources
Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.
- OpenAI — Overview of OpenAI crawlers
- Anthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity — Perplexity crawlers
- Google Search Central — Google common crawlers (Google-Extended)
- Google Search Central — Google user-triggered fetchers
- Apple — About Applebot
- Meta — Meta web crawlers
- Common Crawl — CCBot
- Google Search Central — AI features and your website
- Crawlwise — The AI Crawler Access Report (top-10000 run, 2026-Q4)
Crawlwise crawler check vs checking by hand
Resolves robots.txt groups the way crawlers do (RFC 9309)
- Crawlwise
- Yes, per crawler token
- By hand
- By eye; group precedence is easy to misread
Fetches the page as each AI crawler
- Crawlwise
- Yes, one live request per agent
- By hand
- One curl at a time
Spots CDN and firewall blocks (403, challenge pages)
- Crawlwise
- Yes, with challenge fingerprints
- By hand
- Only if you read the response body
Separates a training opt-out from an accidental block
- Crawlwise
- Labelled separately
- By hand
- Up to you
Sees blocks that only happen in other regions or on real crawler IPs
- Crawlwise
- No — one location, one moment
- By hand
- Only from your own server logs
Meet the author

Adnan Arodiya
Crawlwise
Writes about what Crawlwise actually measures: on-page evidence, crawler access, and performance signals — with the limits stated up front.


