Skip to content
What is an AI crawler? Training, search, and user-triggered fetchers — illustrated banner

What is an AI crawler? Training, search, and user-triggered fetchers

An AI crawler is an automated client run by an AI company that fetches web pages. There are three kinds: training crawlers that gather text for future models, search crawlers that build an index for cited answers, and user-triggered fetchers that load one page because a person asked. Each kind needs a separate decision.

Share article:

Key takeaways

  • Training, search, and user-triggered agents are separate bots with separate robots.txt tokens.
  • Blocking a training crawler does not, by itself, remove you from AI answers.
  • Google-Extended and Applebot-Extended are control tokens, not crawlers you will see in logs.
  • User-triggered fetchers may not follow robots.txt; vendors say so in their own docs.

“AI crawler” gets used for any bot with an AI company’s name on it. That hides the distinction that matters to a site owner: what the fetched page is used for. A page fetched to train a model, a page fetched to build a search index, and a page fetched because a person pasted your URL into a chat are three different events with three different consequences.

What counts as an AI crawler?

An AI crawler is any automated HTTP client operated by an AI company that requests your pages and identifies itself with a documented user agent or robots.txt product token. The vendors that document theirs — OpenAI, Anthropic, Google, Apple, Perplexity, Meta and Common Crawl — each publish a page listing the names and what each one is for. Those pages are the only reliable source; third-party “bot directories” often copy each other and go stale.

What are the three kinds of AI crawler?

KindWhat it does with your pageDocumented examplesrobots.txt
TrainingCollects content that may be used to train future modelsGPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), Meta-ExternalAgent (Meta)Followed, per each vendor
Search / indexBuilds the index an assistant searches before it answers and citesOAI-SearchBot, Claude-SearchBot, PerplexityBotFollowed, per each vendor
User-triggeredFetches one page at the moment a person’s request needs itChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher, Google-AgentOpenAI says rules “may not apply”; Perplexity and Google say these fetchers “generally ignore” them; Meta says it “may bypass” them

Two names sit outside the table because they are not crawlers at all. Google-Extended “doesn’t have a separate HTTP request user agent string”, according to Google; it is a robots.txt token that controls whether content Google already crawls may be used for Gemini training and grounding. Applebot-Extended works the same way for Apple: Apple says it “does not crawl webpages”. You will never see either in a log file, and blocking them does not change what Googlebot or Applebot fetch.

Why does the distinction matter?

Because the trade-offs point in opposite directions. Opting out of training is a reasonable business decision with no direct effect on whether assistants can cite you. Blocking a search crawler removes you from that assistant’s index: OpenAI says sites that opt out of OAI-SearchBot “will not be shown in ChatGPT search answers”, and Anthropic says disabling Claude-SearchBot “prevents our system from indexing your content for search optimization”. Mixing the two up — usually with one broad rule — is how sites disappear from answers they wanted to be in. Which AI crawlers should you allow? walks through that decision vendor by vendor.

Ask what the fetched page is used for, not whose bot it is. That one question sorts almost every AI crawler decision.

Adnan Arodiya, Crawlwise

How is an AI crawler different from Googlebot or Bingbot?

Classic search crawlers are now AI-search crawlers too. Google’s AI Overviews and AI Mode are built on the regular Google index, and Google says a page must be indexed and eligible for a snippet to appear as a supporting link. Microsoft’s Copilot answers are grounded on the Bing index. So the most important “AI crawlers” for many sites are still Googlebot and Bingbot — see how ChatGPT, Claude, Gemini, Perplexity, Copilot and Grok find pages.

How many sites block AI crawlers today?

From the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026): 4,159 of 10,000 domains (41.6%) refused at least one of the 17 AI crawlers we test, either with a robots.txt disallow or with a refused or error response to the live fetch. Only 1,837 of the 10,000 served a robots.txt we could read as one. Among those, 785 (42.7%) disallow GPTBot at the homepage and 736 (40.1%) disallow Claude-SearchBot. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.

How Crawlwise tests this

The free AI crawler checker reads your robots.txt with an RFC 9309 matcher, then requests your URL once with each crawler’s user agent and classifies the answer: allowed, blocked, challenged, or not established. A training opt-out in robots.txt is labelled as a deliberate choice, not a defect. To test a single path against one token, use the robots.txt tester.

Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.

What a plan actually costs

The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.

PlanMonthlyCreditsA good fit when
Starter$4.9910You look after one site and check it now and then
Pro$14.9940A few sites, plus the API and a handful of watched URLs
Studio$49.99150Client work that would burn through Pro mid-month
Agency$99.99400Many locations, reported under your own name

A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.

Frequently asked questions

Is Googlebot an AI crawler?

Not by name, but its index feeds Google’s AI Overviews and AI Mode. Google-Extended, the token people associate with Google AI, is not a separate crawler: it only governs whether content Googlebot already fetched may be used for Gemini training and grounding.

Do AI crawlers run JavaScript?

Most AI vendors do not document JavaScript rendering for their crawlers. The safe assumption is that the text an answer engine can use is the text in your server HTML.

Can I see AI crawlers in my analytics?

Usually not. Crawlers rarely run analytics scripts. Look in server or CDN logs for the user-agent strings, and verify them against the vendor’s published IP ranges where one exists.

Sources

Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.

Crawlwise crawler check vs checking by hand

  • Resolves robots.txt groups the way crawlers do (RFC 9309)

    Crawlwise
    Yes, per crawler token
    By hand
    By eye; group precedence is easy to misread
  • Fetches the page as each AI crawler

    Crawlwise
    Yes, one live request per agent
    By hand
    One curl at a time
  • Spots CDN and firewall blocks (403, challenge pages)

    Crawlwise
    Yes, with challenge fingerprints
    By hand
    Only if you read the response body
  • Separates a training opt-out from an accidental block

    Crawlwise
    Labelled separately
    By hand
    Up to you
  • Sees blocks that only happen in other regions or on real crawler IPs

    Crawlwise
    No — one location, one moment
    By hand
    Only from your own server logs

Meet the author

Photo of Adnan Arodiya