Skip to content
Which AI crawlers should you allow? A vendor-by-vendor decision guide — illustrated banner

Which AI crawlers should you allow? A vendor-by-vendor decision guide

If you want to be cited in AI answers, allow the search and user-triggered agents: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User. Training crawlers and tokens — GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Bytespider, Meta-ExternalAgent — are a separate decision. Blocking them opts out of training, not out of answers.

Share article:

Key takeaways

  • Each major vendor separates training from search in its own documentation.
  • Blocking a training token is a valid opt-out and does not remove you from that vendor’s answers.
  • Most sites that block GPTBot also block OAI-SearchBot — usually by accident of a broad rule.
  • Bytespider and xAI document little; decide on them without assuming what they do.

There is no single right policy. A news publisher negotiating licences and a SaaS company that wants to be recommended in ChatGPT should not have the same robots.txt. What every site needs is a policy that matches its intent — and the most common failure is a rule that blocks answer visibility when the owner only meant to block training.

What should you allow if you want to be cited?

Your goalAllowBlock (optional)
Be cited in AI answers, opt out of trainingOAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot, BingbotGPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Bytespider, Meta-ExternalAgent
Maximum reach, including trainingEverythingNothing
Keep content out of AI products entirelyGooglebot and Bingbot for classic searchAll AI agents, plus snippet controls for AI Overviews and Copilot

Ready-to-paste files for each row are in robots.txt for AI crawlers: copy-paste examples.

What does each vendor’s split look like?

OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User

GPTBot is for training: OpenAI says disallowing it “indicates a site’s content should not be used in training generative AI foundation models.” OAI-SearchBot surfaces sites in ChatGPT search, and opted-out sites “will not be shown in ChatGPT search answers.” ChatGPT-User handles user actions, and “robots.txt rules may not apply.” OpenAI states that “each setting is independent of the others.” Full breakdown: GPTBot vs OAI-SearchBot vs ChatGPT-User.

Anthropic: ClaudeBot, Claude-SearchBot, Claude-User

Blocking ClaudeBot signals that future content “should be excluded from our AI model training datasets.” Claude-SearchBot indexes content to improve search results, and Claude-User fetches pages when a person asks Claude a question. Anthropic says blocking Claude-User prevents retrieval for those user-initiated questions and can reduce visibility. Anthropic says it honors robots.txt and the non-standard Crawl-delay extension.

Perplexity: PerplexityBot, Perplexity-User

PerplexityBot surfaces and links sites in Perplexity search and, per Perplexity, “is not used to crawl content for AI foundation models.” Perplexity-User visits pages when a user asks a question; “this fetcher generally ignores robots.txt rules.” There is no training-only token to block.

Google: Google-Extended

Google-Extended is a token, not a crawler. It controls use of crawled content for Gemini model training and for grounding in Gemini Apps and Vertex AI. Google says it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal.” AI Overviews and AI Mode are part of Search, so they are not covered by it — see what you can control in AI Overviews and AI Mode.

Apple: Applebot-Extended

Also a token, not a crawler. Disallowing it keeps your content out of Apple’s generative model training, and Apple says such pages “can still be included in search results” in Siri, Spotlight and Safari, which Applebot powers.

Common Crawl, ByteDance, Meta

CCBot builds Common Crawl’s open web archive. Common Crawl documents a standard robots.txt block. The archive is free for anyone to download, so blocking CCBot is the broadest single training opt-out — and it also removes you from research uses. Bytespider is ByteDance’s crawler; we could not find English documentation of its purpose or robots.txt policy, and independent log analyses report inconsistent compliance, so treat robots.txt as a request and enforce at the CDN if you care. Meta-ExternalAgent crawls “for use cases such as training foundation AI models or improving products by indexing content directly” — mixed purpose — while Meta-ExternalFetcher fetches links at a user’s request and “may bypass robots.txt rules.”

Blocking training is a business decision. Blocking answer visibility is usually an accident. Write down which one you meant.

Adnan Arodiya, Crawlwise

From the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026), among the 1,837 domains with a readable robots.txt:

  • 785 (42.7%) disallow GPTBot at the homepage. Of those 785, 728 also disallow OAI-SearchBot. Only 57 block GPTBot while allowing OAI-SearchBot — the “opt out of training, stay in answers” policy.
  • 760 (41.4%) disallow ClaudeBot; only 33 block ClaudeBot while allowing Claude-SearchBot.
  • 689 (37.5%) disallow all 17 AI crawlers we test, which in practice means a blanket rule such as User-agent: * with Disallow: /, not an AI policy. Only 140 (7.6%) treat crawlers differently from one another.

In other words, deliberate split policies are rare. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.

How Crawlwise tests this

Paste your homepage into the AI crawler checker. Each row names the crawler, its owner and purpose, the robots.txt result and what the live request returned. Training opt-outs are labelled as deliberate. If a search or user agent shows “blocked” and you did not mean it, the row tells you whether robots.txt or a 403 from your CDN caused it. Confirm the exact rule with the robots.txt tester, and compare with the full report.

Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.

What a plan actually costs

The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.

PlanMonthlyCreditsA good fit when
Starter$4.9910You look after one site and check it now and then
Pro$14.9940A few sites, plus the API and a handful of watched URLs
Studio$49.99150Client work that would burn through Pro mid-month
Agency$99.99400Many locations, reported under your own name

A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.

Frequently asked questions

If I block GPTBot, will ChatGPT stop citing me?

Not because of that rule. OpenAI documents GPTBot, OAI-SearchBot and ChatGPT-User as independent settings. ChatGPT search citations depend on OAI-SearchBot access.

Does blocking Google-Extended remove me from AI Overviews?

No. Google says Google-Extended does not affect inclusion in Google Search. AI Overviews and AI Mode are controlled with indexing and snippet controls instead.

Should I block Bytespider?

That is a business call. We could not find English-language documentation from ByteDance describing its purpose, so we cannot tell you what allowing it buys you. If you block it, enforce the block at your CDN as well as in robots.txt.

Sources

Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.

Crawlwise crawler check vs checking by hand

  • Resolves robots.txt groups the way crawlers do (RFC 9309)

    Crawlwise
    Yes, per crawler token
    By hand
    By eye; group precedence is easy to misread
  • Fetches the page as each AI crawler

    Crawlwise
    Yes, one live request per agent
    By hand
    One curl at a time
  • Spots CDN and firewall blocks (403, challenge pages)

    Crawlwise
    Yes, with challenge fingerprints
    By hand
    Only if you read the response body
  • Separates a training opt-out from an accidental block

    Crawlwise
    Labelled separately
    By hand
    Up to you
  • Sees blocks that only happen in other regions or on real crawler IPs

    Crawlwise
    No — one location, one moment
    By hand
    Only from your own server logs

Meet the author

Photo of Adnan Arodiya