Skip to content
How ChatGPT, Claude, Gemini, Perplexity, Copilot and Grok find web pages — illustrated banner

How ChatGPT, Claude, Gemini, Perplexity, Copilot and Grok find web pages

Each assistant answers from a search index plus live fetches. ChatGPT search uses OpenAI’s OAI-SearchBot and third-party providers including Bing. Copilot is grounded on Bing. Gemini and AI Overviews use Google’s index. Perplexity runs PerplexityBot. Claude uses Claude-SearchBot, Claude-User and third-party search. xAI documents little about how Grok crawls.

Share article:

Key takeaways

  • Being cited starts with being in the index the assistant searches.
  • Google and Bing indexing still matter for Gemini, AI Overviews, Copilot and ChatGPT.
  • Each AI lab’s search crawler must be allowed for its own index.
  • Where a vendor documents nothing, plan around the indexes you can verify.

When an assistant cites a page, it almost always found that page in a search index first, then read it — from the index, or with a live fetch. So the practical question is: whose index does each assistant search? Below is what each vendor documents, with anything we could not confirm marked as such.

Which index does each assistant use?

AssistantIndex and crawlers (documented)What to allow
ChatGPT searchOAI-SearchBot; third-party search providers, with Bing named by OpenAI; ChatGPT-User for user-requested pagesOAI-SearchBot, ChatGPT-User, Bingbot
Microsoft CopilotThe Bing index (Bingbot)Bingbot
Google AI Overviews and AI ModeGoogle Search index (Googlebot)Googlebot; a page must be indexed and snippet-eligible
Gemini appGoogle crawling, with Google-Extended governing training and grounding useGooglebot; leave Google-Extended allowed if you want Gemini grounding
PerplexityPerplexityBot index; Perplexity-User for live fetchesPerplexityBot, Perplexity-User
ClaudeClaude-SearchBot index; Claude-User for live fetches; third-party search providersClaude-SearchBot, Claude-User
GrokxAI says Grok draws on X posts and web pages; no crawler documentation foundNothing documented to target

How does ChatGPT find pages?

OpenAI says OAI-SearchBot is “used to surface websites in search results in ChatGPT’s search features,” and its help center says ChatGPT search uses third-party search providers as well as content from partners, naming Bing. ChatGPT-User fetches pages for specific user requests. Details: GPTBot vs OAI-SearchBot vs ChatGPT-User.

How do Copilot and Gemini find pages?

Copilot answers are grounded on Bing’s index. Microsoft lets publishers limit generative use page by page: with NOCACHE, Microsoft said in 2023, only URL, title and snippet are shown in answers; with NOARCHIVE, content is not included or linked in answers. Gemini and AI Overviews sit on Google’s crawl. Google-Extended controls Gemini training and grounding in Gemini Apps and Vertex AI, but “does not impact a site’s inclusion in Google Search” — AI Overviews and AI Mode are part of Search. See what you can control in AI Overviews and AI Mode.

How do Perplexity and Claude find pages?

Perplexity runs PerplexityBot to surface and link sites in its results, and Perplexity-User to visit pages when a user asks a question. Claude uses Claude-SearchBot to index content for search quality and Claude-User to fetch pages when a person asks Claude something. Anthropic’s subprocessor list has named Brave Search since March 2025, which is widely read as Claude’s web-search provider; Anthropic does not document how its search results are assembled, so we treat that as a reasonable inference rather than a fact.

How does Grok find pages?

xAI says Grok “draws upon posts from X and webpages from the broader internet” and shows citations. We could not find xAI documentation of a crawler user agent, robots.txt token or IP ranges. Third-party lists name several strings, but without a vendor source we do not repeat them as fact.

You cannot be cited from an index you are not in. Work backwards from the index, not from the chatbot.

Adnan Arodiya, Crawlwise

What does this mean for being cited?

  1. Be indexed in Google and Bing: they feed AI Overviews, AI Mode, Gemini, Copilot, and part of ChatGPT search.
  2. Allow each AI lab’s search crawler and user-triggered fetcher, even if you block its training crawler — see which AI crawlers to allow.
  3. Make sure your CDN agrees with your robots.txt.
  4. Put the answer in server-rendered HTML, near the top, in plain language.

Many sites fail step 2 or 3 without knowing. In the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026), 3,947 of 10,000 domains (39.5%) refused at least one retrieval crawler — the agents that feed answers rather than training. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.

How Crawlwise tests this

The AI crawler checker tests the retrieval agents for ChatGPT, Claude and Perplexity on one URL, alongside the training crawlers and Google-Extended, and says for each blocked row which product cannot cite you. It does not test Bingbot or Googlebot indexing; use Bing Webmaster Tools and Google Search Console for that.

Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.

What a plan actually costs

The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.

PlanMonthlyCreditsA good fit when
Starter$4.9910You look after one site and check it now and then
Pro$14.9940A few sites, plus the API and a handful of watched URLs
Studio$49.99150Client work that would burn through Pro mid-month
Agency$99.99400Many locations, reported under your own name

A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.

Frequently asked questions

Do I need to submit my site to ChatGPT or Claude?

There is no documented submission process for either. Make sure their search crawlers are allowed, and that you are indexed by the search providers they draw on.

Does Bing matter if nobody I know uses Bing?

Yes. Microsoft Copilot is grounded on the Bing index, and OpenAI names Bing among the providers ChatGPT search has used. Check your pages in Bing Webmaster Tools.

Can I block Grok?

We could not find xAI documentation of a crawler user agent or robots.txt token, so there is nothing documented to target. A firewall rule can only block what it can identify.

Sources

Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.

Crawlwise crawler check vs checking by hand

  • Resolves robots.txt groups the way crawlers do (RFC 9309)

    Crawlwise
    Yes, per crawler token
    By hand
    By eye; group precedence is easy to misread
  • Fetches the page as each AI crawler

    Crawlwise
    Yes, one live request per agent
    By hand
    One curl at a time
  • Spots CDN and firewall blocks (403, challenge pages)

    Crawlwise
    Yes, with challenge fingerprints
    By hand
    Only if you read the response body
  • Separates a training opt-out from an accidental block

    Crawlwise
    Labelled separately
    By hand
    Up to you
  • Sees blocks that only happen in other regions or on real crawler IPs

    Crawlwise
    No — one location, one moment
    By hand
    Only from your own server logs

Meet the author

Photo of Adnan Arodiya