
How ChatGPT, Claude, Gemini, Perplexity, Copilot and Grok find web pages
Each assistant answers from a search index plus live fetches. ChatGPT search uses OpenAI’s OAI-SearchBot and third-party providers including Bing. Copilot is grounded on Bing. Gemini and AI Overviews use Google’s index. Perplexity runs PerplexityBot. Claude uses Claude-SearchBot, Claude-User and third-party search. xAI documents little about how Grok crawls.
Key takeaways
- Being cited starts with being in the index the assistant searches.
- Google and Bing indexing still matter for Gemini, AI Overviews, Copilot and ChatGPT.
- Each AI lab’s search crawler must be allowed for its own index.
- Where a vendor documents nothing, plan around the indexes you can verify.
When an assistant cites a page, it almost always found that page in a search index first, then read it — from the index, or with a live fetch. So the practical question is: whose index does each assistant search? Below is what each vendor documents, with anything we could not confirm marked as such.
Which index does each assistant use?
| Assistant | Index and crawlers (documented) | What to allow |
|---|---|---|
| ChatGPT search | OAI-SearchBot; third-party search providers, with Bing named by OpenAI; ChatGPT-User for user-requested pages | OAI-SearchBot, ChatGPT-User, Bingbot |
| Microsoft Copilot | The Bing index (Bingbot) | Bingbot |
| Google AI Overviews and AI Mode | Google Search index (Googlebot) | Googlebot; a page must be indexed and snippet-eligible |
| Gemini app | Google crawling, with Google-Extended governing training and grounding use | Googlebot; leave Google-Extended allowed if you want Gemini grounding |
| Perplexity | PerplexityBot index; Perplexity-User for live fetches | PerplexityBot, Perplexity-User |
| Claude | Claude-SearchBot index; Claude-User for live fetches; third-party search providers | Claude-SearchBot, Claude-User |
| Grok | xAI says Grok draws on X posts and web pages; no crawler documentation found | Nothing documented to target |
How does ChatGPT find pages?
OpenAI says OAI-SearchBot is “used to surface websites in search results in ChatGPT’s search features,” and its help center says ChatGPT search uses third-party search providers as well as content from partners, naming Bing. ChatGPT-User fetches pages for specific user requests. Details: GPTBot vs OAI-SearchBot vs ChatGPT-User.
How do Copilot and Gemini find pages?
Copilot answers are grounded on Bing’s index. Microsoft lets publishers limit generative use page by page: with NOCACHE, Microsoft said in 2023, only URL, title and snippet are shown in answers; with NOARCHIVE, content is not included or linked in answers. Gemini and AI Overviews sit on Google’s crawl. Google-Extended controls Gemini training and grounding in Gemini Apps and Vertex AI, but “does not impact a site’s inclusion in Google Search” — AI Overviews and AI Mode are part of Search. See what you can control in AI Overviews and AI Mode.
How do Perplexity and Claude find pages?
Perplexity runs PerplexityBot to surface and link sites in its results, and Perplexity-User to visit pages when a user asks a question. Claude uses Claude-SearchBot to index content for search quality and Claude-User to fetch pages when a person asks Claude something. Anthropic’s subprocessor list has named Brave Search since March 2025, which is widely read as Claude’s web-search provider; Anthropic does not document how its search results are assembled, so we treat that as a reasonable inference rather than a fact.
How does Grok find pages?
xAI says Grok “draws upon posts from X and webpages from the broader internet” and shows citations. We could not find xAI documentation of a crawler user agent, robots.txt token or IP ranges. Third-party lists name several strings, but without a vendor source we do not repeat them as fact.
You cannot be cited from an index you are not in. Work backwards from the index, not from the chatbot.
Adnan Arodiya, Crawlwise
What does this mean for being cited?
- Be indexed in Google and Bing: they feed AI Overviews, AI Mode, Gemini, Copilot, and part of ChatGPT search.
- Allow each AI lab’s search crawler and user-triggered fetcher, even if you block its training crawler — see which AI crawlers to allow.
- Make sure your CDN agrees with your robots.txt.
- Put the answer in server-rendered HTML, near the top, in plain language.
Many sites fail step 2 or 3 without knowing. In the Crawlwise AI Crawler Access Report (top-10000 run, 2026-Q4: the first 10,000 domains of the Cisco Umbrella top-one-million list, each probed once on October 1, 2026), 3,947 of 10,000 domains (39.5%) refused at least one retrieval crawler — the agents that feed answers rather than training. The list ranks hostnames by DNS popularity, so it includes API, CDN and update hosts that never serve a web page, and only the homepage path was probed. Treat the shares as a description of that list, not of the whole web.
How Crawlwise tests this
The AI crawler checker tests the retrieval agents for ChatGPT, Claude and Perplexity on one URL, alongside the training crawlers and Google-Extended, and says for each blocked row which product cannot cite you. It does not test Bingbot or Googlebot indexing; use Bing Webmaster Tools and Google Search Console for that.
Every verdict, what counts as a block, and what the checks cannot see are written up on the methodology page.
What a plan actually costs
The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.
| Plan | Monthly | Credits | A good fit when |
|---|---|---|---|
| Starter | $4.99 | 10 | You look after one site and check it now and then |
| Pro | $14.99 | 40 | A few sites, plus the API and a handful of watched URLs |
| Studio | $49.99 | 150 | Client work that would burn through Pro mid-month |
| Agency | $99.99 | 400 | Many locations, reported under your own name |
A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.
Frequently asked questions
Do I need to submit my site to ChatGPT or Claude?
There is no documented submission process for either. Make sure their search crawlers are allowed, and that you are indexed by the search providers they draw on.
Does Bing matter if nobody I know uses Bing?
Yes. Microsoft Copilot is grounded on the Bing index, and OpenAI names Bing among the providers ChatGPT search has used. Check your pages in Bing Webmaster Tools.
Can I block Grok?
We could not find xAI documentation of a crawler user agent or robots.txt token, so there is nothing documented to target. A firewall rule can only block what it can identify.
Sources
Vendor documentation was read on 2026-10-05. Vendors change these pages without notice, so check the original before you act on a detail.
- OpenAI — Overview of OpenAI crawlers
- OpenAI Help Center — Searching the web with ChatGPT
- Bing Webmaster Blog — New options for webmasters to control usage of their content in Bing Chat (Sept 2023)
- Google Search Central — Google common crawlers (Google-Extended)
- Google Search Central — AI features and your website
- Perplexity — Perplexity crawlers
- Anthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Anthropic Trust Center — subprocessor list
- Simon Willison — Anthropic Trust Center: Brave Search added as a subprocessor (March 2025)
- xAI — Bringing Grok to Everyone (web search and citations)
- Crawlwise — The AI Crawler Access Report (top-10000 run, 2026-Q4)
Crawlwise crawler check vs checking by hand
Resolves robots.txt groups the way crawlers do (RFC 9309)
- Crawlwise
- Yes, per crawler token
- By hand
- By eye; group precedence is easy to misread
Fetches the page as each AI crawler
- Crawlwise
- Yes, one live request per agent
- By hand
- One curl at a time
Spots CDN and firewall blocks (403, challenge pages)
- Crawlwise
- Yes, with challenge fingerprints
- By hand
- Only if you read the response body
Separates a training opt-out from an accidental block
- Crawlwise
- Labelled separately
- By hand
- Up to you
Sees blocks that only happen in other regions or on real crawler IPs
- Crawlwise
- No — one location, one moment
- By hand
- Only from your own server logs
Meet the author

Adnan Arodiya
Crawlwise
Writes about what Crawlwise actually measures: on-page evidence, crawler access, and performance signals — with the limits stated up front.


