AI crawler · research · Common Crawl
What is CCBot?
Common Crawl's bot. It builds an open web archive that many research and training pipelines read indirectly.
How this helps you show up
Letting CCBot through spreads your content into a widely used corpus. That can matter if you want your version of the facts to exist in datasets others train on — a long-game visibility play, not a day-one traffic lever.
If it cannot reach your page
Future Common Crawl snapshots skip you. Models already trained on older snapshots do not forget you overnight — but new training runs might never see your latest pages.
The facts
- robots.txt token
- CCBot
- Owner
- Common Crawl
- Product it serves
- Common Crawl
- Category
- research — an open corpus or knowledge graph
- Training opt-out
- Documented — only affects training-style use, not every AI product
- User agent as sent
CCBot/2.0 (https://commoncrawl.org/faq/)
robots.txt is a request, not a lock. Our checker fetches your page as CCBot and reports what actually came back — status code, challenge page, or real HTML.
Check your site in a minute
The AI crawler checker tests CCBot plus the other answer-engine bots on one URL. No account. Want robots.txt logic only? Use the robots.txt tester.
Other AI crawlers
- GPTBot — OpenAI
- OAI-SearchBot — OpenAI
- ChatGPT-User — OpenAI
- ClaudeBot — Anthropic
- Claude-SearchBot — Anthropic
- Claude-User — Anthropic
- Google-Extended — Google
- PerplexityBot — Perplexity
- Perplexity-User — Perplexity
- Bytespider — ByteDance
- Applebot-Extended — Apple
- Meta-ExternalAgent — Meta
- Amazonbot — Amazon
- cohere-ai — Cohere
- Diffbot — Diffbot
- Timpibot — Timpi