Which websites block Google-Extended?
38 of 244 popular sites we measured block Google-Extended (15.6%), checked 2026-10-01.
Google-Extended is Google’s crawler for gemini / vertex training opt-in. robots.txt token; the live fetch uses googlebot (Gemini). Blocking it opts a site out of training data; it does not stop the retrieval crawlers that fetch pages for cited answers. More about Google-Extended.
The robots.txt rule
To block it:
User-agent: Google-Extended Disallow: /
To allow it explicitly:
User-agent: Google-Extended Allow: /
A server or CDN can still refuse the crawler when robots.txt allows it. That is why the checker fetches as the crawler as well as reading the file.
Sites that block Google-Extended
| Site | How |
|---|---|
| msn.com | robots.txt disallow |
| fbcdn.net | robots.txt disallow |
| amazon.com | robots.txt disallow |
| whatsapp.net | robots.txt disallow |
| instagram.com | robots.txt disallow |
| yahoo.com | robots.txt disallow |
| skybridge.click | robots.txt disallow |
| linkedin.com | robots.txt disallow |
| pubmatic.com | refused at the server |
| whatsapp.com | robots.txt disallow |
| snapchat.com | robots.txt disallow |
| creativecdn.com | robots.txt disallow |
| chatgpt.com | robots.txt disallow |
| twitter.com | robots.txt disallow |
| pinterest.com | robots.txt disallow |
| reddit.com | robots.txt disallow |
| sharethrough.com | robots.txt disallow |
| amazon.co.za | robots.txt disallow |
| netcraze.io | robots.txt disallow |
| juicyscore.com | robots.txt disallow |
| pinimg.com | robots.txt disallow |
| globalsign.com | robots.txt disallow |
| samsung.com | refused at the server |
| t.co | robots.txt disallow |
| cornell.edu | refused at the server |
| gravatar.com | robots.txt disallow |
| x.com | robots.txt disallow |
| wp.com | refused at the server |
| claude.ai | robots.txt disallow |
| wikipedia.com | refused at the server |
| meraki.com | robots.txt disallow |
| wikipedia.org | refused at the server |
| ad-score.com | robots.txt disallow |
| onedrive.com | refused at the server |
| yahoo.co.jp | robots.txt disallow |
| redd.it | refused at the server |
| storygize.net | robots.txt disallow |
| starttest.com | robots.txt disallow |
Check your own site
The free checker runs the same probe on any URL: robots.txt resolved for every AI crawler, then a live fetch as each one, so a CDN or firewall block shows up even when robots.txt looks fine.
Need it across every page? Site crawls start at $4.99 a month, and Pro adds monitoring that re-checks your URLs on a schedule.
Other crawlers
- Sites that block GPTBot
- Sites that block OAI-SearchBot
- Sites that block ChatGPT-User
- Sites that block ClaudeBot
- Sites that block Claude-SearchBot
- Sites that block Claude-User
- Sites that block PerplexityBot
- Sites that block Perplexity-User
- Sites that block CCBot
- Sites that block Bytespider
- Sites that block Applebot-Extended
- Sites that block Meta-ExternalAgent
- Sites that block Amazonbot
- Sites that block cohere-ai
- Sites that block Diffbot
- Sites that block Timpibot