Which websites block Amazonbot?
31 of 237 popular sites we measured block Amazonbot (13.1%), checked 2026-10-01.
Amazonbot is Amazon’s crawler for alexa and product training (Alexa). Blocking it opts a site out of training data; it does not stop the retrieval crawlers that fetch pages for cited answers. More about Amazonbot.
The robots.txt rule
To block it:
User-agent: Amazonbot Disallow: /
To allow it explicitly:
User-agent: Amazonbot Allow: /
A server or CDN can still refuse the crawler when robots.txt allows it. That is why the checker fetches as the crawler as well as reading the file.
Sites that block Amazonbot
| Site | How |
|---|---|
| msn.com | robots.txt disallow |
| netflix.net | robots.txt disallow |
| fbcdn.net | robots.txt disallow |
| whatsapp.net | robots.txt disallow |
| instagram.com | robots.txt disallow |
| skybridge.click | robots.txt disallow |
| linkedin.com | robots.txt disallow |
| criteo.com | refused at the server |
| pubmatic.com | refused at the server |
| rubiconproject.com | refused at the server |
| netflix.com | robots.txt disallow |
| whatsapp.com | robots.txt disallow |
| snapchat.com | robots.txt disallow |
| creativecdn.com | robots.txt disallow |
| twitter.com | robots.txt disallow |
| pinterest.com | robots.txt disallow |
| reddit.com | robots.txt disallow |
| sharethrough.com | robots.txt disallow |
| netcraze.io | robots.txt disallow |
| juicyscore.com | robots.txt disallow |
| pinimg.com | robots.txt disallow |
| blismedia.com | refused at the server |
| t.co | robots.txt disallow |
| weather.com | robots.txt disallow |
| gravatar.com | robots.txt disallow |
| x.com | robots.txt disallow |
| meraki.com | robots.txt disallow |
| ad-score.com | robots.txt disallow |
| bidmachine.io | refused at the server |
| storygize.net | robots.txt disallow |
| starttest.com | robots.txt disallow |
Check your own site
The free checker runs the same probe on any URL: robots.txt resolved for every AI crawler, then a live fetch as each one, so a CDN or firewall block shows up even when robots.txt looks fine.
Need it across every page? Site crawls start at $4.99 a month, and Pro adds monitoring that re-checks your URLs on a schedule.
Other crawlers
- Sites that block GPTBot
- Sites that block OAI-SearchBot
- Sites that block ChatGPT-User
- Sites that block ClaudeBot
- Sites that block Claude-SearchBot
- Sites that block Claude-User
- Sites that block Google-Extended
- Sites that block PerplexityBot
- Sites that block Perplexity-User
- Sites that block CCBot
- Sites that block Bytespider
- Sites that block Applebot-Extended
- Sites that block Meta-ExternalAgent
- Sites that block cohere-ai
- Sites that block Diffbot
- Sites that block Timpibot