crawlwise.

Which websites block Bytespider?

44 of 238 popular sites we measured block Bytespider (18.5%), checked 2026-10-01.

Bytespider is ByteDance’s crawler for training corpus (ByteDance). Blocking it opts a site out of training data; it does not stop the retrieval crawlers that fetch pages for cited answers. More about Bytespider.

The robots.txt rule

To block it:

User-agent: Bytespider
Disallow: /

To allow it explicitly:

User-agent: Bytespider
Allow: /

A server or CDN can still refuse the crawler when robots.txt allows it. That is why the checker fetches as the crawler as well as reading the file.

Sites that block Bytespider

SiteHow
msn.comrobots.txt disallow
facebook.comrobots.txt disallow
netflix.netrobots.txt disallow
fbcdn.netrobots.txt disallow
amazon.comrobots.txt disallow
instagram.comrobots.txt disallow
yahoo.comrobots.txt disallow
skybridge.clickrobots.txt disallow
linkedin.comrobots.txt disallow
pubmatic.comrefused at the server
netflix.comrobots.txt disallow
snapchat.comrobots.txt disallow
spotify.comrefused at the server
amplitude.comrobots.txt disallow
creativecdn.comrobots.txt disallow
chatgpt.comrobots.txt disallow
twitter.comrobots.txt disallow
pinterest.comrobots.txt disallow
reddit.comrobots.txt disallow
sharethrough.comrobots.txt disallow
amazon.co.zarobots.txt disallow
netcraze.iorobots.txt disallow
juicyscore.comrobots.txt disallow
pinimg.comrobots.txt disallow
unity3d.comrobots.txt disallow
globalsign.comrobots.txt disallow
moloco.comrobots.txt disallow
kargo.comrefused at the server
blismedia.comrefused at the server
t.corobots.txt disallow
weather.comrobots.txt disallow
gravatar.comrobots.txt disallow
github.comrobots.txt disallow
x.comrobots.txt disallow
canva.comrobots.txt disallow
forter.comrobots.txt disallow
esxdos.orgrefused at the server
amazonvideo.comrobots.txt disallow
meraki.comrobots.txt disallow
ad-score.comrobots.txt disallow
mediavine.comrefused at the server
vimeo.comrobots.txt disallow
storygize.netrobots.txt disallow
starttest.comrobots.txt disallow

Check your own site

The free checker runs the same probe on any URL: robots.txt resolved for every AI crawler, then a live fetch as each one, so a CDN or firewall block shows up even when robots.txt looks fine.

Check AI crawler access, free

Need it across every page? Site crawls start at $4.99 a month, and Pro adds monitoring that re-checks your URLs on a schedule.

Other crawlers