crawlwise.

Does blismedia.com block AI crawlers?

blismedia.com blocks 15 of 17 AI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Perplexity-User, CCBot, Bytespider, Applebot-Extended, Meta-ExternalAgent, Amazonbot, cohere-ai, Diffbot and Timpibot), and allows 2 (checked 2026-10-01).

Retrieval crawlers are blocked too, so AI assistants that fetch a page live to quote and link it cannot read this site. Its content can still appear in answers from older training data or third-party copies, but not as a fresh, cited source.

Every crawler we tested

CrawlerUsed forVerdictWhat we saw
GPTBot
OpenAI
ChatGPT (training)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
OAI-SearchBot
OpenAI
ChatGPT search (retrieval)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
ChatGPT-User
OpenAI
ChatGPT (retrieval)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
ClaudeBot
Anthropic
Claude (training)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
Claude-SearchBot
Anthropic
Claude search (retrieval)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
Claude-User
Anthropic
Claude (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
Google-Extended
Google
Gemini (training)AllowedAllowed: robots.txt permits it and the page loaded.
PerplexityBot
Perplexity
Perplexity (retrieval)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
Perplexity-User
Perplexity
Perplexity (retrieval)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
CCBot
Common Crawl
Common Crawl (research)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
Bytespider
ByteDance
ByteDance (training)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
Applebot-Extended
Apple
Apple Intelligence (training)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
Meta-ExternalAgent
Meta
Meta AI (training)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
Amazonbot
Amazon
Alexa (training)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
cohere-ai
Cohere
Cohere (retrieval)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
Diffbot
Diffbot
Diffbot (research)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.
Timpibot
Timpi
Timpi (research)BlockedBlocked at the server: robots.txt permits it, but the live fetch was refused.

How that compares

How this was measured

We read blismedia.com’s robots.txt, resolved it with an RFC 9309 matcher for each crawler’s product token, then fetched the homepage once as each crawler. A 403 or a challenge page is a block even when robots.txt allows the crawler; a timeout or rate limit is “not established”, never a block. One probe, from one location, on 2026-10-01, so a rule changed since then is not reflected here. Full methodology.

Check your own site

The free checker runs the same probe on any URL: robots.txt resolved for every AI crawler, then a live fetch as each one, so a CDN or firewall block shows up even when robots.txt looks fine.

Check AI crawler access, free

Need it across every page? Site crawls start at $4.99 a month, and Pro adds monitoring that re-checks your URLs on a schedule.

Other sites