crawlwise.

Does yahoo.com block AI crawlers?

yahoo.com blocks 9 of 17 AI crawlers (GPTBot, ChatGPT-User, ClaudeBot, Google-Extended, PerplexityBot, CCBot, Bytespider, cohere-ai and Diffbot), and allows 5 (checked 2026-10-01).

Retrieval crawlers are blocked too, so AI assistants that fetch a page live to quote and link it cannot read this site. Its content can still appear in answers from older training data or third-party copies, but not as a fresh, cited source.

Every crawler we tested

CrawlerUsed forVerdictWhat we saw
GPTBot
OpenAI
ChatGPT (training)BlockedBlocked: robots.txt disallows it.
OAI-SearchBot
OpenAI
ChatGPT search (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
ChatGPT-User
OpenAI
ChatGPT (retrieval)BlockedBlocked: robots.txt disallows it.
ClaudeBot
Anthropic
Claude (training)BlockedBlocked: robots.txt disallows it.
Claude-SearchBot
Anthropic
Claude search (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
Claude-User
Anthropic
Claude (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
Google-Extended
Google
Gemini (training)BlockedBlocked: robots.txt disallows it.
PerplexityBot
Perplexity
Perplexity (retrieval)BlockedBlocked: robots.txt disallows it.
Perplexity-User
Perplexity
Perplexity (retrieval)Not establishedAllowed by robots.txt; the live fetch was rate limited, so access was not confirmed.
CCBot
Common Crawl
Common Crawl (research)BlockedBlocked: robots.txt disallows it.
Bytespider
ByteDance
ByteDance (training)BlockedBlocked: robots.txt disallows it.
Applebot-Extended
Apple
Apple Intelligence (training)AllowedAllowed: robots.txt permits it and the page loaded.
Meta-ExternalAgent
Meta
Meta AI (training)Not establishedAllowed by robots.txt; the live fetch was rate limited, so access was not confirmed.
Amazonbot
Amazon
Alexa (training)Not establishedAllowed by robots.txt; the live fetch timed out, so access was not confirmed.
cohere-ai
Cohere
Cohere (retrieval)BlockedBlocked: robots.txt disallows it.
Diffbot
Diffbot
Diffbot (research)BlockedBlocked: robots.txt disallows it.
Timpibot
Timpi
Timpi (research)AllowedAllowed: robots.txt permits it and the page loaded.

How that compares

Homepage SEO snapshot

From the Crawlwise SEO Index run, checked 2026-10-05. One fetch of the homepage, so it describes that page and not the whole site.

Homepagehttps://www.yahoo.com/
HTTPSyes
Title71 characters
Meta description127 characters
Canonicalpoints to itself
Noindexno
H1 elements0
Languageen-US
JSON-LD typesWebSite, NewsMediaOrganization, ImageObject, Organization
hreflang alternates0
/sitemap.xmlnot found
Response time210 ms

How this was measured

We read yahoo.com’s robots.txt, resolved it with an RFC 9309 matcher for each crawler’s product token, then fetched the homepage once as each crawler. A 403 or a challenge page is a block even when robots.txt allows the crawler; a timeout or rate limit is “not established”, never a block. One probe, from one location, on 2026-10-01, so a rule changed since then is not reflected here. Full methodology.

Check your own site

The free checker runs the same probe on any URL: robots.txt resolved for every AI crawler, then a live fetch as each one, so a CDN or firewall block shows up even when robots.txt looks fine.

Check AI crawler access, free

Need it across every page? Site crawls start at $4.99 a month, and Pro adds monitoring that re-checks your URLs on a schedule.

Other sites