crawlwise.

Does yandex.ru block AI crawlers?

yandex.ru does not block any of the 17 AI crawlers we tested: each one is allowed by its robots.txt, though 6 live fetches did not complete (checked 2026-10-01).

Nothing in its robots.txt keeps AI products out: its pages can be used for model training and fetched by ChatGPT, Claude and Perplexity when they answer a question and cite sources.

Every crawler we tested

CrawlerUsed forVerdictWhat we saw
GPTBot
OpenAI
ChatGPT (training)AllowedAllowed: robots.txt permits it and the page loaded.
OAI-SearchBot
OpenAI
ChatGPT search (retrieval)Not establishedAllowed by robots.txt; the live fetch timed out, so access was not confirmed.
ChatGPT-User
OpenAI
ChatGPT (retrieval)Not establishedAllowed by robots.txt; the live fetch timed out, so access was not confirmed.
ClaudeBot
Anthropic
Claude (training)Not establishedAllowed by robots.txt; the live fetch timed out, so access was not confirmed.
Claude-SearchBot
Anthropic
Claude search (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
Claude-User
Anthropic
Claude (retrieval)Not establishedAllowed by robots.txt; the live fetch timed out, so access was not confirmed.
Google-Extended
Google
Gemini (training)AllowedAllowed: robots.txt permits it and the page loaded.
PerplexityBot
Perplexity
Perplexity (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
Perplexity-User
Perplexity
Perplexity (retrieval)Not establishedAllowed by robots.txt; the live fetch timed out, so access was not confirmed.
CCBot
Common Crawl
Common Crawl (research)AllowedAllowed: robots.txt permits it and the page loaded.
Bytespider
ByteDance
ByteDance (training)AllowedAllowed: robots.txt permits it and the page loaded.
Applebot-Extended
Apple
Apple Intelligence (training)AllowedAllowed: robots.txt permits it and the page loaded.
Meta-ExternalAgent
Meta
Meta AI (training)AllowedAllowed: robots.txt permits it and the page loaded.
Amazonbot
Amazon
Alexa (training)AllowedAllowed: robots.txt permits it and the page loaded.
cohere-ai
Cohere
Cohere (retrieval)Not establishedAllowed by robots.txt; the live fetch timed out, so access was not confirmed.
Diffbot
Diffbot
Diffbot (research)AllowedAllowed: robots.txt permits it and the page loaded.
Timpibot
Timpi
Timpi (research)AllowedAllowed: robots.txt permits it and the page loaded.

Homepage SEO snapshot

From the Crawlwise SEO Index run, checked 2026-10-05. One fetch of the homepage, so it describes that page and not the whole site.

Homepagehttps://sso.passport.yandex.ru/push?uuid=224a4998-aacb-4eff-88da-620beae4bcee&retpath=https%3A%2F%2Fdzen.ru%2F%3Fyredirect%3Dtrue%26is_autologin_ya%3Dtrue
HTTPSyes
Titlemissing
Meta descriptionmissing
Canonicalnone declared
Noindexno
H1 elements0
Languagenot declared
JSON-LD typesnone
hreflang alternates0
/sitemap.xmlnot measured
Response time756 ms

How this was measured

We read yandex.ru’s robots.txt, resolved it with an RFC 9309 matcher for each crawler’s product token, then fetched the homepage once as each crawler. A 403 or a challenge page is a block even when robots.txt allows the crawler; a timeout or rate limit is “not established”, never a block. One probe, from one location, on 2026-10-01, so a rule changed since then is not reflected here. Full methodology.

Check your own site

The free checker runs the same probe on any URL: robots.txt resolved for every AI crawler, then a live fetch as each one, so a CDN or firewall block shows up even when robots.txt looks fine.

Check AI crawler access, free

Need it across every page? Site crawls start at $4.99 a month, and Pro adds monitoring that re-checks your URLs on a schedule.

Other sites