crawlwise.

Does google-analytics.com block AI crawlers?

google-analytics.com does not block any of the 17 AI crawlers we tested: each one is allowed by its robots.txt (checked 2026-10-01).

Nothing in its robots.txt keeps AI products out: its pages can be used for model training and fetched by ChatGPT, Claude and Perplexity when they answer a question and cite sources.

Every crawler we tested

CrawlerUsed forVerdictWhat we saw
GPTBot
OpenAI
ChatGPT (training)AllowedAllowed: robots.txt permits it and the page loaded.
OAI-SearchBot
OpenAI
ChatGPT search (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
ChatGPT-User
OpenAI
ChatGPT (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
ClaudeBot
Anthropic
Claude (training)AllowedAllowed: robots.txt permits it and the page loaded.
Claude-SearchBot
Anthropic
Claude search (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
Claude-User
Anthropic
Claude (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
Google-Extended
Google
Gemini (training)AllowedAllowed: robots.txt permits it and the page loaded.
PerplexityBot
Perplexity
Perplexity (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
Perplexity-User
Perplexity
Perplexity (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
CCBot
Common Crawl
Common Crawl (research)AllowedAllowed: robots.txt permits it and the page loaded.
Bytespider
ByteDance
ByteDance (training)AllowedAllowed: robots.txt permits it and the page loaded.
Applebot-Extended
Apple
Apple Intelligence (training)AllowedAllowed: robots.txt permits it and the page loaded.
Meta-ExternalAgent
Meta
Meta AI (training)AllowedAllowed: robots.txt permits it and the page loaded.
Amazonbot
Amazon
Alexa (training)AllowedAllowed: robots.txt permits it and the page loaded.
cohere-ai
Cohere
Cohere (retrieval)AllowedAllowed: robots.txt permits it and the page loaded.
Diffbot
Diffbot
Diffbot (research)AllowedAllowed: robots.txt permits it and the page loaded.
Timpibot
Timpi
Timpi (research)AllowedAllowed: robots.txt permits it and the page loaded.

Homepage SEO snapshot

From the Crawlwise SEO Index run, checked 2026-10-05. One fetch of the homepage, so it describes that page and not the whole site.

Homepagehttps://marketingplatform.google.com/about/analytics/
HTTPSyes
Title64 characters
Meta description215 characters
Canonicalpoints to itself
Noindexno
H1 elements1
Languageen
JSON-LD typesWebpage, ImageObject, Organization, BreadcrumbList, ListItem
hreflang alternates0
/sitemap.xmlnot measured
Response time224 ms

How this was measured

We read google-analytics.com’s robots.txt, resolved it with an RFC 9309 matcher for each crawler’s product token, then fetched the homepage once as each crawler. A 403 or a challenge page is a block even when robots.txt allows the crawler; a timeout or rate limit is “not established”, never a block. One probe, from one location, on 2026-10-01, so a rule changed since then is not reflected here. Full methodology.

Check your own site

The free checker runs the same probe on any URL: robots.txt resolved for every AI crawler, then a live fetch as each one, so a CDN or firewall block shows up even when robots.txt looks fine.

Check AI crawler access, free

Need it across every page? Site crawls start at $4.99 a month, and Pro adds monitoring that re-checks your URLs on a schedule.

Other sites