Does pubmatic.com block AI crawlers?
pubmatic.com blocks 14 of 17 AI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Google-Extended, PerplexityBot, CCBot, Bytespider, Applebot-Extended, Meta-ExternalAgent, Amazonbot, cohere-ai, Diffbot and Timpibot), and allows 3 (checked 2026-10-01).
Retrieval crawlers are blocked too, so AI assistants that fetch a page live to quote and link it cannot read this site. Its content can still appear in answers from older training data or third-party copies, but not as a fresh, cited source.
Every crawler we tested
| Crawler | Used for | Verdict | What we saw |
|---|---|---|---|
| GPTBot OpenAI | ChatGPT (training) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| OAI-SearchBot OpenAI | ChatGPT search (retrieval) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| ChatGPT-User OpenAI | ChatGPT (retrieval) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| ClaudeBot Anthropic | Claude (training) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| Claude-SearchBot Anthropic | Claude search (retrieval) | Allowed | Allowed: robots.txt permits it and the page loaded. |
| Claude-User Anthropic | Claude (retrieval) | Allowed | Allowed: robots.txt permits it and the page loaded. |
| Google-Extended Google | Gemini (training) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| PerplexityBot Perplexity | Perplexity (retrieval) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| Perplexity-User Perplexity | Perplexity (retrieval) | Allowed | Allowed: robots.txt permits it and the page loaded. |
| CCBot Common Crawl | Common Crawl (research) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| Bytespider ByteDance | ByteDance (training) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| Applebot-Extended Apple | Apple Intelligence (training) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| Meta-ExternalAgent Meta | Meta AI (training) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| Amazonbot Amazon | Alexa (training) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| cohere-ai Cohere | Cohere (retrieval) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| Diffbot Diffbot | Diffbot (research) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
| Timpibot Timpi | Timpi (research) | Blocked | Blocked at the server: robots.txt permits it, but the live fetch was refused. |
How that compares
- GPTBot is blocked by 38 of the 239 sites on these pages (15.9%).
- OAI-SearchBot is blocked by 26 of the 243 sites on these pages (10.7%).
- ChatGPT-User is blocked by 30 of the 240 sites on these pages (12.5%).
- ClaudeBot is blocked by 36 of the 240 sites on these pages (15%).
Homepage SEO snapshot
From the Crawlwise SEO Index run, checked 2026-10-05. One fetch of the homepage, so it describes that page and not the whole site.
| Homepage | https://pubmatic.com/ |
|---|---|
| HTTPS | yes |
| Title | 56 characters |
| Meta description | 159 characters |
| Canonical | points to itself |
| Noindex | no |
| H1 elements | 1 |
| Language | en-US |
| JSON-LD types | WebPage, ReadAction, BreadcrumbList, ListItem, WebSite, SearchAction, EntryPoint, PropertyValueSpecification |
| hreflang alternates | 2 |
| /sitemap.xml | valid XML sitemap |
| Response time | 19 ms |
How this was measured
We read pubmatic.com’s robots.txt, resolved it with an RFC 9309 matcher for each crawler’s product token, then fetched the homepage once as each crawler. A 403 or a challenge page is a block even when robots.txt allows the crawler; a timeout or rate limit is “not established”, never a block. One probe, from one location, on 2026-10-01, so a rule changed since then is not reflected here. Full methodology.
Check your own site
The free checker runs the same probe on any URL: robots.txt resolved for every AI crawler, then a live fetch as each one, so a CDN or firewall block shows up even when robots.txt looks fine.
Need it across every page? Site crawls start at $4.99 a month, and Pro adds monitoring that re-checks your URLs on a schedule.
Other sites
- criteo.com — blocks 1 of 17
- rubiconproject.com — blocks 2 of 17
- linkedin.com — blocks 14 of 17
- opera.com — blocks none of 17
- adobe.io — blocks none of 17
- pendo.io — blocks none of 17