Does collegeboard.org block AI crawlers?
collegeboard.org serves a bot challenge to 1, and allows 0 (checked 2026-10-01).
The blocks are on research and corpus crawlers rather than on the assistants people use to search, so citation in ChatGPT, Claude and Perplexity is not affected by them.
Every crawler we tested
| Crawler | Used for | Verdict | What we saw |
|---|---|---|---|
| GPTBot OpenAI | ChatGPT (training) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| OAI-SearchBot OpenAI | ChatGPT search (retrieval) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| ChatGPT-User OpenAI | ChatGPT (retrieval) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| ClaudeBot Anthropic | Claude (training) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| Claude-SearchBot Anthropic | Claude search (retrieval) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| Claude-User Anthropic | Claude (retrieval) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| Google-Extended Google | Gemini (training) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| PerplexityBot Perplexity | Perplexity (retrieval) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| Perplexity-User Perplexity | Perplexity (retrieval) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| CCBot Common Crawl | Common Crawl (research) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| Bytespider ByteDance | ByteDance (training) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| Applebot-Extended Apple | Apple Intelligence (training) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| Meta-ExternalAgent Meta | Meta AI (training) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| Amazonbot Amazon | Alexa (training) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| cohere-ai Cohere | Cohere (retrieval) | Challenged | Challenged: robots.txt permits it, but the fetch got a bot challenge page instead of the content. |
| Diffbot Diffbot | Diffbot (research) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
| Timpibot Timpi | Timpi (research) | Not established | Allowed by robots.txt; the live fetch timed out, so access was not confirmed. |
Homepage SEO snapshot
From the Crawlwise SEO Index run, checked 2026-10-05. One fetch of the homepage, so it describes that page and not the whole site.
| Homepage | https://www.collegeboard.org/ |
|---|---|
| HTTPS | yes |
| Title | 59 characters |
| Meta description | 154 characters |
| Canonical | points to itself |
| Noindex | no |
| H1 elements | 1 |
| Language | en |
| JSON-LD types | none |
| hreflang alternates | 0 |
| /sitemap.xml | valid XML sitemap |
| Response time | 26 ms |
How this was measured
We read collegeboard.org’s robots.txt, resolved it with an RFC 9309 matcher for each crawler’s product token, then fetched the homepage once as each crawler. A 403 or a challenge page is a block even when robots.txt allows the crawler; a timeout or rate limit is “not established”, never a block. One probe, from one location, on 2026-10-01, so a rule changed since then is not reflected here. Full methodology.
Check your own site
The free checker runs the same probe on any URL: robots.txt resolved for every AI crawler, then a live fetch as each one, so a CDN or firewall block shows up even when robots.txt looks fine.
Need it across every page? Site crawls start at $4.99 a month, and Pro adds monitoring that re-checks your URLs on a schedule.
Other sites
- hcaptcha.com — blocks none of 17
- vungle.com — blocks none of 17
- richaudience.com — blocks none of 17
- instructure.com — blocks none of 17
- trustarc.com — blocks none of 17
- blismedia.com — blocks 15 of 17