# Crawlwise

> SEO audit tool and AI crawler checker. Paste a URL and see what search engines and AI crawlers such as GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot can actually read.

Crawlwise audits a page for search discoverability, structured data, live crawler access and measured performance. Every finding shows the evidence it was based on: the HTML, the response header, the robots.txt rule or the live fetch. When something cannot be measured, the report says "unmeasured" instead of guessing a value. The free tools below need no account; full audits, site crawls, monitoring and the API use credits.

## Product

- [Home](https://crawlwise.site/): Run an SEO and AI crawler audit on any public URL
- [Plans](https://crawlwise.site/plans): Pricing, credits and what a full audit costs
- [Methodology](https://crawlwise.site/methodology): What every check measures, how, and what it cannot measure
- [Check catalogue (JSON)](https://crawlwise.site/methodology/checks.json): Machine-readable definition of every check
- [Limits](https://crawlwise.site/limits): What an audit cannot see, what the checklist score is not, and why a metric can be missing

## Free tools

- [Schema markup generator](https://crawlwise.site/tools/schema): Build JSON-LD for supported page types, with required-field checks, ready to paste into your head.
- [Meta tag generator](https://crawlwise.site/tools/meta-tags): Title, description, robots directives, and canonical in one HTML snippet.
- [robots.txt generator](https://crawlwise.site/tools/robots-txt): Write crawl rules for search and retrieval bots. Training preferences are a separate decision.
- [Hreflang tag generator](https://crawlwise.site/tools/hreflang): Language and region alternates as HTML, HTTP headers, or XML sitemap markup.
- [Redirect generator](https://crawlwise.site/tools/redirects): 301 and 302 rules for Apache, Nginx, Next.js, and Vercel, including bulk import.
- [XML sitemap generator](https://crawlwise.site/tools/sitemap-generator): A valid urlset of canonical URLs with honest lastmod dates. Bulk URL support included.
- [Disavow file generator](https://crawlwise.site/tools/disavow): Advanced maintenance for confirmed manual actions. Most sites never need this.
- [Search preview](https://crawlwise.site/tools/serp): See how a title and description truncate in Google results on desktop and mobile. Approximate: Google may rewrite either one.
- [Core Web Vitals checker](https://crawlwise.site/tools/vitals): Lab Lighthouse and field Chrome UX Report data from Google PageSpeed Insights.
- [Content phrase inspector](https://crawlwise.site/tools/keyword-density): See which words and phrases your page actually repeats. There is no target percentage.
- [Domain rating checker](https://crawlwise.site/tools/domain-rating): Look up Domain Rating for a host. An Ahrefs metric, not a Google ranking factor.
- [SEO + generative search checklist](https://crawlwise.site/tools/seo-checklist): Work through what applies to your site. Items you mark not applicable are excluded, not counted against you.
- [Generative search readiness review](https://crawlwise.site/tools/visibility): Inspect the source signals a retriever can see. No score: actual citations are not measured here.
- [llms.txt generator](https://crawlwise.site/tools/llms-txt): Create a Markdown overview for tools that support this format. Experimental: it does not improve Google rankings or guarantee AI citations.
- [AI crawler checker](https://crawlwise.site/tools/ai-crawler-check): See which AI crawlers can read your URL — ChatGPT, Claude, Perplexity and more. Find fixable blocks before you miss citations. Free, no sign-up.
- [Schema / Rich Results validator](https://crawlwise.site/tools/schema-validator): Validate JSON-LD schema. A Rich Results Test alternative: eligible, ineligible, or not applicable. Free, no sign-up.
- [Content question planner](https://crawlwise.site/tools/query-fanout): Draft the questions your page should answer. Editorial prompts, not engine queries.
- [Broken link checker](https://crawlwise.site/tools/broken-link-checker): Check every link on one page and see which ones fail. Internal and outbound, status by status. Free, no sign-up.
- [Redirect chain checker](https://crawlwise.site/tools/redirect-chain): Follow every hop a URL takes and see the chain, the loops, and the final destination. Free, no sign-up.
- [Social and search preview](https://crawlwise.site/tools/meta-preview): One URL, previewed as Google, Facebook, X, LinkedIn, WhatsApp, and Slack would show it. Free, no sign-up.
- [robots.txt tester](https://crawlwise.site/tools/robots-tester): Does this URL get blocked, for this crawler? Test the real file against a real path. Free, no sign-up.
- [Sitemap validator](https://crawlwise.site/tools/sitemap-validator): Fetch a sitemap or sitemap index and check every URL entry: the XML, the limits, and where the URLs actually land. Free, no sign-up.
- [Page speed comparison](https://crawlwise.site/tools/speed-compare): Two URLs, the same Lighthouse run each, side by side. Free, no sign-up.
- [HTTP header inspector](https://crawlwise.site/tools/headers): Every response header for a URL, with the security and caching headers explained. Free, no sign-up.
- [Mixed content checker](https://crawlwise.site/tools/mixed-content): Find HTTP resources loaded by an HTTPS page, which browsers block or warn on. Free, no sign-up.
- [Canonical checker](https://crawlwise.site/tools/canonical-checker): Read the canonical a page declares, and what it points at. Free, no sign-up.
- [Structured data extractor](https://crawlwise.site/tools/structured-data): Every JSON-LD entity on a page, nested and readable, with the types and properties as published. Free, no sign-up.
- [Heading checker](https://crawlwise.site/tools/heading-checker): List every H1–H6 on a page in document order, with the H1 count, empty headings, and level jumps shown as observations, not errors. Free, no sign-up.
- [HTTP status code checker](https://crawlwise.site/tools/http-status-checker): Check the HTTP status code of up to 10 URLs at once: the final code, every redirect hop with its Location, response time, and the headers that explain it. Free, no sign-up.

## Research

- [AI Crawler Access Report](https://crawlwise.site/reports/ai-crawler-access): Which AI crawlers the top 10,000 domains allow or block, from robots.txt and a live fetch as each crawler
- [The Crawlwise SEO Index](https://crawlwise.site/reports/seo-index): SEO statistics from 10,000 top hosts: titles, meta descriptions, canonicals, noindex, JSON-LD, sitemaps and HTTPS, each with its denominator

## Developers and agents

- [Agents and MCP](https://crawlwise.site/agents): Use Crawlwise from coding agents through MCP or the HTTP API
- [MCP server](https://crawlwise.site/docs/mcp): Connect the Crawlwise MCP server and the tools it exposes
- [HTTP API](https://crawlwise.site/docs/api): Authentication, endpoints and credit costs
- [OpenAPI specification](https://crawlwise.site/v1/openapi.json): Machine-readable description of the public API
- [Claude Code setup](https://crawlwise.site/docs/claude-code): Add the Crawlwise MCP server to Claude Code
- [Cursor setup](https://crawlwise.site/docs/cursor): Add the Crawlwise MCP server to Cursor
- [Windsurf setup](https://crawlwise.site/docs/windsurf): Add the Crawlwise MCP server to Windsurf

## Guides

- [Guides](https://crawlwise.site/guides): Step-by-step guides for audits, plans and every tool
- [Blog](https://crawlwise.site/blog): Technical SEO, AI crawler access and performance articles
- [Learn](https://crawlwise.site/learn): Reference pages for AI crawlers, schema types and every check
- [Your first site audit: what to fix this week](https://crawlwise.site/blog/site-audits/first-site-audit-checklist): A practical order of operations for your first Crawlwise audit — from crawler access to metadata, performance, and structured data.
- [Five crawler-blocking mistakes that hide your best pages](https://crawlwise.site/blog/answer-engine-crawlers/answer-engine-crawler-blocking-mistakes): Robots.txt surprises, CDN rules, and staging leaks — common ways sites block retrieval bots without knowing it.
- [How to run your first full Crawlwise audit](https://crawlwise.site/guides/getting-started/run-full-audit): From URL to saved report: credits, Turnstile, reading the checklist, and exporting a plan.
- [Which Crawlwise plan fits your workflow?](https://crawlwise.site/guides/plans-and-credits/pick-a-plan): Starter vs Pro vs Studio vs Agency — credits, API limits, and when yearly billing wins.
- [Site crawls, monitors, and rank tracking](https://crawlwise.site/guides/visibility/site-crawls-and-monitoring): When to add a site, how crawls differ from single-page audits, and what monitoring emails mean.
- [What is an AI crawler? Training, search, and user-triggered fetchers](https://crawlwise.site/guides/ai-search/what-is-an-ai-crawler): An AI crawler is a bot run by an AI company that fetches web pages. The three kinds — training, search, and user-triggered — behave differently and need different decisions.
- [Which AI crawlers should you allow? A vendor-by-vendor decision guide](https://crawlwise.site/guides/ai-search/which-ai-crawlers-to-allow): GPTBot vs OAI-SearchBot, ClaudeBot vs Claude-SearchBot, PerplexityBot, Google-Extended, Applebot-Extended, CCBot and more: what blocking each one costs, and what it does not.
- [Does robots.txt control AI crawlers? What it does and does not do](https://crawlwise.site/guides/ai-search/does-robots-txt-control-ai-crawlers): robots.txt is a request the major AI crawlers say they honor — with exceptions for user-triggered fetchers, and with CDN and firewall rules able to override it in either direction.
- [Does llms.txt affect Google Search? What Google says, and what it is for](https://crawlwise.site/guides/ai-search/does-llms-txt-affect-google-search): Google says Search ignores llms.txt and other AI text files. Here is what llms.txt actually is, who might read it, and what to fix instead.
- [How to check whether AI crawlers can access your site (step by step)](https://crawlwise.site/guides/ai-search/check-ai-crawler-access): A five-layer check — robots.txt, CDN and firewall rules, live user-agent requests, server HTML, and logs — to find out whether GPTBot, ClaudeBot, PerplexityBot and others can read your pages.
- [GPTBot vs OAI-SearchBot vs ChatGPT-User: what each OpenAI bot does](https://crawlwise.site/guides/ai-search/gptbot-vs-oai-searchbot-vs-chatgpt-user): GPTBot trains models, OAI-SearchBot powers ChatGPT search, and ChatGPT-User fetches pages for a person’s request. How to allow or block each one independently.
- [How ChatGPT, Claude, Gemini, Perplexity, Copilot and Grok find web pages](https://crawlwise.site/guides/ai-search/how-ai-assistants-find-web-pages): Which index and which crawlers each AI assistant relies on, as documented by the vendors — and what that means for getting your pages cited.
- [robots.txt for AI crawlers: copy-paste examples with explanations](https://crawlwise.site/guides/ai-search/robots-txt-for-ai-crawlers-examples): Three tested robots.txt policies for AI crawlers — allow all, block training but stay in AI answers, and block everything — plus the mistakes that break them.
- [AI Overviews and AI Mode: what site owners can actually control](https://crawlwise.site/guides/ai-search/ai-overviews-ai-mode-controls): Google says AI Overviews and AI Mode need no special markup. Here are the controls that do apply — nosnippet, max-snippet, data-nosnippet, noindex — and what Google-Extended does not do.
- [Structured data in the AI search era: what still matters for a SaaS or tool site](https://crawlwise.site/guides/ai-search/structured-data-ai-search): Structured data helps search engines understand a page and can make it eligible for rich results, but guarantees neither a rich result nor an AI citation. The types still worth maintaining.
- [robots.txt in the wild: what the top 10,000 domains actually serve](https://crawlwise.site/blog/technical-seo/robots-txt-top-10000-domains): We requested robots.txt from the 10,000 most-queried domains on the Umbrella list. 1,837 served a file a crawler could read; 2,503 returned 404 or 410; 5,660 could not be read at all.
- [noindex vs robots.txt disallow: which one removes a page from Google?](https://crawlwise.site/blog/technical-seo/noindex-vs-robots-txt-disallow): noindex removes a page from Google’s index; robots.txt disallow only stops crawling. Combining them fails because a blocked crawler never sees the noindex. Here is the fix.
- [Canonical tags: why Google treats rel=canonical as a hint, and the mistakes that make it ignore yours](https://crawlwise.site/blog/technical-seo/canonical-tags-how-google-treats-them): rel=canonical is a strong signal, not a directive. How Google picks a canonical, how redirects and sitemaps stack with the tag, and the common mistakes that get it overruled.
- [Redirect chains and loops: 301 vs 302 vs 308, and how many hops is too many](https://crawlwise.site/blog/technical-seo/redirect-chains-301-302-308): 301 and 308 are permanent, 302 and 307 temporary; 308 and 307 keep the request method. Googlebot follows up to 10 hops, but Google advises fewer than 5. How to find and flatten chains.
- [XML sitemaps: lastmod honesty, size limits, and what Google ignores](https://crawlwise.site/blog/technical-seo/xml-sitemaps-lastmod-limits): Google ignores priority and changefreq, uses lastmod only when it is accurate, and caps each sitemap at 50,000 URLs or 50MB. What to put in a sitemap, and what to leave out.
- [HTTP status codes for SEO: 404 vs 410, soft 404s, 5xx, and 429](https://crawlwise.site/blog/technical-seo/http-status-codes-for-seo): How Google treats each status class: 404 and 410 alike, soft 404s on 200 pages, 5xx and 429 slowing crawling, and persistent server errors dropping URLs from the index.

## Tool playbooks

- [Schema markup generator: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/schema-playbook): How to use Schema markup generator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Meta tag generator: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/meta-tags-playbook): How to use Meta tag generator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [robots.txt generator: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/robots-txt-playbook): How to use robots.txt generator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Hreflang tag generator: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/hreflang-playbook): How to use Hreflang tag generator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Redirect generator: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/redirects-playbook): How to use Redirect generator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [XML sitemap generator: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/sitemap-generator-playbook): How to use XML sitemap generator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Disavow file generator: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/disavow-playbook): How to use Disavow file generator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Search preview: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/serp-playbook): How to use Search preview on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Core Web Vitals checker: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/vitals-playbook): How to use Core Web Vitals checker on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Content phrase inspector: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/keyword-density-playbook): How to use Content phrase inspector on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Domain rating checker: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/domain-rating-playbook): How to use Domain rating checker on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [SEO + generative search checklist: a calm step-by-step playbook](https://crawlwise.site/guides/getting-started/seo-checklist-playbook): How to use SEO + generative search checklist on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Generative search readiness review: a calm step-by-step playbook](https://crawlwise.site/guides/visibility/visibility-playbook): How to use Generative search readiness review on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [llms.txt generator: a calm step-by-step playbook](https://crawlwise.site/guides/visibility/llms-txt-playbook): How to use llms.txt generator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [AI crawler checker: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/ai-crawler-check-playbook): How to use AI crawler checker on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Schema / Rich Results validator: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/schema-validator-playbook): How to use Schema / Rich Results validator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Content question planner: a calm step-by-step playbook](https://crawlwise.site/guides/visibility/query-fanout-playbook): How to use Content question planner on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Broken link checker: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/broken-link-checker-playbook): How to use Broken link checker on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Redirect chain checker: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/redirect-chain-playbook): How to use Redirect chain checker on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Social and search preview: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/meta-preview-playbook): How to use Social and search preview on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [robots.txt tester: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/robots-tester-playbook): How to use robots.txt tester on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Sitemap validator: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/sitemap-validator-playbook): How to use Sitemap validator on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Page speed comparison: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/speed-compare-playbook): How to use Page speed comparison on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [HTTP header inspector: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/headers-playbook): How to use HTTP header inspector on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Mixed content checker: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/mixed-content-playbook): How to use Mixed content checker on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Canonical checker: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/canonical-checker-playbook): How to use Canonical checker on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Structured data extractor: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/structured-data-playbook): How to use Structured data extractor on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Heading checker: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/heading-checker-playbook): How to use Heading checker on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [HTTP status code checker: a calm step-by-step playbook](https://crawlwise.site/guides/tool-playbooks/http-status-checker-playbook): How to use HTTP status code checker on a real page, what the output means, and when a full Crawlwise audit is the kinder next step.
- [Using Schema markup generator before you publish](https://crawlwise.site/blog/free-tools/using-schema-before-you-publish): We built Product JSON-LD with the schema generator and ran it through our validator: without a price it was ineligible for product snippets, though the generator flagged nothing missing.
- [Using Meta tag generator before you publish](https://crawlwise.site/blog/free-tools/using-meta-tags-before-you-publish): We fed GitHub’s and Wikipedia’s live title and description into the meta tag generator. Its length hints are character guidelines; Google says titles have no length limit.
- [Using robots.txt generator before you publish](https://crawlwise.site/blog/free-tools/using-robots-txt-before-you-publish): The “allow search, block training” template disallows ten training crawlers and leaves search open. 785 of 1,837 readable robots.txt files in our top-10,000 data disallow GPTBot.
- [Using Hreflang tag generator before you publish](https://crawlwise.site/blog/free-tools/using-hreflang-before-you-publish): The hreflang generator writes link tags, headers or sitemap entries with x-default. It accepts en-UK and es-419, which Google does not support, and cannot check return links.
- [Using Redirect generator before you publish](https://crawlwise.site/blog/free-tools/using-redirects-before-you-publish): We wrote the redirect MDN serves for its moved 404 reference page with the redirect generator, and tested its loop and query-string checks. It writes rules; it cannot test them.
- [Using XML sitemap generator before you publish](https://crawlwise.site/blog/free-tools/using-sitemap-generator-before-you-publish): The XML sitemap generator writes a valid urlset and refuses mixed origins and non-ISO dates. GOV.UK splits its sitemap into 35 files of up to 25,000 URLs each.
- [Using Disavow file generator before you publish](https://crawlwise.site/blog/free-tools/using-disavow-before-you-publish): Google says most sites never need a disavow file. The generator formats one correctly — domain: prefixes, comments, UTF-8 — but cannot see your links or decide what to disavow.
- [Using Search preview before you publish](https://crawlwise.site/blog/performance/using-serp-before-you-publish): We measured real titles in the search preview’s pixel model: GitHub’s 61-character title fits the 580 px desktop budget; its 186-character description is cut near 990 px.
- [Using Core Web Vitals checker before you publish](https://crawlwise.site/blog/performance/using-vitals-before-you-publish): We ran the Core Web Vitals checker on GOV.UK: lab LCP 2.0 s on Google’s test hardware, field LCP 0.70 s from real Chrome users. Here is why both numbers are true.
- [Using Content phrase inspector before you publish](https://crawlwise.site/blog/performance/using-keyword-density-before-you-publish): The phrase inspector counted Wikipedia’s SEO article: “search” 190 times (3.26%), but also “retrieved” 59 and “archived” 52 — citation text. Google publishes no ideal density.
- [Using Domain rating checker before you publish](https://crawlwise.site/blog/performance/using-domain-rating-before-you-publish): Domain Rating is Ahrefs’ score for a backlink profile “compared to the others in our database”. What the checker looks up, what it refuses to invent, and what DR cannot tell you.
- [Using SEO + generative search checklist before you publish](https://crawlwise.site/blog/free-tools/using-seo-checklist-before-you-publish): The technical SEO checklist has 45 items in six sections; 13 carry a condition for when they apply. Its Core Web Vitals and heading items match Google’s own documentation.
- [Using Generative search readiness review before you publish](https://crawlwise.site/blog/answer-engine-crawlers/using-visibility-before-you-publish): The generative search readiness review read Wikipedia’s SEO article and GOV.UK’s homepage. It reports 12 source signals with no score, because Google says AI features need nothing special.
- [Using llms.txt generator before you publish](https://crawlwise.site/blog/answer-engine-crawlers/using-llms-txt-before-you-publish): We generated an llms.txt for MDN: five entries, all HTML reference pages, because it takes the first four internal links. Google says Search ignores the file.
- [Using AI crawler checker before you publish](https://crawlwise.site/blog/performance/using-ai-crawler-check-before-you-publish): We fetched GOV.UK, GitHub and The New York Times as nine AI crawlers. Two let every bot in; the Times disallows all nine in robots.txt and answered every request 403.
- [Using Schema / Rich Results validator before you publish](https://crawlwise.site/blog/performance/using-schema-validator-before-you-publish): We validated live JSON-LD from Wikipedia and BBC Good Food: Article, Recipe, VideoObject and BreadcrumbList all eligible. A generated Product without offers was ineligible.
- [Using Content question planner before you publish](https://crawlwise.site/blog/answer-engine-crawlers/using-query-fanout-before-you-publish): The content question planner turns a topic into 9 to 15 editorial prompts from fixed templates. Google’s AI features run their own query fan-out, which no outside tool can see.
- [Using Broken link checker before you publish](https://crawlwise.site/blog/performance/using-broken-link-checker-before-you-publish): The broken link checker found 46 links on GOV.UK’s vehicle tax page and checked the first 25: all answered 2xx. Why it checks 25, and why a 403 is not reported as broken.
- [Using Redirect chain checker before you publish](https://crawlwise.site/blog/performance/using-redirect-chain-before-you-publish): The redirect checker followed http://gov.uk/ through three hops and an old MDN URL through four, including a 302 in the middle. What Google says about chains.
- [Using Social and search preview before you publish](https://crawlwise.site/blog/performance/using-meta-preview-before-you-publish): The social preview read GitHub, GOV.UK and Wikipedia. GitHub declares every Open Graph tag; GOV.UK only og:image; Wikipedia’s article declared no og:image at all.
- [Using robots.txt tester before you publish](https://crawlwise.site/blog/performance/using-robots-tester-before-you-publish): We tested real paths against live robots.txt files: Wikipedia disallows /wiki/Special:, GitHub disallows /search for Googlebot, GOV.UK disallows /search/all. With the rules quoted.
- [Using Sitemap validator before you publish](https://crawlwise.site/blog/performance/using-sitemap-validator-before-you-publish): The sitemap validator read GOV.UK’s and MDN’s sitemap indexes cleanly. GOV.UK’s first child sitemap, 5.5 MB with 25,000 URLs, was over the tool’s 4 MB read limit.
- [Using Page speed comparison before you publish](https://crawlwise.site/blog/performance/using-speed-compare-before-you-publish): We compared GOV.UK and MDN with the page speed comparison: Lighthouse scores 97 and 91, lab LCP 2.0 s and 3.2 s. Field data said both were fast for real users.
- [Using HTTP header inspector before you publish](https://crawlwise.site/blog/performance/using-headers-before-you-publish): The HTTP header inspector read Wikipedia, GOV.UK and MDN: GOV.UK sent all six security headers it looks for, MDN five, Wikipedia three. Caching ranged from max-age=0 to 3600.
- [Using Mixed content checker before you publish](https://crawlwise.site/blog/performance/using-mixed-content-before-you-publish): The mixed content checker scanned Wikipedia, GOV.UK and MDN and found no http:// resources in their source HTML. What browsers block, what they upgrade, and what a scan misses.
- [Using Canonical checker before you publish](https://crawlwise.site/blog/performance/using-canonical-checker-before-you-publish): The canonical checker on three real URLs: a Wikipedia URL with utm_source points to the clean URL, an MDN page is self-referencing, and a GOV.UK browse page declares none.
- [Using Structured data extractor before you publish](https://crawlwise.site/blog/performance/using-structured-data-before-you-publish): The structured data extractor found one Article on Wikipedia, a Recipe with 56 properties on BBC Good Food, and no JSON-LD on GitHub’s or GOV.UK’s homepages.
- [Using Heading checker before you publish](https://crawlwise.site/blog/performance/using-heading-checker-before-you-publish): The heading checker on Wikipedia, GOV.UK and MDN: one H1 each, no skipped levels, and on two of them an H2 before the H1. Google says heading order doesn’t matter for Search.
- [Using HTTP status code checker before you publish](https://crawlwise.site/blog/performance/using-http-status-checker-before-you-publish): Six URLs through the HTTP status code checker: two real 404s, a three-hop redirect to wikipedia.org and a moved MDN page. What each code means to Google.

## AI crawler reference

- [GPTBot](https://crawlwise.site/learn/crawler/gptbot): OpenAI (ChatGPT): Foundation-model training
- [OAI-SearchBot](https://crawlwise.site/learn/crawler/oai-searchbot): OpenAI (ChatGPT search): Search product index
- [ChatGPT-User](https://crawlwise.site/learn/crawler/chatgpt-user): OpenAI (ChatGPT): On-demand browse for a user
- [ClaudeBot](https://crawlwise.site/learn/crawler/claudebot): Anthropic (Claude): Model training crawler
- [Claude-SearchBot](https://crawlwise.site/learn/crawler/claude-searchbot): Anthropic (Claude search): Search discovery crawler
- [Claude-User](https://crawlwise.site/learn/crawler/claude-user): Anthropic (Claude): User-triggered fetch
- [Google-Extended](https://crawlwise.site/learn/crawler/google-extended): Google (Gemini): Gemini / Vertex training opt-in. robots.txt token; the live fetch uses Googlebot
- [PerplexityBot](https://crawlwise.site/learn/crawler/perplexitybot): Perplexity (Perplexity): Answer-engine index
- [Perplexity-User](https://crawlwise.site/learn/crawler/perplexity-user): Perplexity (Perplexity): On-demand citation fetch
- [CCBot](https://crawlwise.site/learn/crawler/ccbot): Common Crawl (Common Crawl): Open web corpus
- [Bytespider](https://crawlwise.site/learn/crawler/bytespider): ByteDance (ByteDance): Training corpus
- [Applebot-Extended](https://crawlwise.site/learn/crawler/applebot-extended): Apple (Apple Intelligence): Apple Intelligence training opt-in. robots.txt token; the live fetch uses Applebot
- [Meta-ExternalAgent](https://crawlwise.site/learn/crawler/meta-externalagent): Meta (Meta AI): Llama-product crawler
- [Amazonbot](https://crawlwise.site/learn/crawler/amazonbot): Amazon (Alexa): Alexa and product training
- [cohere-ai](https://crawlwise.site/learn/crawler/cohere-ai): Cohere (Cohere): Search and training fetch
- [Diffbot](https://crawlwise.site/learn/crawler/diffbot): Diffbot (Diffbot): Knowledge graph
- [Timpibot](https://crawlwise.site/learn/crawler/timpibot): Timpi (Timpi): Search index

## Comparisons

- [Comparisons](https://crawlwise.site/vs): How Crawlwise compares with other audit tools, including what each does better
- [Crawlwise vs Screaming Frog SEO Spider](https://crawlwise.site/vs/screaming-frog): The desktop crawler most technical SEOs already have open.
- [Crawlwise vs Ahrefs Site Audit](https://crawlwise.site/vs/ahrefs-site-audit): An audit bolted to the largest independent backlink index.
- [Crawlwise vs Semrush Site Audit](https://crawlwise.site/vs/semrush-site-audit): An enterprise SEO suite with an audit attached to a very large keyword database.
- [Crawlwise vs open-seo](https://crawlwise.site/vs/open-seo): An open-source DataForSEO interface you can self-host.
- [Crawlwise vs Sitebulb](https://crawlwise.site/vs/sitebulb): A desktop crawler built around prioritised hints and visual reports.
- [Crawlwise vs Seobility](https://crawlwise.site/vs/seobility): A cloud SEO suite with a site audit, rank tracking and backlinks in one subscription.
- [Crawlwise vs SEO Site Checkup](https://crawlwise.site/vs/seo-site-checkup): A long-running instant checkup that has grown into a monitoring suite.
- [Free alternatives](https://crawlwise.site/alternatives): Where Crawlwise is a free alternative to another tool, and where it is not
- [A free alternative to Screaming Frog for auditing one page](https://crawlwise.site/alternatives/screaming-frog-single-page): The first thing people open Screaming Frog to do: point it at a URL and read what is wrong with that page.
- [A free alternative to Ahrefs Site Audit for a single URL](https://crawlwise.site/alternatives/ahrefs-site-audit-single-page): Auditing one URL without a subscription, and seeing the finding list rather than a health score.
- [A free alternative to the Rich Results Test, that says why](https://crawlwise.site/alternatives/rich-results-test): Testing JSON-LD against the feature requirements, with the failing property named and the rule quoted.
