
WordPress SEO audit: robots.txt, AI crawlers and plugin conflicts
WordPress serves a virtual robots.txt unless a physical file exists in the site root. SEO plugins edit that output and security plugins or CDNs can block crawlers without touching it. Read the live /robots.txt, fetch a page as each AI crawler, then audit titles, canonicals and structured data on each template.
Key takeaways
- The live /robots.txt is the truth; it may come from a file, from core or from a plugin.
- Since WordPress 5.3, discouraging search engines uses a noindex meta tag, not robots.txt.
- Security plugins and CDNs block crawlers without leaving a trace in robots.txt.
- Audit each template: posts, pages, categories and any product pages.
On a WordPress site three different things can write the rules a crawler obeys: WordPress core, your SEO plugin, and whatever sits in front of the site, from a security plugin to a CDN. An audit has to check what comes out at the end, not what one settings screen claims.
Where does WordPress’s robots.txt come from?
When nothing else answers, WordPress builds a virtual file with do_robots(): a User-agent: * group that disallows the admin path and allows admin-ajax.php. If a physical robots.txt exists in the site root, the web server serves it first and none of that runs. SEO plugins usually offer a robots.txt editor that changes the virtual output.
How do you add an AI crawler rule in code?
Use the robots_txt filter, which receives the output and whether the site is public:
add_filter( 'robots_txt', function ( $output, $public ) {
if ( $public ) {
$output .= "\nUser-agent: GPTBot\nDisallow: /\n";
}
return $output;
}, 10, 2 );
Put it in a small site plugin rather than a theme, so a theme switch does not silently remove it. The trade-offs between training and retrieval crawlers are in which AI crawlers to allow, and ready-made rule sets are in robots.txt examples for AI crawlers.
On WordPress the settings screen is a claim. The response a crawler gets is the evidence.
Adnan Arodiya, Crawlwise
What can block crawlers without touching robots.txt?
Security and firewall plugins, bot-protection features and CDNs can refuse a user agent with a 403 or serve it a challenge page. Since WordPress 5.3, the “Discourage search engines from indexing this site” setting also works through a noindex robots meta tag rather than a Disallow line, so robots.txt can look open while every page asks not to be indexed. A staging copy that was pushed live with that box still ticked is one of the most common causes of a site vanishing from search.
Which templates should the audit cover?
Pick one URL from each template: the homepage, a post, a page, a category archive, and a product if you run WooCommerce. Check that each has its own title and meta description, a canonical pointing to itself, valid structured data from your SEO plugin, and the main content in the server HTML. Then check the sitemap: core WordPress publishes /wp-sitemap.xml, and SEO plugins usually replace it with their own index, so make sure robots.txt points to the one that actually exists.
Check it on your own WordPress site
Run the free AI crawler checker on your homepage and one product or article page. It reads your robots.txt the way each crawler does, then fetches the page as GPTBot, ClaudeBot, PerplexityBot and 14 others, so a block added by a CDN, firewall or app shows up even when the file looks right. The robots.txt tester shows which rule decides a given path, and a full audit adds titles, canonicals, structured data, hreflang and performance for the same URL.
To see how popular sites set these rules, look any of them up in the site-by-site AI crawler results.
What a plan actually costs
The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.
| Plan | Monthly | Credits | A good fit when |
|---|---|---|---|
| Starter | $4.99 | 10 | You look after one site and check it now and then |
| Pro | $14.99 | 40 | A few sites, plus the API and a handful of watched URLs |
| Studio | $49.99 | 150 | Client work that would burn through Pro mid-month |
| Agency | $99.99 | 400 | Many locations, reported under your own name |
A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.
Frequently asked questions
Where is the WordPress robots.txt file?
Often nowhere on disk. WordPress answers /robots.txt with a virtual file built by do_robots(). If a real robots.txt file sits in the site root, the web server serves that file instead and WordPress’s version never runs.
Does “Discourage search engines” add Disallow: / to robots.txt?
Not since WordPress 5.3. The setting now adds a noindex robots meta tag to pages instead of a Disallow line in robots.txt.
Can a security plugin block AI crawlers?
Yes, and it will not show up in robots.txt. Firewall and bot-protection plugins, and CDNs in front of WordPress, can refuse crawler user agents outright. Only a live fetch as the crawler shows that.
Sources
Platform documentation was read on 2026-10-08. Platforms change these pages and settings without notice, so check the original before you act on a detail.
Crawlwise vs checking inside WordPress
Shows the robots.txt your visitors and crawlers receive
- Crawlwise
- Yes, fetched live
- WordPress admin
- Depends on which plugin wrote it
Fetches the page as each AI crawler
- Crawlwise
- Yes, 17 crawlers
- WordPress admin
- No
Catches CDN and firewall blocks
- Crawlwise
- Yes, with challenge fingerprints
- WordPress admin
- No
Edits your WordPress settings for you
- Crawlwise
- No — it tells you what to change
- WordPress admin
- Yes, it is where you make the change
Try it on your URL
Run the checks on your live site
Free tools answer one question. A full audit scores the page, tests answer-engine crawlers, and saves to history on a plan.
- Live crawler probes
- Core Web Vitals
- One free full audit
Meet the author

Adnan Arodiya
Crawlwise
Writes about what Crawlwise actually measures: on-page evidence, crawler access, and performance signals — with the limits stated up front.


