crawlwise.
Original data · SEO statistics

SEO statistics 2026: The Crawlwise SEO Index

Of 1,547 homepages that served HTML (from 10,000 hosts), 58.2% have a meta description and 34.3% carry JSON-LD structured data.

Sample of 10,000 hosts, collected 2026-10-05.

Headline findings

Of 10,000 hosts requested, 6,083 answered with an HTTP response and 1,577 served an HTML page. 30 of those could not be parsed; they count as unmeasured, not as missing anything. On-page shares below are over the 1,547 HTML homepages that were parsed.

FindingHostsOfShare
Final URL is HTTPS
over hosts that served HTML
1,5071,57795.6%
Has a <title>1,3121,54784.8%
Has a meta description9011,54758.2%
Declares a canonical link7691,54749.7%
Canonical points to the page itself6471,54741.8%
Marked noindex (meta tag or X-Robots-Tag)831,5475.4%
Carries JSON-LD structured data5311,54734.3%
JSON-LD with at least one parse error
over homepages that carry JSON-LD
65311.1%
Mobile viewport declared1,1151,54772.1%
html lang declared1,0771,54769.6%
hreflang alternates declared3301,54721.3%
robots.txt readable
over hosts that served HTML and whose robots.txt was requested
9931,57463.1%
/sitemap.xml is a sitemap
over hosts that served HTML and whose /sitemap.xml answered
6311,52141.5%

Who answered

Every host requested, including the ones that never served a page. Dropping failures is how a sample silently becomes favourable, so they are counted here and excluded from on-page figures by name.

OutcomeHostsShare of sample
Served an HTML page1,57715.8%
Answered with something other than HTML5005%
Answered with an HTTP error3,58635.9%
Answered with a bot challenge4044%
Redirected too many times, or without a target160.2%
Could not be reached2,88128.8%
Timed out2892.9%
Asked us not to fetch (robots.txt)7477.5%
Why a homepage was not measured
ReasonHosts
dns: no such host2,139
http 4042,064
robots.txt disallows our user-agent747
http 403656
tls error628
http 400434
bot challenge page404
timeout289
non-html response194
non-html response: text/plain109
http 500107
http 40192
http 50369
non-html response: application/json64
connection reset57

Titles and meta descriptions

Lengths in characters, over homepages that have one. Length is guidance, not a Google limit, so no length is counted as a defect here.

MeasureHosts25th pctMedian75th pct90th pct
Title length1,31215355564
Meta description length901121146160190

Canonical links

CanonicalHomepages
Points to the page itself647 (41.8%)
Another URL on the same host78 (5%)
The www / non-www twin of the host17 (1.1%)
A different host25 (1.6%)
Empty or not a valid http(s) URL2 (0.1%)
No canonical link778 (50.3%)

11 homepage(s) declare more than one canonical; the first is classified. A homepage canonical pointing elsewhere is often a language or regional redirect target, not a mistake.

Indexing directives

FindingHostsOfShare
noindex or none in robots/googlebot meta, or in X-Robots-Tag
over parsed HTML homepages
831,5475.4%
X-Robots-Tag noindex on the response
over every host that answered, HTML or not
536,0670.9%

Structured data

531 of 1,547 (34.3%) parsed homepages carry at least one JSON-LD block. The most common types, counted once per homepage:

@typeHomepages
Organization420 (27.1%)
WebSite335 (21.7%)
ImageObject322 (20.8%)
WebPage274 (17.7%)
SearchAction213 (13.8%)
ListItem163 (10.5%)
BreadcrumbList152 (9.8%)
PostalAddress149 (9.6%)
ContactPoint146 (9.4%)
EntryPoint141 (9.1%)
ReadAction111 (7.2%)
PropertyValueSpecification98 (6.3%)
Person68 (4.4%)
SoftwareApplication61 (3.9%)
Offer59 (3.8%)

Presence and JSON syntax only. Whether the markup describes visible content or qualifies for a rich result is not assessed here.

Observations: headings, language, Open Graph

Reported because they are real and widely asked about. None of them is scored by Crawlwise as a search defect: Google sets no H1 count, the lang attribute is an accessibility matter, and Open Graph is a social-sharing format.

ObservationHomepages
No H1664 (42.9%)
Exactly one H1744 (48.1%)
More than one H1139 (9%)
html lang declared and well-formed1,071 (69.2%)
Any og: meta tag802 (51.8%)

Most declared languages: en (1,026), ja (7), zh (7), es (6), tr (4), fr (3), hi (3), pl (3), ru (3), uk (2).

robots.txt and sitemaps, all hosts

FileStateHosts
robots.txtReadable robots document2,068 (30.3%)
robots.txtMissing (404 or 410), which allows everything2,668 (39.1%)
robots.txtUnread: HTML, a challenge, or another error2,088 (30.6%)
robots.txtUnmeasured: no response at all3,176
robots.txtDeclares at least one Sitemap: line (of readable files)827 (40%)
/sitemap.xmlA sitemap (urlset or sitemapindex)661 (11%)
/sitemap.xmlReturned HTML instead263 (4.4%)
/sitemap.xmlReturned something else191 (3.2%)
/sitemap.xmlMissing (404 or 410)3,098 (51.6%)
/sitemap.xmlAnother HTTP error1,792 (29.8%)
/sitemap.xmlUnmeasured (not requested or no response)3,995

A missing /sitemap.xml is not a missing sitemap: many sites publish theirs at another path and declare it in robots.txt, which is why the Sitemap: line is counted separately.

HTTPS, redirects, weight and response time

FindingHostsOfShare
Final URL is HTTPS
over every host that answered
5,2416,06786.4%
No redirect before the final response5,0126,08382.4%
One redirect7116,08311.7%
Two redirects2516,0834.1%
Three or more redirects1096,0831.8%
MeasureHosts25th pctMedian75th pct90th pct
HTML size (decoded)1,5755 KB85 KB291 KB584 KB
Response time to headers1,57767 ms191 ms333 ms650 ms

HTML only, excluding images, scripts and styles. Response time is one request from one location, timed to the response headers of the final hop; it is a server-side sample, not a visitor’s experience.

By rank band

Shares over the HTML homepages parsed in each band, except HTTPS (over hosts that served HTML) and sitemap (over hosts whose /sitemap.xml answered).

RanksHostsHTML parsedHTTPSTitleDescriptionCanonicalJSON-LDViewportnoindex/sitemap.xml
1–1001003497.1%94.1%55.9%52.9%32.4%73.5%11.8%32.3%
101–1,00090017496%85.6%60.3%51.7%37.4%74.7%5.7%36.1%
1,001–5,0004,00065295.2%84.7%58.7%51.4%33.9%71.9%4.9%45.9%
5,001–10,0005,00068795.7%84.3%57.4%47.5%34.1%71.5%5.4%39.1%

What changed

No earlier run of a comparable sample exists to compare against, so nothing is claimed here.

Methodology, in full

Sample. This run requested the homepage of 10,000 hosts, ranks 1 through 10,000 of the Cisco Umbrella top-one-million list (the same list as the AI Crawler Access Report). Umbrella ranks hostnames by DNS query volume, so it includes API, CDN, telemetry and ad hosts that never serve a web page. That is why every on-page figure is stated over the hosts that actually served an HTML page, and why that denominator is printed next to each one. It is a popularity-ranked sample, not a random sample of the web, so every share describes this list and not the internet.

Requests. Per host: robots.txt first, then the homepage with redirects followed (at most five), then /sitemap.xml on the final origin, plus the final origin’s robots.txt when the homepage redirected to another host. One request at a time per host, four hosts in parallel, a pause after each host. The user-agent names Crawlwise and links here (Mozilla/5.0 (compatible; CrawlwiseResearch/1.0; +https://crawlwise.site/reports/seo-index)). A host whose robots.txt disallows that user-agent is not fetched and is reported as such. A host that refuses https connections is retried once over http.

Definitions. Each measurement uses the definition the Crawlwise audit uses, documented on the methodology page (methodology version 6, index version 1): the title is the text of the head title element; the description is the first meta description; the canonical is the first link rel=canonical in the head, resolved against the final URL; noindex is noindex or none in a robots or googlebot meta tag or the X-Robots-Tag header; the viewport counts when it sets width=device-width; JSON-LD is every script of type application/ld+json, parsed with JSON.parse and walked for @type; robots.txt is readable when it is a robots document and not an HTML or challenge page, and missing on 404 or 410, by the same rules as the AI crawler checker; a sitemap is a response whose root element is urlset or sitemapindex.

Unmeasured is an answer. A host that timed out, refused the connection, answered with an error or a bot challenge, or returned something other than HTML is recorded with the reason and kept out of on-page denominators. HTML larger than the audit’s parse limits (1.5 MB or 50,000 tags) is counted as served but unmeasured. Nothing is inferred for a host we could not read.

What this cannot tell you

Check one site yourself, free

Related research: the AI Crawler Access Report, which hosts on the same list allow or block GPTBot, ClaudeBot and other AI crawlers.

Reproduce it

node scripts/fetch-crawler-report-tranco.mjs --limit 10000
node scripts/seo-index.mjs --list scripts/fixtures/crawler-report-tranco.csv --limit 10000 --label top-10000 --concurrency 4 --delay 350

About four to five requests per host. The page states whichever sample size it is reading, and nothing here is extrapolated from a smaller sample.

Audit one page, free