SEO statistics 2026: The Crawlwise SEO Index
Of 1,547 homepages that served HTML (from 10,000 hosts), 58.2% have a meta description and 34.3% carry JSON-LD structured data.
Sample of 10,000 hosts, collected 2026-10-05.
Headline findings
Of 10,000 hosts requested, 6,083 answered with an HTTP response and 1,577 served an HTML page. 30 of those could not be parsed; they count as unmeasured, not as missing anything. On-page shares below are over the 1,547 HTML homepages that were parsed.
| Finding | Hosts | Of | Share |
|---|---|---|---|
| Final URL is HTTPS over hosts that served HTML | 1,507 | 1,577 | 95.6% |
| Has a <title> | 1,312 | 1,547 | 84.8% |
| Has a meta description | 901 | 1,547 | 58.2% |
| Declares a canonical link | 769 | 1,547 | 49.7% |
| Canonical points to the page itself | 647 | 1,547 | 41.8% |
| Marked noindex (meta tag or X-Robots-Tag) | 83 | 1,547 | 5.4% |
| Carries JSON-LD structured data | 531 | 1,547 | 34.3% |
| JSON-LD with at least one parse error over homepages that carry JSON-LD | 6 | 531 | 1.1% |
| Mobile viewport declared | 1,115 | 1,547 | 72.1% |
| html lang declared | 1,077 | 1,547 | 69.6% |
| hreflang alternates declared | 330 | 1,547 | 21.3% |
| robots.txt readable over hosts that served HTML and whose robots.txt was requested | 993 | 1,574 | 63.1% |
| /sitemap.xml is a sitemap over hosts that served HTML and whose /sitemap.xml answered | 631 | 1,521 | 41.5% |
Who answered
Every host requested, including the ones that never served a page. Dropping failures is how a sample silently becomes favourable, so they are counted here and excluded from on-page figures by name.
| Outcome | Hosts | Share of sample |
|---|---|---|
| Served an HTML page | 1,577 | 15.8% |
| Answered with something other than HTML | 500 | 5% |
| Answered with an HTTP error | 3,586 | 35.9% |
| Answered with a bot challenge | 404 | 4% |
| Redirected too many times, or without a target | 16 | 0.2% |
| Could not be reached | 2,881 | 28.8% |
| Timed out | 289 | 2.9% |
| Asked us not to fetch (robots.txt) | 747 | 7.5% |
Why a homepage was not measured
| Reason | Hosts |
|---|---|
| dns: no such host | 2,139 |
| http 404 | 2,064 |
| robots.txt disallows our user-agent | 747 |
| http 403 | 656 |
| tls error | 628 |
| http 400 | 434 |
| bot challenge page | 404 |
| timeout | 289 |
| non-html response | 194 |
| non-html response: text/plain | 109 |
| http 500 | 107 |
| http 401 | 92 |
| http 503 | 69 |
| non-html response: application/json | 64 |
| connection reset | 57 |
Titles and meta descriptions
Lengths in characters, over homepages that have one. Length is guidance, not a Google limit, so no length is counted as a defect here.
| Measure | Hosts | 25th pct | Median | 75th pct | 90th pct |
|---|---|---|---|---|---|
| Title length | 1,312 | 15 | 35 | 55 | 64 |
| Meta description length | 901 | 121 | 146 | 160 | 190 |
Canonical links
| Canonical | Homepages |
|---|---|
| Points to the page itself | 647 (41.8%) |
| Another URL on the same host | 78 (5%) |
| The www / non-www twin of the host | 17 (1.1%) |
| A different host | 25 (1.6%) |
| Empty or not a valid http(s) URL | 2 (0.1%) |
| No canonical link | 778 (50.3%) |
11 homepage(s) declare more than one canonical; the first is classified. A homepage canonical pointing elsewhere is often a language or regional redirect target, not a mistake.
Indexing directives
| Finding | Hosts | Of | Share |
|---|---|---|---|
| noindex or none in robots/googlebot meta, or in X-Robots-Tag over parsed HTML homepages | 83 | 1,547 | 5.4% |
| X-Robots-Tag noindex on the response over every host that answered, HTML or not | 53 | 6,067 | 0.9% |
Structured data
531 of 1,547 (34.3%) parsed homepages carry at least one JSON-LD block. The most common types, counted once per homepage:
| @type | Homepages |
|---|---|
| Organization | 420 (27.1%) |
| WebSite | 335 (21.7%) |
| ImageObject | 322 (20.8%) |
| WebPage | 274 (17.7%) |
| SearchAction | 213 (13.8%) |
| ListItem | 163 (10.5%) |
| BreadcrumbList | 152 (9.8%) |
| PostalAddress | 149 (9.6%) |
| ContactPoint | 146 (9.4%) |
| EntryPoint | 141 (9.1%) |
| ReadAction | 111 (7.2%) |
| PropertyValueSpecification | 98 (6.3%) |
| Person | 68 (4.4%) |
| SoftwareApplication | 61 (3.9%) |
| Offer | 59 (3.8%) |
Presence and JSON syntax only. Whether the markup describes visible content or qualifies for a rich result is not assessed here.
Observations: headings, language, Open Graph
Reported because they are real and widely asked about. None of them is scored by Crawlwise as a search defect: Google sets no H1 count, the lang attribute is an accessibility matter, and Open Graph is a social-sharing format.
| Observation | Homepages |
|---|---|
| No H1 | 664 (42.9%) |
| Exactly one H1 | 744 (48.1%) |
| More than one H1 | 139 (9%) |
| html lang declared and well-formed | 1,071 (69.2%) |
| Any og: meta tag | 802 (51.8%) |
Most declared languages: en (1,026), ja (7), zh (7), es (6), tr (4), fr (3), hi (3), pl (3), ru (3), uk (2).
robots.txt and sitemaps, all hosts
| File | State | Hosts |
|---|---|---|
| robots.txt | Readable robots document | 2,068 (30.3%) |
| robots.txt | Missing (404 or 410), which allows everything | 2,668 (39.1%) |
| robots.txt | Unread: HTML, a challenge, or another error | 2,088 (30.6%) |
| robots.txt | Unmeasured: no response at all | 3,176 |
| robots.txt | Declares at least one Sitemap: line (of readable files) | 827 (40%) |
| /sitemap.xml | A sitemap (urlset or sitemapindex) | 661 (11%) |
| /sitemap.xml | Returned HTML instead | 263 (4.4%) |
| /sitemap.xml | Returned something else | 191 (3.2%) |
| /sitemap.xml | Missing (404 or 410) | 3,098 (51.6%) |
| /sitemap.xml | Another HTTP error | 1,792 (29.8%) |
| /sitemap.xml | Unmeasured (not requested or no response) | 3,995 |
A missing /sitemap.xml is not a missing sitemap: many sites publish theirs at another path and declare it in robots.txt, which is why the Sitemap: line is counted separately.
HTTPS, redirects, weight and response time
| Finding | Hosts | Of | Share |
|---|---|---|---|
| Final URL is HTTPS over every host that answered | 5,241 | 6,067 | 86.4% |
| No redirect before the final response | 5,012 | 6,083 | 82.4% |
| One redirect | 711 | 6,083 | 11.7% |
| Two redirects | 251 | 6,083 | 4.1% |
| Three or more redirects | 109 | 6,083 | 1.8% |
| Measure | Hosts | 25th pct | Median | 75th pct | 90th pct |
|---|---|---|---|---|---|
| HTML size (decoded) | 1,575 | 5 KB | 85 KB | 291 KB | 584 KB |
| Response time to headers | 1,577 | 67 ms | 191 ms | 333 ms | 650 ms |
HTML only, excluding images, scripts and styles. Response time is one request from one location, timed to the response headers of the final hop; it is a server-side sample, not a visitor’s experience.
By rank band
Shares over the HTML homepages parsed in each band, except HTTPS (over hosts that served HTML) and sitemap (over hosts whose /sitemap.xml answered).
| Ranks | Hosts | HTML parsed | HTTPS | Title | Description | Canonical | JSON-LD | Viewport | noindex | /sitemap.xml |
|---|---|---|---|---|---|---|---|---|---|---|
| 1–100 | 100 | 34 | 97.1% | 94.1% | 55.9% | 52.9% | 32.4% | 73.5% | 11.8% | 32.3% |
| 101–1,000 | 900 | 174 | 96% | 85.6% | 60.3% | 51.7% | 37.4% | 74.7% | 5.7% | 36.1% |
| 1,001–5,000 | 4,000 | 652 | 95.2% | 84.7% | 58.7% | 51.4% | 33.9% | 71.9% | 4.9% | 45.9% |
| 5,001–10,000 | 5,000 | 687 | 95.7% | 84.3% | 57.4% | 47.5% | 34.1% | 71.5% | 5.4% | 39.1% |
What changed
No earlier run of a comparable sample exists to compare against, so nothing is claimed here.
Methodology, in full
Sample. This run requested the homepage of 10,000 hosts, ranks 1 through 10,000 of the Cisco Umbrella top-one-million list (the same list as the AI Crawler Access Report). Umbrella ranks hostnames by DNS query volume, so it includes API, CDN, telemetry and ad hosts that never serve a web page. That is why every on-page figure is stated over the hosts that actually served an HTML page, and why that denominator is printed next to each one. It is a popularity-ranked sample, not a random sample of the web, so every share describes this list and not the internet.
Requests. Per host: robots.txt first, then the homepage with redirects followed (at most five), then /sitemap.xml on the final origin, plus the final origin’s robots.txt when the homepage redirected to another host. One request at a time per host, four hosts in parallel, a pause after each host. The user-agent names Crawlwise and links here (Mozilla/5.0 (compatible; CrawlwiseResearch/1.0; +https://crawlwise.site/reports/seo-index)). A host whose robots.txt disallows that user-agent is not fetched and is reported as such. A host that refuses https connections is retried once over http.
Definitions. Each measurement uses the definition the Crawlwise audit uses, documented on the methodology page (methodology version 6, index version 1): the title is the text of the head title element; the description is the first meta description; the canonical is the first link rel=canonical in the head, resolved against the final URL; noindex is noindex or none in a robots or googlebot meta tag or the X-Robots-Tag header; the viewport counts when it sets width=device-width; JSON-LD is every script of type application/ld+json, parsed with JSON.parse and walked for @type; robots.txt is readable when it is a robots document and not an HTML or challenge page, and missing on 404 or 410, by the same rules as the AI crawler checker; a sitemap is a response whose root element is urlset or sitemapindex.
Unmeasured is an answer. A host that timed out, refused the connection, answered with an error or a bot challenge, or returned something other than HTML is recorded with the reason and kept out of on-page denominators. HTML larger than the audit’s parse limits (1.5 MB or 50,000 tags) is counted as served but unmeasured. Nothing is inferred for a host we could not read.
What this cannot tell you
- Homepages only. A site’s other pages can differ completely.
- Source HTML only. Tags injected by JavaScript after load are not seen, so a JavaScript-rendered site can look emptier here than in a browser.
- One request from one location at one moment. Geo-targeted redirects, A/B tests and CDNs answering differently elsewhere are invisible.
- Presence, not quality. A title can exist and be useless; JSON-LD can parse and still describe nothing on the page.
- Nothing here is a ranking factor study. Whether these signals change rankings is not measured.
Check one site yourself, free
- Meta tag generator: title, description, robots and canonical tags
- Canonical checker: where a page’s canonical points
- Schema validator: JSON-LD syntax and types
- Sitemap validator: whether a sitemap parses
- robots.txt tester: what a robots.txt allows
- HTTP status checker: status codes and redirect hops
- Heading checker: the H1–H6 outline
Related research: the AI Crawler Access Report, which hosts on the same list allow or block GPTBot, ClaudeBot and other AI crawlers.
Reproduce it
node scripts/fetch-crawler-report-tranco.mjs --limit 10000 node scripts/seo-index.mjs --list scripts/fixtures/crawler-report-tranco.csv --limit 10000 --label top-10000 --concurrency 4 --delay 350
About four to five requests per host. The page states whichever sample size it is reading, and nothing here is extrapolated from a smaller sample.