Original research

The State of Structural SEO

Every RankForge audit contributes one anonymised row to this dataset — no URL, no domain, no page titles, just the structural measurements. This is what those rows say so far.

Free to quote and republish with attribution and a link to this page. Download the dataset (CSV).

What’s broken, and how often

24.1%

of sites have orphan pages

At least one indexable page with no internal links pointing to it — invisible to crawlers that arrive by following links, and invisible in most tools, because a page with no inbound links isn't an error anywhere.

2.6%

of a typical site's pages are orphaned

Averaged per site rather than across all pages, so a handful of very large crawls can't dominate the figure.

n = 116#

have a flat architecture

Nine in ten pages sit one click or less from the homepage. This verdict needs a complete crawl of 50+ pages to be decidable at all, so it's measured on that subset — but the underlying depth measurement below uses every audit.

broken pages in a typical crawl

Internal links resolving to an error status. Every one of them is authority spent on nothing.

AI and agent readiness

of sites publish an llms.txt

The proposed convention for telling AI agents which pages matter. Measured only on audits where the check actually ran, so an un-probed crawl never counts as an absence.

average GEO readiness score

Agent-discoverability signals combined: llms.txt, identity schema, and agent-readable link headers. Low scores here are the norm, not an outlier.

of sites have an XML sitemap

The oldest discovery mechanism there is, and still not universal.

How sites score

78

average overall health score

The weighted composite across every scored module, out of 100.

n = 116#

76

average authority distribution

How evenly internal authority reaches the pages that matter. Consistently the weakest module.

64

average topical clustering

Whether content forms coherent topic clusters with supporting pages around a pillar.

89

average content visibility

Share of pages whose raw HTML carries real text — how much of a site a crawler can actually read before running any JavaScript.

average on-page SEO

Titles, meta descriptions and heading structure. The area most teams have already worked on.

The shape of a typical site

129

pages in a typical audit

Average crawl size. Bounded by each plan's page budget, so this describes the audits, not the sites.

n = 116#

of a typical site sits one click from home

Share of crawled pages at depth 0 or 1. This is the raw measurement behind the flat-architecture verdict, and unlike that verdict it's computed on every audit regardless of size.

average click depth

Clicks from the homepage to a typical page.

words on a typical page

Average body word count across crawled pages.

n = 116#

What the numbers suggest

Content visibility and on-page SEO score far higher than authority distribution across this dataset. That gap is the whole argument for structural SEO: most sites have solved “can a crawler reach and read the page” and have not solved “does the page receive enough internal authority to rank once it is found”. Being reachable is table stakes; being supported is what moves positions.

Link saturation is the counter-intuitive one. A site where navigation links everything to everything scores well on any metric that counts links, and badly on the one that matters — whether authority concentrates anywhere in particular. More internal links is not the same as better internal linking.

Method, and what this is not

  • Sample size: 116 completed audits. That is enough to be indicative and too small to be definitive — treat these as directional.
  • Self-selected, not random. These are sites someone chose to audit, which skews toward sites whose owners already suspect a structural problem. It is not a sample of the web.
  • 0% of the sample comes from free public audits — the single-purpose tools and the anonymous audit form — which crawl far smaller page budgets than a signed-in audit. Page-count-derived figures — average pages, click depth, links per page — describe those crawls, not the whole sites behind them.
  • Small audits count. Everything a short crawl genuinely measures — scores, orphan counts, link saturation, depth share, llms.txt, sitemaps, on-page — is recorded from every audit regardless of size. The one exception is the flat-architecture verdict, which the analyzer only reaches on a complete crawl of 50+ pages; below that it returns the same answer for a flat site and a deeply nested one, so recording it would be recording nothing. The measurement underneath it (share of pages one click from home) is kept from every crawl instead.
  • 0% of crawls finished without hitting their page budget. Site-wide claims are only strictly honest on those, so any figure that truncation could fake — flat architecture, for one — is only ever recorded as present when the crawl completed. That makes those shares a floor.
  • Each share is reported against the number of samples that actually carry that measurement (the n under each figure), not against the whole sample. Measurements added to the dataset recently have a smaller n, and would otherwise read as absent when they were simply never recorded.
  • Fully anonymised. Each row stores structural measurements only — no URL, no domain, no page titles, no account identifier. Repeat audits are deduplicated with a salted one-way hash of the domain that is never published.
  • Nothing is published early: the endpoint behind this page returns no figures at all until the dataset passes 100 samples.
  • Updated continuously as more audits complete, so these figures move. The CSV is generated from the same query as the page.

Quoting a figure? Link the specific stat — every card on this page has its own permanent anchor — or download the full dataset as CSV.

See where your site sits

Run a free structural audit and compare your scores against the averages above. No account required, and your audit adds one anonymised row to this dataset.

Want the methodology behind each score? Read how the scores are calculated.