Original research
The State of Structural SEO
Every RankForge audit contributes one anonymised row to this dataset — no URL, no domain, no page titles, just the structural measurements. This is what those rows say so far.
Free to quote and republish with attribution and a link to this page. Download the dataset (CSV).
What’s broken, and how often
24.1%
of sites have orphan pages
At least one indexable page with no internal links pointing to it — invisible to crawlers that arrive by following links, and invisible in most tools, because a page with no inbound links isn't an error anywhere.
—
of sites are link-saturated
Navigation links nearly every page to nearly every other, so internal authority lands almost uniformly. It looks like a well-linked site and behaves like an unlinked one — no page is prioritised over any other. Detected on sites of every size: small sites register through near-identical authority across their top pages.
—
have a flat architecture
Nine in ten pages sit one click or less from the homepage. This verdict needs a complete crawl of 50+ pages to be decidable at all, so it's measured on that subset — but the underlying depth measurement below uses every audit.
—
broken pages in a typical crawl
Internal links resolving to an error status. Every one of them is authority spent on nothing.
—
of internal links are editorial
Links inside content rather than in navigation, headers or footers. Editorial links carry the topical signal; boilerplate links mostly carry noise.
AI and agent readiness
—
of sites publish an llms.txt
The proposed convention for telling AI agents which pages matter. Measured only on audits where the check actually ran, so an un-probed crawl never counts as an absence.
—
average GEO readiness score
Agent-discoverability signals combined: llms.txt, identity schema, and agent-readable link headers. Low scores here are the norm, not an outlier.
—
of sites have an XML sitemap
The oldest discovery mechanism there is, and still not universal.
How sites score
78
average overall health score
The weighted composite across every scored module, out of 100.
—
average internal link strategy
Whether internal links are deliberate — pointing at pages that need support — or incidental.
64
average topical clustering
Whether content forms coherent topic clusters with supporting pages around a pillar.
89
average content visibility
Share of pages whose raw HTML carries real text — how much of a site a crawler can actually read before running any JavaScript.
—
average on-page SEO
Titles, meta descriptions and heading structure. The area most teams have already worked on.
The shape of a typical site
129
pages in a typical audit
Average crawl size. Bounded by each plan's page budget, so this describes the audits, not the sites.
—
of a typical site sits one click from home
Share of crawled pages at depth 0 or 1. This is the raw measurement behind the flat-architecture verdict, and unlike that verdict it's computed on every audit regardless of size.
—
internal links per page
Counting links between crawled pages only. High numbers here usually mean navigation, not editorial linking.
What the numbers suggest
Content visibility and on-page SEO score far higher than authority distribution across this dataset. That gap is the whole argument for structural SEO: most sites have solved “can a crawler reach and read the page” and have not solved “does the page receive enough internal authority to rank once it is found”. Being reachable is table stakes; being supported is what moves positions.
Link saturation is the counter-intuitive one. A site where navigation links everything to everything scores well on any metric that counts links, and badly on the one that matters — whether authority concentrates anywhere in particular. More internal links is not the same as better internal linking.
Method, and what this is not
- Sample size: 116 completed audits. That is enough to be indicative and too small to be definitive — treat these as directional.
- Self-selected, not random. These are sites someone chose to audit, which skews toward sites whose owners already suspect a structural problem. It is not a sample of the web.
- 0% of the sample comes from free public audits — the single-purpose tools and the anonymous audit form — which crawl far smaller page budgets than a signed-in audit. Page-count-derived figures — average pages, click depth, links per page — describe those crawls, not the whole sites behind them.
- Small audits count. Everything a short crawl genuinely measures — scores, orphan counts, link saturation, depth share, llms.txt, sitemaps, on-page — is recorded from every audit regardless of size. The one exception is the flat-architecture verdict, which the analyzer only reaches on a complete crawl of 50+ pages; below that it returns the same answer for a flat site and a deeply nested one, so recording it would be recording nothing. The measurement underneath it (share of pages one click from home) is kept from every crawl instead.
- 0% of crawls finished without hitting their page budget. Site-wide claims are only strictly honest on those, so any figure that truncation could fake — flat architecture, for one — is only ever recorded as present when the crawl completed. That makes those shares a floor.
- Each share is reported against the number of samples that actually carry that measurement (the n under each figure), not against the whole sample. Measurements added to the dataset recently have a smaller n, and would otherwise read as absent when they were simply never recorded.
- Fully anonymised. Each row stores structural measurements only — no URL, no domain, no page titles, no account identifier. Repeat audits are deduplicated with a salted one-way hash of the domain that is never published.
- Nothing is published early: the endpoint behind this page returns no figures at all until the dataset passes 100 samples.
- Updated continuously as more audits complete, so these figures move. The CSV is generated from the same query as the page.
Quoting a figure? Link the specific stat — every card on this page has its own permanent anchor — or download the full dataset as CSV.
See where your site sits
Run a free structural audit and compare your scores against the averages above. No account required, and your audit adds one anonymised row to this dataset.
Want the methodology behind each score? Read how the scores are calculated.