What an Audit of Your Own Estate Finds
We counted our own websites properly for the first time last week. Seven sites, 2,310 published pages.
384 of them were duplicates of each other — 102 groups sharing an identical title, with bodies 77% to 98% token-identical. One title existed fourteen times.
650 pages had no editorial inbound links at all. Not few. None. Reachable only from a paginated listing.
Twenty commercial pages — the ones that exist to win work — had no inbound links either. Including the main service page of one brand, and a division launched the week before.
None of this was hidden. Every page was visible, every file was in version control, and every individual piece of it had been reviewed by somebody at the time. It had simply never been counted.
Why nobody finds this from inside a project
The structural reason is worth stating, because it is the same in any organisation.
Each individual decision was reasonable. A post gets written. A page gets published. Nobody links to it from an old article because nobody is reading old articles while writing a new one. Multiply by four years.
Every review is local. A project review asks whether this piece of work is good. It has no mechanism for asking whether the estate now contains eleven versions of it, because the estate is not in scope of any project.
The failure mode is absence, and absence produces no signal. A page with no inbound links does not error. A duplicate does not conflict. Nothing anywhere fails, so nothing surfaces.
Which means these problems are found by counting or not at all — and counting is nobody’s job, because it belongs to no project.
The three findings, and what each cost
Duplication. 384 pages, 282 of them redundant, from a historical bulk-generation run. The cost is not storage. It is that search engines had to choose between near-identical pages, that inbound links were split across copies rather than accumulating on one, and that a substantial share of what we published described a business where content is generated rather than written.
Orphaned content. 650 posts with no editorial link pointing at them. Each of those cost something to produce and is now reachable only by someone paging through an archive. The work exists; the distribution never happened.
Orphaned commercial pages. The expensive one. Twenty pages whose entire purpose is to win work, with nothing pointing at them from any article. Including, in one case, the primary commercial page of a brand.
That last one is worth dwelling on, because it inverts the usual assumption. The pages we publish most often — blog posts — were linking to each other and to nothing that sells. Four years of content marketing, pointing inward.
★ Insight ─────────────────────────────────────
The three findings have one cause: we measured production and never measured the estate. Posts published, words written, pages shipped — all tracked. Duplicates, orphans and inbound distribution — none of them. A metric that counts output will always look healthy, because output is the thing being done. It takes a different class of measurement to notice that the output is accumulating into something nobody would have designed.
─────────────────────────────────────────────────
Two of our own measurements were wrong first
Worth reporting, because an audit that is not itself audited is just a more confident opinion.
The link analysis initially classified social redirect stubs — /facebook/, /gbp/ — as commercial pages. That filled the list of orphaned commercial pages with things that were never meant to have inbound links, and buried the suburb pages that were the actual finding.
The keyword analysis resolved ownership by whichever site declared a term first in a configuration file, disregarding which field it was declared in. That produced 183 reported violations against a site for using a keyword it owns. After the fix the real number was a small fraction of that.
Both were caught the same way: the result was surprising, so it was checked before it was acted on. Had either been reported as-is, the remediation would have been a quarter of work aimed at nothing.
What we did
Consolidated 244 duplicates onto 102 canonical pages, with internal links repointed first so nothing routes through a redirect, and 158 pre-existing redirects repointed because they targeted pages being retired.
Twenty-five groups were held back rather than consolidated: same title, genuinely different content, which is a rewrite rather than a redirect. Automating past that distinction would have destroyed content while the report claimed a tidy-up.
Then added links to every orphaned commercial page, from published posts where the link is genuinely relevant.
Why we could not simply delete the duplicates
Worth explaining, because “delete the duplicates” sounds like the obvious response and it is the wrong one.
Those 384 pages held 214 inbound links between them, spread across copies. Deleting them discards that; consolidating with redirects accumulates it onto one page. The difference is the entire commercial value of the exercise.
That determined the keep rule. Not the newest copy, not the longest — the one holding the most inbound links, with date and length as tiebreakers. Somewhat counter-intuitively that occasionally meant keeping an older page, because links accumulate over time.
It also determined the order of operations. Internal links pointing at retiring pages were repointed at the survivor before anything was deleted, so no surviving page links through a redirect. Once the file is gone, the mapping that tells you what to repoint to is gone with it.
And a check we nearly missed: 158 existing redirect rules already pointed at pages on the retirement list. Deleting those pages without touching the rules would have turned 158 working redirects into redirects into errors.
None of that is difficult. All of it has to happen before the deletion, which is the part that makes a bulk cleanup a project rather than a command.
The executive version
Three questions that apply to any accumulated estate — content, applications, data stores, spreadsheets, integrations.
1. When did anyone last count it? Not audit for quality. Count. How many are there, and is that number higher than anyone would have guessed.
2. What percentage is duplicate? Ours was 11% by page count. For an application portfolio the equivalent question is how many systems do substantially the same job.
3. What has nothing pointing at it? In content, inbound links. In an application portfolio, systems no process depends on. In data, tables nothing queries. The answer is always more than expected, and each one costs something to maintain.
The part that should be uncomfortable
We publish detailed engineering notes about our own defects. We run gates on our own content. We had written, that same fortnight, about controls that cannot fail and records that disagree with reality.
And we had never counted our own pages.
The general lesson is not that audits are good. It is that the work of a system is always more visible than the shape of it, and every incentive in every organisation points at the work. Somebody has to periodically stop and count, it will be uncomfortable the first time, and nothing in the normal operation of the business will ever prompt it.
Estate audits, content governance and the measurement work underneath them are part of the advisory we do through Ganda Tech Services, with web and content operations through Cosmos Web Tech.
Digital Transformation Roadmap 2026
A 12-month framework for Australian SMBs ready to modernise — phases, tools, and milestones.
Almost done
Check your inbox and click the confirmation link to get your download.