Every B2B site I audit that is more than about five years old has the same layer of sediment underneath it.

Landing pages from a trade show in 2022, still live, still indexed, still mentioning a product name that changed. A URL structure from before the rebrand, redirecting to a second URL structure from after the rebrand, which redirects again to the current one. Twelve city pages that differ by one word. Schema markup listing an office the company moved out of two years ago.

Nobody built this on purpose. It happened one campaign at a time, and the cleanup was never anybody's job.

How the debt accumulates

The mechanism is always the same. A campaign needs a page, so someone builds one fast. The campaign ends. Nobody removes the page, because removing things feels risky and leaving them feels free.

Multiply that by five years, two agencies, one rebrand, and a CMS migration, and you get a site where the number of live URLs is two or three times the number of pages anyone would say the company has.

The individual items are small. The combined effect is a site that is expensive to crawl, confusing to categorise, and inconsistent about what the business does.

The reason it stays hidden is that none of it breaks anything visibly. The site loads. Forms work. The homepage looks fine. Technical debt on a website does not announce itself, it just slowly makes everything else you do less effective, so the symptom people notice first is usually "our content is not working any more" rather than anything technical at all.

Redirect chains from old campaigns

A redirect chain is a URL that redirects to a URL that also redirects. Sometimes three or four deep on a site that has been through a rebrand and a platform change.

Each hop costs the user time and costs you crawl budget. Google follows a limited number of hops before it stops, so a long chain can mean the final page is never reached from that starting URL. And chains make everything harder to debug: when a link breaks three steps into a chain, tracing it back to the rule that created it takes real time.

Find them by crawling the site with Screaming Frog or Sitebulb and opening the redirect report. Then fix them at the source. Point the first URL directly at the final destination instead of patching a new rule on top of the old ones. Also read your redirect configuration as a whole file rather than line by line, because that is where you see rules from three different years contradicting each other.

One rule while you are in there: stop redirecting dead pages to the homepage. Search engines treat an irrelevant redirect as a soft 404, so you get the confusion of a redirect with none of the benefit. A genuine 404 for a page that should not exist is a correct answer.

Orphaned pages nobody links to

An orphaned page has no internal links pointing to it. It sits in the sitemap, or it is reachable if you know the URL, and nothing on the site connects to it.

Internal links are how search engines work out which of your pages matter. A page nothing links to reads as unimportant, no matter how good the content is. B2B sites collect orphans faster than most because campaign pages are designed to be linked from ads and emails, never from the site itself, and then the campaign ends.

Compare your crawl results against your sitemap. Anything in the sitemap that the crawler could not reach by following links is an orphan. Then decide per page: link it properly if it deserves traffic, merge it into a stronger page if it overlaps with one, or remove it. Building the habit of auditing content on a schedule is what stops orphans from accumulating again after you clean them up.

Thin and duplicate location pages

This is the item most likely to cause real damage, and the one companies defend hardest.

The pattern: one template, twelve cities, the city name swapped in the heading, the intro, and the meta title. Everything else identical. It is fast to build and it sometimes works for a while.

Google's documentation on doorway pages describes this directly, and the risk is not limited to the pages themselves. A large cluster of near-identical pages affects how the whole site is assessed.

Here is the test I use. Remove the city name from the page. Can a reader tell which page they are on? If the answer is no, it is one page wearing twelve hats.

Fixing it means either making each page genuinely different or reducing the count. Genuinely different means specific clients in that place, work you actually did there, questions that market actually asks, local conditions that change your advice. A page about serving manufacturers in Burnaby should read differently from one about professional services firms downtown, because the businesses are different and so is the advice. If you cannot write that for a city, you should not have a page for that city.

The uncomfortable version of this fix is deleting nine pages and writing three good ones. That is usually the right answer.

Stale schema and outdated structured data

Structured data is a set of statements about your business that machines read as fact. Nobody looks at it, which is exactly why it goes wrong quietly.

What I find most often on B2B sites: an old address after a move, opening hours from before the schedule changed, a service list including things the company stopped selling, review markup for reviews that no longer exist anywhere, author markup naming people who left two years ago, and product or pricing markup left from a page that was rewritten.

This matters more now than it did five years ago, because AI assistants read structured data when answering questions about a company. Wrong markup produces confidently wrong answers, and it is worse than no markup, because absent information invites a check while present information gets repeated.

Run your key page types through Google's Rich Results Test and Schema Markup Validator, and read the output as claims rather than as code. Is every statement still true? Fix the false ones. Delete markup for things that no longer exist.

Content only JavaScript can see

This is the newest item on the list and the one most teams have not checked.

Google renders JavaScript for its index. It costs Google resources, but it works. Several AI crawlers do not render at all. They fetch the raw HTML document and read what comes back, and anything your JavaScript inserts after page load does not exist as far as they are concerned.

So you can have a page that ranks perfectly well in Google and is functionally blank to a system that only reads raw HTML.

Check it in one command. Run curl against a page and read the response. That is the document before any script runs. If your headings, body copy, and JSON-LD are present, you are fine. If you get an empty container and a bundle reference, your content is invisible to non-rendering crawlers.

The fix is server-side rendering or prerendering, which turns a single-page application into real HTML at the URL. This is a development project rather than a settings change, so it needs to be scoped honestly. But it is worth checking before you spend another quarter writing content those crawlers cannot read. If you are already thinking about how AI assistants find and cite your business, this is the technical floor underneath all of it.

How to audit and prioritise

Start with a crawl, then run three checks the crawler will not do for you.

  1. Crawl the whole site. Screaming Frog is free up to 500 URLs, which covers most B2B sites. Export the redirect report, the 4xx report, the orphan comparison against your sitemap, and the duplicate title and meta description lists.
  2. Curl your five most important pages. Confirm the main content and the structured data are in the raw HTML.
  3. Read your structured data as statements. Check each claim against reality.
  4. Open Search Console page indexing. Look at what Google chose not to index and why. That report tells you which of your problems Google has already noticed.

Then fix in this order, which is by cost rather than by any tool's severity label.

First: anything blocking indexing on pages that make money. A stray noindex tag left from staging, a robots.txt rule that was meant to be temporary, a canonical pointing at the wrong URL. These are usually quick and occasionally enormous.

Second: content crawlers cannot read at all. No amount of writing helps if the page is empty to the reader.

Third: thin and duplicate clusters. This is the site-wide quality risk, and it takes the longest, which is why it should start early.

Fourth: redirect chains, orphans, and stale schema. Real problems, rarely urgent, and satisfying to clear in one focused week.

Then keep it from coming back. A short monthly crawl checking for new 404s, new chains, and new orphans takes under an hour and catches most of what would otherwise become next year's audit.

The takeaway

Technical debt on a website behaves like technical debt in code. It accumulates from reasonable decisions made under time pressure, and it stays invisible until it is expensive.

Crawl the site. Curl your top pages. Read your schema as claims. Fix what is blocking money first and the tidy-up items last. Most B2B sites do not need more content. They need the content they already have to be reachable.