All posts
internal linking

Internal Link Audit: Are Your Content Clusters Connected?

A reproducible internal link audit for small SaaS blogs: export your links, check hub links, broken targets, redirects and isolated clusters, then fix them.

By

TL;DR An internal link audit is not a hunt for one bad link. It is a check that the cluster map in your head matches the link graph Google actually crawls: every spoke links to its hub, the hub links back, no link points at a 404 or a redirect, and no cluster sits on an island. One crawl export, five checks, and a findings-to-action table get you there in an afternoon.

Last month I pulled a link export for a 60-post SaaS blog that I was sure was well linked. I had planned every cluster. I had a linking step in my publishing checklist. The export said otherwise: four spokes in one cluster linked to a pillar URL I had renamed in spring, so every one of those "hub links" went through a redirect. One cluster of six posts linked tightly to itself and received exactly one link from the rest of the site. Nothing was broken in a way that would ever show up as an error. The clusters were just quietly less connected than I believed.

That gap between the plan and the graph is what this post is about. Designing clusters is covered in my internal linking strategy for topic clusters. Finding pages with zero inbound links is covered in the orphan pages SEO guide. This is the broader verification pass that sits on top of both.

What an internal link audit actually checks

Most guides treat an internal link audit as "run a crawler, fix the red rows." That catches broken links, and broken links matter. But on a small site, the expensive failures are structural, not technical. A link that returns 200 can still be the wrong link.

Google's own guidance is short and blunt: Google can only crawl a link that is a proper anchor element with an href, and "every page you care about should have a link from at least one other page on your site," per the Search Central link best practices. The orphan check covers the "at least one" part. The audit asks the harder question: is it the right one, pointing at the final URL, from inside the right cluster?

I check five things:

  1. Missing hub links: spokes that do not link to their pillar, or pillars that do not link to all their spokes.
  2. Broken targets: internal links pointing at 4xx or 5xx URLs.
  3. Redirected targets: internal links pointing at a URL that redirects somewhere else.
  4. Isolated sections: clusters that link well internally but receive almost nothing from the rest of the site.
  5. Weak links: links that exist but carry no meaning: generic anchors, nofollow on internal links, or links that are not crawlable anchors at all.

On a low-authority site, this matters more than it does for an incumbent. You have almost no external links to lean on, so the internal graph is the main way you tell Google which page is the hub and which pages belong together. If that graph has holes, your cluster is a list of posts that happen to share a theme.

Step 1: Get a link export and a cluster map

You need two files. Neither costs anything.

The link export. Any internal links checker that can export source, destination, anchor text and the destination's status code will do. The Screaming Frog SEO Spider is the default for a reason: it is free for crawls up to 500 URLs, which covers almost every bootstrapped SaaS blog, and its Bulk Export menu includes All Inlinks, which is one row per link found on the site. For a Screaming Frog internal link audit, crawl your domain, then export Bulk Export, Links, All Inlinks to CSV. Filter that file to hyperlinks only, so images and stylesheets do not pollute the counts.

If you do not want a crawler, a static site generator can usually produce the same thing by parsing your Markdown files for internal links. The tool matters far less than the shape of the output.

The cluster map. A two-column sheet: URL and cluster name, with the pillar marked. If you built a keyword mapping template, you already have this. If you do not, this is the moment you find out your clusters were never written down, which is a finding in itself.

The audit is a join between those two files. The crawl tells you what is linked. The map tells you what should be.

Step 2: Run the five checks

Here is a small script I use against a Screaming Frog All Inlinks export. Column names can differ slightly between versions, so check your header row first.

import csv
from collections import defaultdict

# cluster_map.csv columns: url, cluster, is_pillar (yes/no)
clusters, pillars = {}, {}
for row in csv.DictReader(open("cluster_map.csv")):
    clusters[row["url"]] = row["cluster"]
    if row["is_pillar"] == "yes":
        pillars[row["cluster"]] = row["url"]

links = [r for r in csv.DictReader(open("all_inlinks.csv"))
         if r["Type"] == "Hyperlink"]

edges = {(r["Source"], r["Destination"]) for r in links}
inbound_from_outside = defaultdict(int)

for r in links:
    code = r["Status Code"]
    if code.startswith(("4", "5")):
        print("BROKEN", r["Source"], r["Destination"], code)
    elif code.startswith("3"):
        print("REDIRECT", r["Source"], r["Destination"])
    if r["Follow"].lower() == "false":
        print("NOFOLLOW", r["Source"], r["Destination"])
    src, dst = clusters.get(r["Source"]), clusters.get(r["Destination"])
    if dst and src != dst:
        inbound_from_outside[dst] += 1

for url, c in clusters.items():
    hub = pillars.get(c)
    if hub and url != hub:
        if (url, hub) not in edges:
            print("SPOKE MISSING HUB LINK", url)
        if (hub, url) not in edges:
            print("HUB MISSING SPOKE LINK", url)

for c in set(clusters.values()):
    print("CLUSTER", c, "inbound from other sections:", inbound_from_outside[c])

Every line it prints is a finding. Here is how to read each type.

Missing hub links

This is the check almost nobody runs, because a crawler cannot know your intended structure. A crawler sees links. It does not know that post X is supposed to be a spoke of pillar Y. Only the join with your cluster map surfaces it.

Pay special attention to the "hub missing spoke link" rows. Pillars get written early and spokes get added for months afterwards. If you never go back to update the pillar, every new spoke is connected in one direction only. The pillar is the page that should be gathering relevance from the whole cluster, and it is usually the one that is most out of date.

Broken targets

Any internal link to a 4xx or 5xx URL. The Screaming Frog Response Codes tab splits these into client errors and server errors if you prefer the UI to a script. On a small blog these almost always come from one of two places: a post you deleted or merged, or a slug you edited after publishing. Fix the link at the source. Do not just add a redirect and walk away, because then you have moved the problem into check 3.

Redirected targets

Internal links that hit a 3xx before reaching the final page. They work for users, so nobody notices them. But you control both ends of an internal link, so there is no reason for one to route through a redirect. Google's redirects documentation explains how it treats permanent and temporary redirects as signals about which URL is canonical. Your internal links should send that signal directly by pointing at the canonical URL, not by asking Google to follow a hop. If you also have duplicate URL and canonical issues, redirected internal links make them harder to diagnose.

Redirected hub links are the worst version: the most important link in the cluster pointing at a URL that is no longer the hub.

Isolated sections

A cluster can be internally perfect and still be an island. The script's last block counts how many links each cluster receives from pages outside it. If one cluster gets dozens and another gets one, the second cluster is depending entirely on your sitemap and whatever its pillar earns on its own. That is not orphaning: every page has inbound links, so the orphan page check will not catch it.

The fix is not to spray links. It is to find the two or three posts in other clusters where the isolated topic genuinely comes up, and link from there to the pillar with a descriptive anchor.

Weak links

Three sub-checks. First, anchors like "here", "this post" or "read more" pointing at cluster pages. Google's link guidance says anchor text "tells people and Google something about the page you're linking to" (source), so a generic anchor is a wasted signal. Second, nofollow on internal links, which is almost always a plugin or theme default nobody chose. Third, navigation built with JavaScript click handlers rather than real anchor elements, which Google says it cannot crawl as links per the same page.

How many internal links per page for SEO?

This is the question I get most, so here is the honest answer: there is no target number. Google says it directly: "there's no magical ideal number of links a given page should contain. However, if you think it's too much, then it probably is" (Google Search Central).

Any internal link audit tool that flags a page for having "too few" or "too many" links is applying its own made-up threshold. I ignore those flags. What I check instead is coverage: does each spoke link to its hub, does the hub link to each spoke, and do siblings link where the reader would actually want to go next. A spoke with three links that hit those marks is better linked than one with twenty links to random posts.

The findings-to-action table

This is the table I paste into the audit notes. Each finding has one default action and one owner check, so the audit ends in a to-do list rather than a spreadsheet nobody opens again.

FindingWhat it meansDefault actionCheck after fixing
Spoke missing hub linkSpoke does not pass relevance to its pillarAdd a contextual link to the pillar with a keyword-descriptive anchorRe-crawl shows the edge exists
Hub missing spoke linkPillar is out of date with its clusterAdd the spoke to the pillar where its subtopic is discussedPillar links to every spoke in the map
Broken target, page deletedLink points at nothingRepoint to the closest live page, or remove the linkNo 4xx rows in the export
Broken target, slug changedLink uses an old URLUpdate the link to the current slug, then add a redirect for external visitorsLink returns 200 directly
Redirected targetLink routes through a hopReplace with the final destination URLNo 3xx rows in the export
Isolated clusterSection receives almost no links from the rest of the siteAdd 2 to 3 links from genuinely related posts in other clusters to the pillarCross-cluster inbound count rises
Generic anchorLink carries no topical signalRewrite the anchor to describe the targetNo "click here" style anchors to cluster pages
Internal nofollowLink signal suppressed by default settingRemove nofollow from internal links at the template levelFollow is true across the export
Page in no clusterPost has no home in the mapAssign it to a cluster, merge it, or retire itEvery URL in the crawl appears in the map

That last row is where audits turn into content decisions. A post with no cluster is usually a sign of a one-off that never fit the plan. Treat it the way a content audit would: keep and integrate, merge, or remove. If two posts in the same cluster keep fighting for the same query, that is a keyword cannibalization problem, and no amount of linking will fix it.

How often to run an internal links check

For a blog shipping several posts a week, I run the full audit monthly and a narrow version on every publish. The narrow version is two questions in the pre-publish checklist: does the new post link to its hub, and have I added it to the hub? That catches most missing hub links before they exist.

The monthly pass catches everything the checklist cannot: slugs changed after publishing, posts merged during a refresh, clusters that grew lopsided. I also run it after any template, CMS or URL structure change, because those break links in bulk.

If the audit keeps finding the same isolated cluster, look upstream. Sometimes the problem is not linking discipline but planning: a cluster that has no natural overlap with anything else on your site, picked because it had volume rather than because it connected to your product. That is why I score topics for winnability and cluster fit before writing them, not after. It is also the gap I built Boomranq to close: when the 30-day calendar already defines the hub, the spokes and the cross-links for each post, the cluster map exists before the first word is written, and the audit becomes a quick confirmation instead of an archaeology project.

The point of an internal link audit is not a clean report. It is making sure the structure you planned is the structure Google crawls. On a site without authority, that structure is most of what you have.

Want this planned for your site?

Boomranq turns your product into a sequenced 30-day calendar built around what your site can actually rank for. Join the waitlist for early access.

No spam · One email when we launch · Unsubscribe anytime