DeepSmith

Aug 26 · Content Strategy

18 min read

How to Audit Your Existing Content Inventory Before Planning a Cluster

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
An abstract monochrome diagram of layered page cards linked into clusters, with one card unconnected and one empty outlined slot, beside the cover line Map what you already own.

You are about to plan a topic cluster, and something keeps nagging at you. Do we already have a page on this? Probably. Where is it? No idea.

That nagging feeling is the whole reason for a content inventory audit. When you audit existing content before you plan a single new spoke, you build around your real coverage instead of quietly publishing a second version of a page you forgot about.

By the end of this you will have every relevant URL in one sheet, each page tied to a topic, an intent, and a funnel stage, overlapping pages grouped together, and a short list of gaps you can actually defend.

Eight steps. Take them one at a time.

Step 1: Set your scope and decide what counts as one row

Start by naming what is in and what is out. Blog, resource center, documentation, product pages, glossary, templates, and any PDFs or videos that serve the same audience.

Then decide if this is a full pass or a partial one. A partial pass can cover one subdirectory, one content type, one topic, or one funnel stage. On a large site, a narrow first pass beats a master sheet nobody maintains.

Now the part people skip: decide what one row means. For cluster planning, one row is one canonical, indexable page or asset. URL variants, parameters, redirects, and duplicate versions go in a separate technical exceptions tab, not in your working map.

Last, write the purpose at the top of your content mapping spreadsheet: map existing coverage before planning new pages. That one line will stop three arguments later.

Done when: your team can say which properties are included, what date range applies to performance data, what a row represents, and where excluded assets live.

Common mistake: starting from a list of keywords you want to rank for. That builds a topic backlog, not a content audit before cluster work begins. Start with URLs you already have. Nothing gets called a gap until the existing inventory is cleaned up and classified.

Step 2: Build the content mapping spreadsheet

You need somewhere to put all of this, and a plain spreadsheet is fine. Freeze the header row, keep a separate raw-import tab, and keep the working map readable.

Five columns carry the whole method:

FieldWhat to record
URLThe canonical URL being audited, one per asset
TopicThe durable subject the page covers, like "email deliverability," not a phrase lifted from the title
IntentThe job the reader wants done: informational, commercial investigation, transactional, or navigational
StageAwareness, Consideration, or Decision
Gap flagYes, No, or Review

Those five are the minimum. Around them, add the working fields you will actually use: page title and H1, content type, page purpose, audience or persona, topic definition, parent hub or cluster, author or owner, created and last-updated dates, status, indexability and canonical URL, clicks and impressions and CTR and average position, page views and conversions, internal inlinks and link depth, quality notes, accessibility notes, overlap group, gap reason, evidence source and date checked, and who owns the next action.

That is a long list, and you do not need all of it on the first pass.

Pro tip: resist the urge to fill every optional column before you have finished a single one of the required five. A content inventory audit can take hours on a small site and weeks on a big one. A focused sheet that gets finished is worth more than a beautiful one that stalls in week two.

Done when: the five required columns exist, the header is frozen, and raw imports sit on their own tab.

Step 3: Collect every existing page

No single source knows about all your pages. So use several and reconcile them.

  1. Export the CMS or sitemap URL list, and include every sitemap you have: blog, product, docs.
  2. Crawl the site to collect status codes, titles, headings, canonicals, metadata, and internal links.
  3. Pull pages from analytics and Search Console, including pages that get traffic but never show up in your navigation or your crawl.
  4. Add assets that are not ordinary HTML pages, like PDFs, videos, and downloads, if they serve the same audience.
  5. Keep the raw URL and its source on every row before you clean anything.

Why all four? Because a sitemap helps search engines discover URLs, but it does not guarantee that Google crawls or indexes any of them. Google also finds pages by following links from pages it already knows. So the sitemap is one input, not the truth.

A few tools do most of the lifting. Search Console's Performance report gives you clicks, impressions, CTR, and average position, and the Pages dimension is the one that joins cleanly to your rows. Its default view covers the past three months, and you can change that. Watch your exports: values shown as unavailable can download as zeros, which will quietly lie to you if you sort by them. URL Inspection tells you whether Google has a specific URL indexed and why or why not. A crawler like Screaming Frog surfaces broken links, redirects, metadata, orphan candidates, and internal-link relationships.

This is also where a platform can save you a week. DeepSmith's Content Map crawls your site, keeps it current by re-checking sitemaps every 24 hours, and folds newly published pages in automatically under Recently published. You still reconcile it against analytics and Search Console, because those answer questions a crawl cannot.

Done when: every included URL carries a source, and pages that showed up in only one source are still on the sheet, marked for investigation.

Where people go wrong:

  • Treating the sitemap as the complete inventory.
  • Crawling only the blog and missing product, docs, and resource pages.
  • Dropping redirecting, noindex, orphaned, or low-traffic pages before deciding their role.
  • Mixing URL variants so one page looks like four.
  • Recording only pages that rank. A page with no traffic is still content you own, and it might be the page you should update.

Step 4: Normalize URLs and remove technical duplicates

Right now your sheet has the same page in it more than once. That is normal, and this step fixes it.

Normalize protocol, hostname, trailing slashes, case where it matters, parameters, fragments, and tracking variants. Resolve every redirect and record where it lands. Read each declared canonical and compare it to the URL that actually loads and the URL your internal links point at.

Then group the duplicates. Keep two kinds of grouping apart: technical duplication, which is one page reachable at several URLs, and editorial overlap, which is two genuinely different pages covering related ground. Only the first one belongs in this step.

Pick one representative row for each canonical asset and park the rest as aliases.

A little context helps here. Google treats canonicalization as a deduplication process, choosing one version of otherwise duplicate content to show. Signals differ in strength: a redirect and a rel-canonical annotation are strong, and being listed in a sitemap is weak. Contradictory signals should be fixed rather than layered. Do not use robots.txt or the removal tool to pick a canonical, do not use a fragment, and do not reach for noindex to choose between two pages on your own site. Link consistently to the URL you want to win.

Done when: each real asset has exactly one canonical row, every alternate URL points back to it, and your sheet knows the difference between technical duplication and two related pages.

Where people go wrong:

  • Merging two pages because their titles share a word.
  • Keeping URL variants as if they were separate cluster pages.
  • Assuming a declared canonical guarantees the one Google picks.
  • Canonicalizing pages that serve different audiences, jobs, or stages.
  • Deleting a page before checking its links, backlinks, conversions, and history.

Step 5: Classify each page by topic, intent, and stage

Here is where the audit stops being a list and starts being a map. Open the pages. Really open them, not just their titles or their target keywords.

For each row, record six things:

  1. Topic. The durable subject, plus a one-line definition of what belongs in it. Write the definition and "email deliverability" stops collapsing into "email subject lines" just because both say email.
  2. Intent. What the reader is trying to get done. Does the page teach, compare, help someone choose, help someone implement, or close a sale? Record the dominant one, then note a meaningful secondary.
  3. Stage. Awareness for problem and concept education, Consideration for approaches and alternatives, Decision for product, vendor, and implementation questions.
  4. Cluster role. Existing hub or pillar, supporting page, decision page, glossary entry, or outside the cluster entirely.
  5. Audience. Persona, role, industry, or maturity, wherever that changes the page.
  6. Coverage depth. Does the page answer the topic, or does it just mention it?

Remember what a cluster actually is. It is not a pile of pages with similar keywords. It is a relationship between a central topic, a broad hub, and more specific pages underneath. Your classification has to capture the job, the stage, and the audience, or the relationship will not hold.

Need a shortcut when a page is hard to place? Finish this sentence for it: "This page helps [audience] [do what] about [topic] at [stage]."

If two pages produce the same sentence, put them in an overlap review group. If the job, the audience, or the stage differs, they can both stay.

Classifying a few hundred pages by hand is slow, and this is the second place a platform earns its keep. DeepSmith's Content Map enriches and classifies every page onto one granular topic taxonomy with auto-named topics and definitions, plus the Awareness, Consideration, and Decision stages, so you can open a topic and see how your coverage splits by stage and which pages sit at each one. Treat that as your first draft of the map, then spot-check it against the live pages.

The DeepSmith Content Map opens a single topic and splits your own coverage of it across Awareness, Consideration and Decision, listing the specific pages sitting at each stage so a thin stage is visible at a glance. The figures shown are demo data.

Done when: every row has a topic, a topic definition, an intent, a stage, and a cluster role, and you can filter by topic and see the whole existing path from awareness to decision.

Where people go wrong:

  • Using the raw title as the topic.
  • Treating each keyword variation as its own page.
  • Assigning a stage from the URL folder.
  • Calling a page a pillar because it is long.
  • Classifying from metadata without reading what the page actually promises.

Step 6: Judge quality and performance, then give every page an action

An inventory tells you what exists. It cannot tell you what condition any of it is in. That is what the audit half adds, and it is the half that keeps you from calling a tired page a gap.

For each page, check accuracy and whether its claims, screenshots, and product details are still current. Check completeness against what the page set out to do. Check whether it adds enough to justify its own URL. Check clarity: plain language, headings, useful link text, chunking, whitespace. Check accessibility, including contrast, alternative text, and text buried inside images. Check search performance and business value. Check the internal links pointing to it and out of it. Check that it is indexable and discoverable.

Then give it one status:

  • Keep. Accurate, useful, distinct, and in the right place.
  • Update. The intent or the authority is worth having, but the page needs real work.
  • Consolidate. Two or more overlapping pages should become one stronger resource.
  • Redirect. The URL is obsolete or replaced, and readers should land somewhere better.
  • Remove. No purpose, no history worth keeping, no reason to hold the asset.
  • Review. The evidence is thin, or someone else has to decide.

Be careful with the numbers while you do this. Clicks are clicks from Google results. Impressions are appearances in results. CTR is clicks divided by impressions. Average position is the average position of your topmost result. All useful, none of them a quality score. A decision-stage page can carry the whole pipeline on a fraction of the traffic your best awareness post gets.

Pro tip: keep quality status and gap status in two separate columns. A topic can have coverage that is poor. Another topic can have nothing at all. Blend those two cases and your gap list will inflate, and you will end up commissioning a page that already exists.

Done when: every page has evidence-based notes and one action status, and nobody is using the word gap to mean bad page. This is the step that makes people glad they chose to audit existing content first.

Step 7: Deduplicate on topic plus intent, then find content gaps

This is the step the whole audit was building toward, and it comes in two parts.

Part A: Dedup on topic plus intent, not raw title

Here is the rule to write on a sticky note: the dedup key is topic plus intent, never the title.

Work it like this:

  1. Group your rows by normalized topic.
  2. Inside each topic, group by dominant intent.
  3. Compare audience, stage, promise, depth, and conversion role across the group.
  4. Mark pages distinct when they serve different jobs or different stages.
  5. Mark them for consolidation or repositioning when they make the same promise to the same audience at the same stage.
  6. Use ranking data as a clue, not as the verdict. Several URLs showing up for one query is worth a look, not proof.

Cannibalization gets used loosely, and it costs teams good pages. It is usually described as two pages targeting the same keyword. The more useful test is whether they compete for the same intent. A shared keyword is harmless when one page is a definition, one is a comparison, and one is a decision page. And two pages can absolutely overlap while using completely different words in their titles.

Part B: Flag gaps only after the dedup pass

Now build a coverage matrix. Rows are your normalized topics. Columns are intent and stage. Mark every cell:

  • Covered. A distinct, useful page does the job.
  • Covered, weak. A page exists but needs an update, a consolidation, or better internal support.
  • Partial. The topic is mentioned, or a neighboring intent is served, but the actual job is not done.
  • Uncovered. Nothing serves this combination.
  • Review. The classification or the business relevance is still unclear.

A true gap is an uncovered or materially undercovered combination, and it only counts after canonical cleanup, overlap review, and quality assessment. That order is what lets you find content gaps you can defend instead of a list of things that felt missing.

One page serving a topic, intent, and stage well is coverage you keep and plan around. One weak page there is an update, not a gap. More than one page there, all of them decent, is a question about the job, the audience, and the stage before it is anything else. More than one weak page there is a consolidation into one page. No page at all is the only true gap.

A two-by-two matrix with pages serving this topic, intent and stage across the top and quality of that coverage down the side: one strong page means keep it and plan around it, one weak page means update it rather than call it a gap, several strong pages means checking the job, audience and stage, several weak pages means consolidating into one page, and no page at all is the only true gap.

Competitors come in here, and they come in carefully. A topic a competitor covers heavily and you do not is evidence worth investigating, not an instruction to publish. Compare like with like: same topic definition, same audience, same intent, same stage, and a real reason the topic matters to your business.

If you want that comparison without building it by hand, DeepSmith's Content Map maps competitor sites onto your own taxonomy, then separates coverage gaps, where a competitor publishes more than you do on a shared topic, from untapped topics, where they publish and you have nothing. The judgment call stays yours. The tool narrows where you have to make it.

Done when: every proposed gap carries a reason, like "no Consideration page for this topic" or "the existing page serves Awareness only," and every overlap group has an action or a named owner.

Where people go wrong:

  • Counting competitor URLs and calling the difference a gap.
  • Commissioning a new page where an existing one should be expanded.
  • Treating a missing keyword as a missing topic.
  • Ignoring decision-stage and product pages because the blog looks healthy.
  • Calling every multiple-ranking-URL situation cannibalization.
  • Applying volume or traffic thresholds the audit never established.

Step 8: Validate the map and hand it to planning

Almost done. One pass to make sure the sheet is trustworthy, because this is what the next planner will build on.

Filter for blanks first: any row missing a topic, intent, stage, canonical, status, or gap reason. Then filter for trouble: duplicate URLs, duplicate topic-plus-intent pairs, contradictory canonicals, orphans, redirects, and anything parked in your exceptions tab.

Check that pages you care about are actually reachable. Record internal inlinks and link depth for each one. Unique inlinks is the number you want, since it counts a linking page once even when it links to the target several times.

Look at your orphan candidates before you touch them. An orphan is a link-architecture condition, not a verdict on value. It can still be indexed, still get external traffic, still matter. Link to it, update it, or retire it on purpose, but decide rather than default.

Then sample. Pull ten rows at random and open the live pages. If the sheet and the page disagree, your classification pass needs another look.

Finish by locking the audit date and assigning owners to every update, consolidation, redirect, removal, and unresolved review. Hand forward only the validated coverage matrix and the open decisions.

One note on AI search, since it is on everyone's mind. Google's guidance on its AI features does not create a separate inventory to maintain. The fundamentals still carry: crawlable pages, important content available as text, structured data that matches what is visible. Do that well and you are eligible. Nobody can promise you a citation.

Done when: the map is internally consistent, every included URL has a disposition, every gap has evidence behind it, and the next planner can see what already exists before proposing anything new.

What to do next

Your inventory is canonicalized, classified, deduplicated, and gap-flagged. That is the hard part of a content audit before cluster planning, and you have done it.

Now the cluster planning can start, and it starts from a much better place. You know which hub already exists, which spokes are written, which pages need updating instead of replacing, and which combinations of topic, intent, and stage genuinely have nothing serving them.

Keep the sheet alive. Update it when pages change materially, and record the date and the owner each time. A stale map sends you right back to guessing.

If the manual version of this is more than your team can carry, DeepSmith maps your site and competitor sites onto one shared topic taxonomy, with funnel stages and coverage gaps built in, and then produces the content that closes the gaps you decide are real. You can start a free trial and see your own map before you pay for anything.

Frequently asked questions

Is a content inventory the same as a content audit?

No. An inventory is the quantitative list of what exists and its characteristics, like URL, title, format, author, dates, and metadata. An audit is the qualitative evaluation on top of it: quality, performance, usefulness, accessibility, and the action each asset needs. Build the inventory first, then audit the rows.

How do I know whether I have a content gap or just a bad existing page?

Search your normalized inventory by topic, intent, and stage before you decide anything. If a distinct page exists but is thin, outdated, poorly linked, or inaccurate, mark it for update or consolidation. Call it a gap only when nothing serves that combination, or when the coverage that exists is materially incomplete.

Do two pages with the same keyword always cannibalize each other?

No. Judge them on topic plus intent, audience, and stage. Two pages can share a keyword and still do different jobs, like a definition, a comparison, and a decision page. Investigate overlap when two pages make the same promise to the same audience at the same stage, whatever words their titles use.

Should I delete orphan or low-traffic pages before planning a cluster?

Not automatically. An orphan can still be indexed, still receive external traffic, still hold useful history, and still support a business need. Record its status, look at its links and its value, then keep, update, link, consolidate, redirect, remove, or flag it for review based on what you find.