DeepSmith

Aug 26 · Content Operations

19 min read

How to Run a Content Decay Audit at Enterprise Scale: Ownership, Governance, and Thousands of Pages

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
A monochrome cover showing a large grid of dark page cards in six clusters, a handful picked out in lighter grey as flagged pages, with thin connector lines running up to a single node above, under the white cover line Auditing a Content Estate at Scale.

You have tens of thousands of pages, a dozen teams publishing into different systems, and nobody who can say what most of it is doing. An enterprise content audit at that size rarely breaks on the data. It breaks because "decay," "owner," and "done" mean something different on every team. Work through the eight steps below and you can audit thousands of pages with a queue you can defend, one accountable owner per decision, and a review rhythm that keeps running after the first wave closes.

If that feels like a lot, take a breath. You don't start with the spreadsheet. You start with one page of definitions.

What you'll need: Search Console, your analytics, a crawler, somewhere durable to keep the data, and a named owner for every content domain in scope.

Step 1: Write the audit charter before you export a single row

A content governance audit starts with decision rights, not with data. Write a one-page charter first.

Put these in it: the business objective, the domains and systems covered, the asset types, the time window, the exclusions, the completion date, the decision categories, the risk rules, and the escalation path. Name one central governance lead, and one steward for every domain you are including.

Then write the glossary. Define page, asset, canonical URL, owner, steward, last reviewed, baseline period, decay flag, technical exception, business value, and completed. Decide what the unit of work is, and how you will handle pages with no traffic, pages behind a login, non-HTML assets, redirects, and regional duplicates.

You know it's done when the charter has a scope list, an owner map, definitions, a RACI, a decision vocabulary, and an escalation route. Every included domain has a steward before anyone pulls data.

Where this usually breaks: a central team exports every URL and asks the business units to "review when you can." Nobody set decision rights. Every team reads decay differently, and the queue turns into a report instead of an operating process.

Common mistake: treating the CMS author as the accountable owner. Authors, editors, legal reviewers, product SMEs, publishers, and business owners are often different people. Enterprise content ownership describes who can decide a page's future, not who last touched it.

Step 2: Build one canonical inventory, not a list of URLs

A content inventory answers "what do we have?" An audit answers "what is it doing, what shape is it in, and what happens next?" You need the first before you can do the second.

Key the inventory to a stable content identity, because the URL string alone is not enough. Normalize protocol, host, trailing slashes, tracking parameters, redirects, canonicals, and locale variants. Keep the observed URL and the resolved canonical URL as separate fields so migrations stay auditable later.

Capture seven groups of fields:

  • Identity: stable content ID, observed URL, canonical URL, host, property, locale, CMS record ID.
  • Content: title, content type, topic, funnel stage, audience, author, owner, steward, source system.
  • Lifecycle: publish date, last modified, last reviewed, next review date, workflow state.
  • Search and technical: indexability, canonical target, HTTP status, redirect target, robots state, sitemap presence, links, schema status.
  • Performance: baseline and current clicks and impressions, CTR, average position, engagement rate, key events.
  • Business and risk: business goal, value tier, conversion role, revenue association, legal and accessibility risk.
  • Audit: decay flag, diagnosis, recommended action, priority, accountable owner, reviewers, due date, exception reason.

Keep raw extracts separate from the curated inventory, and store the extraction date, source, date range, and transformation version. That is what lets you reproduce a result six months from now. Inventory everything that serves a user too, not just blog HTML: documentation, product pages, PDFs, videos, forms.

You know it's done when every in-scope asset has one canonical identity or a written exclusion reason, every row has a steward, and you can trace any row back to its source.

Where this usually breaks: teams merge rows by title, count URL variants as separate pages, skip the PDFs and the docs, or overwrite last-modified with the date of the export. Each one inflates the estate and sends ownership to the wrong person.

Step 3: Cut the estate into owned audit waves

Here is the shift that makes content decay at scale manageable. You don't make a giant spreadsheet smaller. You make waves that are independently owned and still comparable to each other.

Partition the inventory before you score anything. Use dimensions that change the decision or the evidence you will need: business unit, region or brand; content type and source system; topic and funnel stage; business value and risk; lifecycle state; and governance status such as a missing owner.

Then order the waves. Start with the content people use most, the pages carrying real business goals, the regulated material, and the domains where ownership is already clear enough to act on. Run one pilot wave first to test your definitions and your data joins.

Give each wave a fixed population, a steward, decision owners, dates, expected evidence, and exit criteria. Freeze the population snapshot when the wave opens, and record pages published or changed mid-wave separately.

You know it's done when every in-scope row sits in exactly one wave or has a documented exclusion, and no page is quietly duplicated across two teams.

Where this usually breaks: the central team ranks all pages globally with no business context, or each unit invents its own fields and thresholds. Neither result can be compared, and your highest-value content ends up waiting behind low-value volume.

Pro tip: publish the rubric before you publish the rankings. Ask each steward to classify the same small sample, compare where they disagree, and fix the glossary before the full wave runs. You will catch four different readings of "owner" while the correction is still cheap.

Step 4: Freeze a baseline and collect evidence at scale

Pick comparison periods that match the content and the business cycle. A practical default is the most recent three months against the same three months a year earlier. Shorten or lengthen only when seasonality or publishing volume forces it, and record the exact dates, the timezone, the data state, and why.

Match each source to the question it can actually answer:

  • Search Console gives clicks, impressions, CTR, and average position by page, query, country, and device. The API's row limit defaults to 1,000 and can be set as high as 25,000 per request, but Google is clear that internal limits mean it does not guarantee every row. It generally returns top rows, so pagination is not proof you have everything. CTR comes back as a value between 0 and 1.
  • The Search Console BigQuery export is the stronger source for a large estate. It runs daily, is not subject to the daily row limit, and gives you URL-level and property-level tables, minus anonymized queries. Only property owners can set it up, the first export can take up to 48 hours, and if you set partition expiration Google's guidance is at least 14 days. Search Console keeps 16 months of history, so the warehouse is where you go longer.
  • GA4 gives landing page, users, engagement rate, and key events. An engaged session lasts longer than 10 seconds, includes a key event, or includes at least two page or screen views. Use the same property, date range, consent assumptions, and event definitions on both sides.
  • A crawler gives status codes, canonicals, robots directives, indexability, metadata, links, and duplicates. Plan its capacity honestly. One widely used crawler documents roughly two million URLs at 4 GB of RAM, five million at 8 GB, and around ten million with 16 GB and a large SSD. Those are tool-specific planning figures, not a promise about your site.
  • Business systems give conversions, support demand, and revenue association. Document the join rather than inferring it from page traffic.

For a warehouse workflow, land the raw data, normalize URLs, deduplicate by content identity and date, then join to the inventory. Keep page-level and query-level facts apart, and add data-quality tests for row counts, date coverage, null canonicals, duplicate keys, and missing owners.

This is also where a shared taxonomy saves you weeks of arguing. DeepSmith's Content Map crawls and enriches every page on your site, classifies it onto a granular topic and a funnel stage of Awareness, Consideration, or Decision, and re-checks your sitemaps every 24 hours so new pages fold in on their own. It maps competitor sites onto the same taxonomy, which gives your segmentation a like-for-like reference. Use it for topic and coverage context, not as your canonical inventory. Owner, canonical identity, and performance joins still live in your own data model.

You know it's done when every wave has a frozen baseline, source extracts, documented joins, and a written list of data gaps. An analyst can rebuild any page's numbers without emailing a team for an export.

Where this usually breaks: someone mixes Search Console and GA4 date definitions, treats missing data as zero, or downloads the top rows from a limited interface and calls it the estate.

Step 5: Flag decay, then rule out the false positives

Content decay is a gradual loss of organic traffic or search visibility after a page has peaked. A page is not decayed because it's old, because it's quiet, or because it had one bad week.

Calculate change at page level and keep the underlying values next to it: absolute click change, percentage click change with a separate state for a tiny baseline, impressions and CTR change, average position change with the direction written down, engagement and conversion change, and time since the last meaningful content change.

Use several signals together instead of one threshold. Some published heuristics can seed your policy: a traffic drop of more than 20 percent over 90 days, a drop of more than 20 percent year over year, or a loss of more than five positions on target keywords. Those are triage signals, not Google rules and not proof. Calibrate them on your pilot wave.

Then read the pattern before you name a cause:

What you seeWorking interpretationWhat to check next
Clicks and impressions both downVisibility and demand may both be downSeasonality, indexability, rankings, sitewide events
Impressions down, CTR upThe page may have lost positions or eligible visibilityAverage position, query mix, competitors
Impressions flat, CTR downClicks may be going to the title, snippet, or a SERP featureSearch appearance, query mix, title history
Position down, demand stableMore likely a ranking or relevance problemIntent, competitors, links, cannibalization
GA4 down, Search Console stableMeasurement or landing-page behavior may be involvedTagging, consent, redirects, attribution
AI citations down, organic stableAnswer-source preference has shiftedTracked prompts, cited competitor pages, answer history

Before you blame the content, check seasonality, demand, algorithm and SERP changes, migrations, template changes, tagging and consent changes, indexation, robots and canonical changes, redirects, internal linking, and cannibalization. Look at whether the page changed shortly before the decline. A drop right after an edit is degradation, not decay.

That last row in the table deserves its own workstream. A page can hold its rankings and still stop being cited in AI answers, and an AI answer can satisfy someone without ever sending a click. Keep two reporting layers: an organic layer, and an AI visibility layer measured on its own terms. DeepSmith's AI Visibility module tracks mention rate, citation rate, share of voice, sentiment, and trend against the prompts you choose, and it shows which of your pages get cited and which competitor pages win the ones you lose. Engine coverage rises by plan, from ChatGPT on Pro up to all ten engines on Enterprise. Report that layer beside organic decay, never blended into it.

The DeepSmith AI Visibility overview reports mention rate, citation rate and share of voice as three separate top-line metrics, with a per-engine breakdown and a competitor leaderboard, so answer-engine visibility stays a layer of its own instead of folding into the organic numbers.

You know it's done when every flagged page carries its periods, source values, change calculations, diagnostic pattern, false-positive checks, and an evidence note.

Where this usually breaks: an analyst sorts by percentage loss, floats a pile of tiny-baseline pages to the top, and treats every decline as a writing problem.

Step 6: Score the queue by value, not by percentage drop

Separate three things that keep getting mashed together: urgency, value, and effort or risk.

A transparent internal model works better than a black box. One workable shape:

priority = impact x confidence x business value x risk multiplier / estimated effort

That is an operating device, not an industry standard. Define each factor on a small ordinal scale and write the rubric down. Impact is the size and persistence of the loss, with absolute values kept beside percentages. Confidence is how many independent signals agree. Business value covers conversion role, revenue association, strategic topic, and reach. The risk multiplier covers legal, regulatory, accessibility, and reputational exposure. Effort is the work it will take across analyst, SME, engineering, legal, and publishing.

Sort the result into four tiers:

  • P0, immediate investigation: severe risk, misleading information, a major business page, or a likely technical incident.
  • P1, high-value recovery: material loss, strong evidence, real business value.
  • P2, planned review: credible decay with moderate value, or a page that needs SME review first.
  • P3, monitor: weak evidence, low value, insufficient baseline, or a seasonal explanation.

Attach the reason and the evidence to every item, plus the decision owner, the proposed action, the reviewers, the due date, and what would change the priority.

Think about three pages. Page A has a large absolute click loss, strong business value, and high confidence. Page B has a bigger percentage loss but a tiny baseline and no business goal. Page C holds its organic visibility while its AI citation rate falls. Page A goes first, Page B goes to monitor, and Page C goes to the AI visibility workstream. Sorting by percentage alone would have put them in exactly the wrong order.

You know it's done when a leader can look at the queue and explain why the first page is ahead of the second.

Where this usually breaks: someone publishes a "top 1,000 declining URLs" list with no business value attached, and treats the score as objective truth. A score should speed up judgment, not stand in for it.

Step 7: Route every page to one accountable owner

Use a finite decision vocabulary. Seven outcomes cover almost everything:

DecisionUse whenRecord
Keep and monitorEvidence is stable or the decline is explainedReason, review date, owner
Update or refreshStill strategically useful, but inaccurate, thin, or behind competitorsOwner, SME, scope, approvals, due date
ConsolidateSeveral pages overlap and a stronger page should absorb themSurviving page, source pages, redirect plan
RedirectAn old URL holds equity or users and has a genuinely relevant destinationSource, destination, rationale, validation
Archive or removeNo continuing user or business purposeRisk check, approval, replacement or none
Investigate technical causeThe loss may sit in indexability, templates, migration, or analyticsTechnical owner, task, evidence deadline
Monitor or deferEvidence is thin, seasonal, or below thresholdReason, next observation period, owner

One page gets one accountable decision owner. Others can be responsible for evidence or implementation, and consulted for accuracy, legal, accessibility, or regional review. Shared ownership with nobody accountable is not governance, it's an escalation gap. Record a decision even when the decision is "keep," because a blank status isn't neutral, it's ungoverned.

Run assignments through a review queue or workflow board, but keep the canonical audit table and the decision history somewhere durable. The queue needs status, owner, reviewer, action, priority, due date, dependency, evidence link, exception reason, and approval state.

When the decision is "update" or "produce a replacement," that is a production handoff. The work stops being governance and starts being writing. DeepSmith's Deep IQ holds your company, product, persona, brand voice, and content type context as structured records, and Content Studio takes an approved idea through research, internal and external linking, metadata, and a cover image to a publish-ready article. Your owner still makes the editorial call and the risk approval. The tool carries the production load, not the accountability.

You know it's done when every flagged page has one accountable owner, an action or a documented monitor state, reviewers where the rubric calls for them, a due date, and an escalation path.

Where this usually breaks: ownership gets assigned after the recommendations are already written, one central editor is expected to approve everything, or "update" becomes a vague bucket with no scope, reviewer, or due date.

Step 8: Report, sign off, and set the next review date

Close every wave with two reports. The executive report covers estate coverage, owner coverage, flagged pages, risk and value tiers, decision distribution, backlog aging, and the trend by domain. The operating report covers the page-level queue, the evidence, the decisions, the reviewers, the due dates, the exceptions, and anything that changed mid-wave.

Quality-check before you present. Reconcile inventory counts against the source systems. Confirm every page sits in one wave, every flagged item has an owner, every percentage has a real or explicitly low baseline, every redirect points somewhere relevant, and every removal has an approval. The domain steward signs off on the population and the decision distribution, and the governance lead signs off on method and data quality.

Then set the rhythm. A quarterly review is a reasonable starting point for broad accuracy and relevance checks, more often for high-risk content, less often for stable low-risk material. Some editorial teams flag material year-over-year loss quarterly and review their highest-value articles annually. Write your cadence into the charter and adjust it from observed volume and capacity.

Keep a change log of glossary revisions, threshold changes, ownership changes, and exceptions. Measure the process itself, not just the pages: inventory completeness, owner coverage, time to decision, share of pages with a next review date, and false-positive rate.

One caution worth holding onto. A content governance audit is not permission to mass-rewrite. Google's guidance on people-first content asks for original, substantial work made to help readers, and its scaled content abuse policy applies when pages are produced mainly to move rankings, whether a person or a machine made them. Google also says the same foundational SEO practices apply to AI Overviews and AI Mode, with no special schema required, though a page has to be indexed and eligible for a snippet to be used as a supporting link. Changing a date without improving anything is not an audit outcome.

You know it's done when leadership sees a decision and risk view, teams see an actionable queue, the steward has signed off, and the next review date is on the calendar.

Where this usually breaks: the enterprise content audit ends when the spreadsheet is delivered. No owner coverage metric, no exception log, no next review date, and no route for new or migrated content. A one-time cleanup is not governance.

A flow diagram showing central standards feeding two parallel evidence lanes, an organic layer and an AI visibility layer, which stay separate until they meet in one prioritized queue that routes each page to one accountable owner, with a return line reading next review date reopens the cycle running back to the standards.

What to do next

Don't try to launch all eight steps across the whole estate at once. Pick one domain where ownership is already clear. Run the charter, the inventory, and a single pilot wave. Calibrate your thresholds against what you find, publish the owner map, and schedule the next review before you close the current queue.

That's the whole trick with content decay at scale. You are not trying to finish. You are trying to build something that keeps going without you.

If part of your queue turns into "update this page" and your team is the bottleneck on production, that's the piece DeepSmith is built for. Track how AI answers describe your brand, see which pages earn citations, and produce the replacements from stored brand context. You can start a free 7-day trial and see real data before you pay.

Frequently asked questions

How do we audit thousands of pages without reviewing every one by hand?

Automate the inventory, the URL normalization, the data collection, the joins, and the flagging. Use owned waves and value tiers to aim human attention. Every page still needs a disposition, even if that disposition is "monitor" or "excluded," but not every page needs the same depth of review.

What's the right threshold for content decay?

There isn't a universal one. A drop of more than 20 percent over 90 days or year over year, and a loss of more than five target-keyword positions, are useful starting heuristics. Calibrate them on a pilot, keep absolute values beside percentages, and combine them with impressions, CTR, business value, and seasonality.

Who should own a decaying page?

The person or team with authority to decide its future. Authors, editors, SMEs, SEO analysts, legal reviewers, and publishers can each be responsible or consulted for parts of the work. Enterprise content ownership means one accountable decision owner per page, with the implementation owner recorded separately.

Should we score organic rankings and AI citations together?

No. Report them as related but separate layers. Organic traffic can hold steady while AI citations fall, and an AI answer can satisfy a reader without a click. Use the AI layer to spot answer visibility and competitor citation patterns, not to redefine organic decay.