DeepSmith

Aug 26 · Content Operations

19 min read

Content Pruning for Ecommerce: Discontinued Products, Out-of-Stock Pages, and Faceted Duplicates

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
A monochrome cover showing a grid of small product-page cards fanning out into three separate stacks with curved redirect arrows, under the centered line Keep, Redirect, or Remove.

You open the crawl report and there are 40,000 URLs. Half of them are filter combinations nobody asked for. A few hundred are products you stopped selling two years ago. And somewhere in there sits a bestseller that is out of stock until next month.

Feeling buried? That is normal, and you are closer than you think. Ecommerce content pruning goes wrong when teams treat all three piles as one. A temporary stockout, a permanently dead product, and a duplicate filter URL each need a different answer.

By the end of this guide you will have a decision for every URL: keep it, redirect it, or remove it. One class of page at a time.

Sort your URLs into three buckets before you touch anything

Here is the rule that saves you from most bad calls: separate temporary availability, permanent discontinuation, and duplicate navigation URLs before you change a single status code.

Those three states look identical in a CMS. They need completely different treatment.

  • Temporarily out of stock. The item is coming back. Keep the URL live, keep it returning a 200, and usually keep it indexable.
  • Permanently discontinued. The item is never coming back. Now you choose: keep a useful archive, redirect to a real successor, or remove it truthfully.
  • Duplicate faceted URL. A filter or sort combination that repeats a product set you already have. Consolidate it or keep it out of search.

Here is the whole framework in one table. If you only read one thing today, read this.

SituationDefault actionKeep in search?
Temporary stockout, item expected backKeep the product URL liveUsually yes
Backorderable or preorderKeep the URL live, explain the waitYes, if the page is useful
Discontinued, close successor existsPermanent redirect to that successorConsolidates into the target
Discontinued, no successor, but real demand or linksKeep a useful archive pageOnly if it still answers a searcher
Discontinued, no traffic, links, or valueReturn 404 or 410No
Facet with distinct inventory and real demandKeep it as a deliberate landing pageYes, on purpose
Facet that duplicates a category or another facetConsolidate to the preferred URLNo
Facet with zero products or infinite valuesDo not generate the URL at allNo

Notice what is missing from that table: a number of days. There is no defensible "after 90 days out of stock, delete it" rule. There is no universal traffic threshold either. Your category, your supplier lead times, your seasonality, and your margins decide that, and you should describe those as your own internal rules rather than an industry law.

Done when: you can name which of the three buckets any URL in your catalogue belongs to without opening a spreadsheet.

Build one URL inventory that covers products, variants, and facets

You cannot prune product pages you have never listed. So the first real job is a single inventory with one row per candidate URL.

Each row needs a few things:

  1. The URL and its page type. Product, variant, category, internal search, filter, sort, or a tracking parameter.
  2. The product or SKU identifier, where there is one.
  3. The lifecycle state. Active, temporary stockout, backorder, preorder, or permanently discontinued.
  4. The value signals. Organic clicks, conversions, assisted revenue, and the recent trend. Backlinks, internal links, reviews, and any genuinely useful product information.
  5. The technical state. HTTP response, canonical target, noindex, robots treatment, sitemap presence, and whether the page still shows up in category pages or internal search.
  6. The commerce state. Product structured data, the availability value, the merchant feed status, and what checkout actually does.
  7. For facets: every parameter, its values, how many combinations exist, the product count, and how much it overlaps the parent page.

No single tool gives you all of that. Your product feed knows the stock state but not the backlink value. Your crawler finds canonical problems but has no idea whether the buying team plans to reorder. Pull from server logs, a crawl, Search Console, analytics, the CMS, the feed, and the XML sitemap together.

Done when: every candidate URL has an owner, a lifecycle state, a value assessment, and a proposed disposition, and the inventory tells a temporary stockout apart from a discontinued product.

Common mistake: starting with a bulk delete or a bulk noindex export. That collapses a temporary inventory condition and a permanent content decision into the same event, and it throws away reviews, links, and rankings the business still needs.

For the content side of that inventory, DeepSmith's Content Map crawls your site, classifies every page onto a shared topic and funnel-stage taxonomy, and rechecks sitemaps every 24 hours. It shows which product and collection content exists and where it is thin. It is a discovery layer, not a stock-of-record: lifecycle and value fields still come from your ecommerce, analytics, and SEO systems.

Ask merchandising one question before you touch a URL

This is the smallest step in the guide and it prevents the most damage.

Ask the merchandising or inventory owner: will this exact sellable item return, be orderable later, or be replaced? Record the answer. Do not infer it from a stock count of zero.

Then file the product into one of six states:

  1. In stock. Normal handling.
  2. Temporarily unavailable, expected back. Keep the URL live and useful.
  3. Backorderable. The customer can order now for later delivery. Explain the wait.
  4. Preorder. Not available yet, but there is a real release date coming.
  5. Permanently discontinued. No planned replenishment. Move to the discontinued decision tree.
  6. Unknown. Do not silently delete. Escalate, and use a truthful temporary state until the business decides.

That last state matters more than it looks. "Unknown" is where most bad deletions come from.

Done when: the state is documented by the business, the page copy matches it, and your feed and schema can express it. A product marked discontinued is not still going out as an ordinary out-of-stock offer.

Where people go wrong: using "out of stock" as a catch-all for maintenance, holiday closures, products they no longer want to advertise, and items that will never return. Google's Merchant Center guidance treats a temporarily unavailable product and a permanently absent one differently, and it expects your landing page and your feed to agree.

Choose keep, redirect, or remove for discontinued product pages

For every permanently dead product, work through this order.

Keep a useful 200 page when the URL still answers a need

Keep it live when it has meaningful search demand, external links, reviews, model recognition, or genuinely useful historical product information.

Then make the status unmistakable. Say the item is discontinued and will not return. Strip the add-to-cart path, any stock messaging that implies a comeback, and active merchandising. Point the shopper somewhere real: a replacement, a relevant alternative, or the parent category.

A useful archive is not a fake store listing. It answers what the product was, who it suited, what replaced it, and where to go next. If the page helps customers but has no reason to appear in search, a crawlable noindex is a fair option.

Redirect permanently when there is a close, relevant replacement

Redirect only when the target preserves the old searcher's intent. Good targets look like this:

  • The direct successor model.
  • A replacement with the same primary use case and product type.
  • A stable equivalent in a similar category and price band, when no model successor exists.
  • A tightly relevant parent category, when there is no equivalent product but the category genuinely answers the old need.

A server-side permanent redirect is the strongest URL consolidation signal you have. Point it at the final target, kill any chains, and update internal links and sitemaps so they lead to the destination rather than through the old URL.

A redirect to an unrelated product can be treated as a poor substitute or a soft 404. "Any in-stock item" is not a relevance test.

Return a 404 or 410 when nothing is left

Use a 404 or 410 when the item is permanently gone and the URL has no meaningful traffic, links, customer value, or relevant successor.

A 410 says the removal was intentional and permanent. A 404 fits a resource that is missing or could come back. Industry guidance calls 410 potentially faster for removal, but nobody has shown a ranking benefit. Pick the code that tells the truth, not the one you hope is magic.

A good custom not-found page can recommend real categories and products. It should never suggest the old product is still available.

Clean up the old product's footprint

The URL is only half the job. For a permanently discontinued product, remove or update:

  • Product and shopping feeds.
  • XML sitemap entries, if the URL is no longer your chosen canonical resource.
  • Category, collection, recommendation, and internal-search placements.
  • Promotional modules and paid landing-page references.
  • Internal links that still present the product as buyable.
  • Product structured data claiming a purchasable offer that does not exist.

Common mistake: redirecting every discontinued product to the homepage, or to whatever the newest product happens to be. The destination has to represent the old page's intent. If nothing does, keep a useful archive or return a truthful not-found response instead.

Keep temporary out-of-stock pages live and honest

Good out of stock page SEO starts with a decision you have already made: the item is coming back, so the URL stays. Keep it live and returning a 200, then tell the customer what is happening right next to the buy button, not buried in a feed or hidden in structured data.

When each of these is true and your operations support it, the page should carry:

  • A clear "out of stock" or "sold out" message.
  • A realistic restock or delivery expectation. Never invent a date.
  • A back-in-stock email option or a save-to-list.
  • A genuinely comparable alternative, not an accessory.
  • A backorder or preorder option, if you accept that order state.
  • A disabled or greyed-out purchase control when the item cannot be ordered.
  • Customer service or delivery information where the uncertainty is material.

Baymard's ecommerce UX research makes a point worth sitting with: if you can accept an order for later delivery, a stockout should not force the customer to leave. Explain the longer wait close to the buy section and show prominent alternatives. A notification form is not the same as letting someone buy.

Pro tip: every surface has to agree. Page text, product structured data, the merchant feed, cart behaviour, checkout, and regional shipping settings should all describe the same state.

Merchant Center supports in_stock, out_of_stock, preorder, and backorder for ordinary products. Use out_of_stock when the item is temporarily unavailable and you are not taking orders. Use backorder or preorder when you genuinely accept that order type. Do not use out_of_stock for maintenance, a holiday, or a product you would rather not show.

The way to keep that in sync is to derive page messaging, structured data, feed state, and purchase controls from one lifecycle source, with separate business states for temporarily out of stock, backorder, preorder, and discontinued.

Done when: a customer can tell what happened and what they can do next, and a crawler, a structured data test, a feed review, and a real checkout test all see the same state.

Where people go wrong: pulling the product off the site on day one of a stockout. That breaks a stable URL, discards the signals it accumulated, and cuts off the restock path. The opposite mistake is just as common: leaving a discontinued product marked "out of stock" for three years, which sets a false expectation and pollutes your product data.

Control faceted navigation duplicates with an allowlist

Faceted navigation duplicates are a URL generation problem first and an indexing problem second. Fix the generation and the indexing problem shrinks.

For each parameter or combination, ask six questions:

  1. Does it expose a materially distinct set of products?
  2. Is there real search demand or a customer task behind it?
  3. Does the page have enough inventory and stable content to satisfy that intent?
  4. Is the combination durable, or does it vanish when a session, price range, or stock level changes?
  5. Does it duplicate a category, another facet, or a sort order?
  6. Can it be linked, described, canonicalized, and maintained on purpose?

A useful facet is a deliberate landing page. It is never a by-product of every click.

Build the allowlist small

Start with a handful of facet types that create unique inventory and clear intent, like a stable brand or category collection where your data supports it. Google's faceted navigation guidance treats parameters like item, category, name, and brand identifiers as potentially valuable. Session IDs, arbitrary user-generated values, and most price-range refinements usually are not.

Give each chosen page a descriptive title, useful content, a consistent URL pattern, a self-referential canonical, and intentional internal links. Add it to the sitemap only if it really is a preferred indexable page.

Consolidate the rest deliberately

For a duplicate or low-value facet shoppers still need to reach, point a relevant canonical at the preferred page, or use a crawlable noindex when the page should never appear in search.

Two rules keep people out of trouble here.

Canonical and noindex are not synonyms. Canonical says which duplicate is preferred and can consolidate signals toward it. Noindex says an accessible page should not appear in results. Pick based on whether you want signals to merge or the page to disappear.

And robots.txt is not a deindexing tool. Google has to crawl a page to see a noindex tag or header. Block it and the instruction never gets read, and the URL can still show up. Robots rules are for crawl control on known parameter traps, nothing more.

Stop creating combinations you never wanted

Do not expose crawlable links for filters that return zero products. Offer only valid refinements, grey out or suppress unavailable combinations in the interface, and never link internally to an empty page.

Use conventional parameter syntax too. A stable key-value format joined with ampersands is easier to reason about than a bespoke encoding, and session information should stay out of crawlable paths wherever you can manage it.

Done when: you have a written facet allowlist and a consolidation policy. Every indexable facet has distinct inventory and intent, every duplicate has a preferred target or a crawlable noindex, and zero-result combinations are not generated as links.

Common mistake: one blanket rule. "Noindex every filter" deletes valuable brand and category landing pages from search. "Index every filter" creates an unbounded duplicate space. The right policy is selective and based on your own evidence.

Every candidate URL leaves one undifferentiated grid and fans into three lanes with three different dispositions: temporarily out of stock pages keep the URL live, permanently discontinued pages get a redirect, an archive, or removal, and duplicate faceted URLs converge on a single consolidate or keep out of search outcome.

Ship the signals together, not one at a time

Ecommerce content pruning fails when the signals live in separate systems and get changed on separate days. Here is what has to move as one release.

Permanent redirects. Use one when a URL has genuinely been replaced or consolidated. Point straight at the final target. Remove chains, and update internal links, canonicals, and sitemaps to the final URL.

Temporary redirects. Use one only when the original URL is really coming back. It does not carry the same canonical implication.

Canonical tags. Put them in the HTML head with an absolute target and no fragment, and make sure the target represents the content. Self-canonicalize your preferred pages. Google can still choose a different canonical, so inspect what it picked after you deploy.

Noindex. Use a meta noindex or X-Robots-Tag when a page stays accessible but should leave search. Keep it crawlable so the instruction gets seen.

Sitemaps and internal links. Include only preferred canonical URLs. Remove deleted, redirected, and non-preferred facet URLs, and repoint internal links to the chosen canonical or successor.

Schema and feeds. Availability should reflect what a customer can actually buy. If a product is permanently discontinued, take it out of shopping data rather than disguising it as a temporary stockout.

Variants. Give each meaningful variant a unique identifier such as a SKU or GTIN. If it has its own URL, that page should render its own image, price, availability, and add-to-cart state. Do not collapse a real variant into its parent just because a parameter is involved, and do not build indexable pages for cosmetic or session-only selections. If price or availability only appears after JavaScript runs, crawlers may miss the current state.

One platform note: Shopify makes the product unavailable through its sales channels or moves it to Draft, then you verify the old address returns a broken-page response, then you create the redirect in the admin. Shopify will not create a redirect from a URL that still resolves. So the order is: change the product state, verify the response, create the redirect, test the destination.

Where people go wrong: fixing the HTML tag and leaving the old URL sitting in the feed, the sitemap, the navigation, and a live paid campaign.

QA the changes and watch what happens next

Test a representative sample from every decision class, then monitor the whole set.

For discontinued product pages, check the old response code and, if redirected, the direct destination. Confirm the destination is relevant and there are no chains, that the URL is gone from feeds, sitemaps, navigation, and internal search, and that archived pages say discontinued, carry no purchase actions, and still render useful reviews and alternatives.

For temporary stockouts, check that stock status is visible near the purchase control, that restock and backorder wording is truthful with no invented dates, and that the notification and alternative-product paths work. Then test structured data, feed state, cart, checkout, and regional availability.

For facets, check the preferred URL and canonical target, that noindex is visible to crawlers where you used it, that robots rules are doing crawl control only, and that sitemaps and internal links contain only intended indexable facets. Test the weird combinations too: zero-result, sort, session, and tracking parameters.

Then watch the business, not just the crawler. Organic clicks, indexed page status, crawl errors, feed warnings, product conversions, alternative-product clicks, notification signups, and customer service complaints all tell you something. Use the trend to revisit decisions rather than to set a permanent delete threshold.

For noindex changes, request a recrawl and inspect the URL after Google has processed it. For redirects, check both ends.

Common mistake: declaring victory when a page disappears from a crawler report, without testing Google's chosen canonical, the merchant feed, mobile UX, or the checkout path. Technical deletion and a good customer outcome are not the same measurement.

What to do next

You do not need to prune product pages all at once. Pick one class, ship it, then move to the next.

Turn the inventory into a recurring lifecycle review rather than a one-off cleanup. Give stock state, SEO signals, and feed consistency separate owners, because those three break in different ways and on different schedules. Set a review trigger for temporary stockouts and for the high-value archive pages you kept.

The content side of this work compounds. Once you know which categories, replacements, and buying guides your catalogue needs, you still have to write them. DeepSmith stores your company, product, persona, voice, and content type context in Deep IQ, then Content Studio's Writer produces publish-ready articles from it with SEO and AEO structure, internal links, and metadata built in during creation. It will not deploy your redirects or fix your feed. It handles the content your pruning decisions expose.

Want to see it on your own catalogue? Start a free trial and give it a week.

Frequently asked questions

Should I delete a product page when it is temporarily out of stock?

No. If the item is coming back or can be ordered for later delivery, keep the URL live and returning a 200. Show the stock status clearly, give a realistic restock expectation, offer a back-in-stock alert, and show comparable alternatives. Good out of stock page SEO protects the links, reviews, and rankings the URL has accumulated, and none of that comes back automatically when the stock does.

Should a discontinued product redirect to the homepage?

No. Redirect only when the destination preserves the old searcher's intent: a direct successor model, a product with the same use case, a close equivalent, or a tightly relevant parent category. A redirect to the homepage or an unrelated in-stock item can be treated as a poor substitute or a soft 404. When nothing relevant exists, keep a useful archive page or return a truthful 404 or 410.

Is noindex or robots.txt better for duplicate faceted URLs?

Noindex, if your goal is keeping pages out of search. Google has to crawl a page to see a noindex tag or header, so blocking it in robots.txt means the instruction is never read and the URL can still surface. Use robots rules for deliberate crawl control on parameter traps. For faceted navigation duplicates whose signals should consolidate, use a relevant canonical instead, and a crawlable noindex when the page should stay reachable for shoppers but leave search.

Do product variants need separate URLs?

Only when the variant is genuinely a different sellable thing. A variant with its own identifier and a distinct image, price, availability, or add-to-cart state deserves an accurate page and should not be collapsed into its parent just because a parameter is in the URL. A cosmetic or session-only selection with no distinct search value does not need its own indexable page.