You can generate three hundred titles in an afternoon. Turning them into three hundred pages that each deserve to exist is the hard part, and it is where most teams quietly go thin. This guide gives you a repeatable way to scale content clusters so every planned URL has its own job, its own evidence, and its own reason to be selected by a reader or an answer engine. By the end you will have a page ledger and a distinct-topic audit you can run again next quarter.
Here is the rule the whole method rests on. Do not multiply keywords into pages. Multiply distinct questions, decisions, and evidence-backed jobs.
A URL earns a place in a hundreds of pages topic cluster only when all seven of these are true:
- It belongs to one clearly defined central entity.
- It sits on a meaningful axis value, such as a different use case, audience decision, lifecycle stage, component, context, or outcome.
- It answers a primary intent that no other planned URL answers better.
- Its page type matches what the target results page is asking for.
- It has its own evidence bundle, not a swapped noun.
- It has a clear relationship to the hub and to its neighbours.
- It can be written with original analysis and attributable sources.
Miss one of those and you have a candidate, not a page. Let's walk through how to get there.
Step 1: Define one central entity and draw the boundary
Start by writing your subject as an entity, not as a keyword.
An entity is the underlying thing: a product, a concept, a role, a process. A keyword is one way somebody asks about it. "Payroll software" can be the entity. "Best payroll software 2026" is a question about that entity. Entity planning asks what the thing is, what relates to it, what people do with it, and what decisions surround it. Keyword planning just asks what people typed.
Write three things down before you enumerate a single page:
- The central entity, in plain words.
- The reader, the business problem, and the outcome the map serves.
- A one-sentence boundary saying what belongs inside this map and what belongs in a neighbouring one.
Use one central entity per map. If your proposed cluster quietly contains two different products or two different audiences, split it into two maps now. It is much cheaper than untangling it at page 180.
How you know it is done: your team can name the entity without reciting a keyword list, and every candidate connects back to it in one sentence.
Where people go wrong: they start with a number. "We need 500 URLs" is a production target, not a coverage definition, and a page count target will happily invent pages to hit itself.
Step 2: Enumerate the axes that change the answer
An axis is a dimension that changes the question, the answer, the evidence, the decision, or the output. Axes are how a cluster grows honestly.
Build an inventory around your entity. These are candidates to consider, not a checklist to complete:
| Candidate axis | It is valid when it changes |
|---|---|
| Audience or role | The reader's inputs, constraints, or decision |
| Use case or job | What the reader is trying to accomplish |
| Lifecycle or funnel stage | The question and the next action |
| Industry or operating context | The evidence, terminology, or examples |
| Geography or regulation | Local facts, legal context, or availability |
| Component or sub-entity | The object being analysed or operated |
| Mechanism or integration | The workflow, the input, or the output |
| Outcome or metric | The success decision and its proof |
| Alternative or comparison | The choice between options that are not interchangeable |
Here is the part that matters. An axis value is not valid just because it can be typed after a keyword. "For marketers," "for agencies," and "for SaaS" are three phrases, not three pages, unless the workflow, constraints, examples, or decision genuinely change.
For each axis, record the allowed values, the question each value creates, the evidence that would support it, and the page role it would produce. Do not generate the full grid of every axis crossed with every other axis. Ask first whether that combination shows up in a real buyer question, has a distinct task, and can be supported with distinct evidence.
Pro tip: run the answer-swap test early. Take one candidate's planned answer and drop it under a neighbouring candidate's title. If it still satisfies that second reader, the axis has not earned a new URL yet.
This is also the moment to check what you already have. DeepSmith's Content Map crawls your site and your competitors' sites into one shared topic taxonomy, classifies every page onto a granular topic and a funnel stage, and reports coverage gaps and untapped topics with per-topic page counts. You can see whether an axis is genuinely missing or whether a page already does that job. Sitemaps are rechecked every 24 hours, so pages you published last week fold in without a manual re-import.

How you know it is done: every candidate has an entity, an axis and value, a proposed reader job, and a reason the axis changes the answer. Anything with only a new modifier gets marked as a query variant.
Where people go wrong: they treat every persona label, filter, location string, and year as an indexable page. That gets you a big number fast and a deduplication nightmare later.
Step 3: Build the candidate inventory from demand and evidence
Now pull candidates. Not from keyword expansion alone, because keyword expansion only tells you what words exist.
Pull from all of these:
- Live results pages for the entity and its core questions.
- Competitor sitemaps and pages, used to find gaps rather than to copy wording.
- Entity extraction or semantic research tools.
- Sales calls, support tickets, interviews, and community threads.
- Your own existing pages, analytics, and search-console queries.
- Buyer-stage questions across awareness, consideration, and decision.
For each candidate, write down the primary query, the underlying intent, the entity and axis value, the likely page type, the unique reader job, the parent hub, and the evidence the page will need. Keep close query variants as supporting terms under one candidate until an audit says they deserve separating.
Then look at the top three results for the target query before you assign a page type. If those results are consistently how-to guides, a product page is a bad starting assumption. If they are comparisons or definitions or category pages, your planned format needs to match, or you need to write down why you are deliberately going against it.
Every candidate also needs a first pass at an evidence path: the facts that will be unique to this page, the source types that can support them, and the first-hand experience or analysis you can add. If a row has no evidence path, park it. That single habit prevents most thinness.
If you want the candidates to arrive with reasons attached, DeepSmith's Opportunity Agents read your AI visibility data and your Content Map gaps and return ideas with the data point that justified each one. You choose a 30, 90, or 180 day window, how many ideas you want, and any extra instructions. The idea shows up defensible instead of brainstormed. Treat that output as inventory input, not as permission to skip the audit coming next.
How you know it is done: every row has a demand signal, an intent, a page type, a unique job, a parent, and at least one plausible evidence path.
Where people go wrong: they use search volume as the only eligibility test, treat a competitor's outline as the required outline, or approve the page first and go looking for evidence afterwards.
Step 4: Audit every candidate for a distinct topic
This is the step that decides whether you get distinct pages at scale or a pile of near-duplicates. Teams that scale content clusters successfully run it before any large-scale drafting, not after.
Group candidates by intent first. Then compare every pair that shares an entity and sits on a nearby axis value. Six binary questions are usually enough:
- Different primary job? Yes or no.
- Different axis value that changes the answer? Yes or no.
- Different expected page type or intent? Yes or no.
- Different evidence or analysis? Yes or no.
- Different output or decision? Yes or no.
- Different role in the cluster? Yes or no.
All "no" answers means this is not a new page. Meaningful "yes" answers still need the page-type and source check.
Notice what is missing: a similarity percentage. There is no reliable universal number of different words that makes two pages distinct, and inventing one will not make your audit more rigorous. Two pages can share a lot of vocabulary and stay genuinely distinct. Two pages can use completely different wording and still be duplicates.
Search results are a strong separation signal here. If many of the same URLs and the same page types rank for both queries, start with one page. If the results show different intents and different formats, look at splitting. Results shift by location, device, and time, so treat this as evidence rather than law.
Give every candidate an explicit disposition:
- Keep: distinct intent, axis, evidence, and role.
- Fold: it is a subquestion the parent page can answer completely.
- Reframe: the entity is useful but the job is too close to something else. Change the job, not just the title.
- Park: no distinct evidence, no clear need, or no suitable page type yet.
- Consolidate: similar pages already exist, so pick the strongest destination rather than adding another URL.
That last one deserves a word. Canonicalization is deduplication, a way for Google to show one version of duplicate or very similar pages. It is a repair tool for URLs that already exist. It is not a permit to create a set of pages that should never have been separate.

How you know it is done: every row carries a disposition, and no two "keep" rows are separated only by a changed keyword or a swapped label.
Where people go wrong: they audit after publishing. By then the duplicates are live, the internal links are wired, and the fix costs ten times what the audit would have.
Step 5: Give every page a role in the hub and spoke map
A cluster is not a spreadsheet of related keywords. The spreadsheet is an input. The cluster is the page-level architecture: what each page targets, how it relates to the central subject, and how the pages link.
Start with hub and spoke. The hub covers the broad topic and routes readers outward, each spoke links back to the hub, and relevant spokes cross-link to each other. As a subject grows past what one hub can sensibly introduce, add sub-hubs. Large mature sites often end up looking more like a mesh, where related pages cross-link through shared entities without a strict hierarchy.
There is no correct number of spokes per hub. Let the number of legitimate subtopics and your reader's navigation needs decide where a sub-hub is warranted. A sound content cluster scale strategy grows sub-hubs because a branch got genuinely deep, not because a chart looked lopsided.
In your ledger, record for each page:
- Its role: central hub, sub-hub, or spoke.
- Its parent.
- Two or three relevant siblings.
- The anchor concepts that connect them.
Then hand those relationships to whoever places the links, rather than improvising anchors mid-draft. Google's guidance on links is refreshingly plain: make them crawlable, use descriptive anchor text, and cross-reference related content.
One caution. Links clarify relationships. They do not rescue a page that has no distinct contribution. A weak page with fifteen internal links is still a weak page, now with better distribution.
How you know it is done: the hub can explain why each spoke belongs, and each spoke can explain what it adds that the hub does not.
Where people go wrong: they hang every page off one broad hub, link everything to everything, or let the architecture get decided after the drafts are already written.
Step 6: Write an evidence packet for each page before anyone drafts
This is the step that separates real programmatic content without thin pages from a variable-filled shell.
An evidence packet is not a reusable article template. It is the research boundary for one page. Write one for every accepted row:
- Page identity: entity, axis and value, intent, audience, stage, page type, parent hub.
- One-sentence job: what the reader can decide, do, compare, or verify after reading.
- Answer block: the direct answer this page gives near the top, in plain language, limited to what the evidence supports.
- Unique evidence: the facts, data, examples, first-hand observations, or analysis that belong to this page and not to a neighbour.
- Source map: which source supports which claim, what date or context matters, and which claims are interpretation rather than fact.
- Boundary notes: what this page will not repeat from the hub or its siblings, and where it links instead.
- Link targets: the internal and external destinations to pass to the linking step.
- Metadata inputs: a descriptive title and summary that do not oversell the result.
Feeling like that is a lot per page? It is less than it looks once the axis work is done, and it is the difference between a page a reader bookmarks and a page that reads like every other result. The packet is what stops one master prompt from producing hundreds of lookalikes.
For citation-ready structure, put the direct answer close to the top, use descriptive headings, make the scope explicit, separate fact from recommendation, and use a table when it exposes a relationship more clearly than prose. Add a statistic or a quotation because it clarifies the answer, not as decoration.
And do not make "contains a citation" your pass condition. A page can link to sources and still be derivative. A citable page has a clear answer, sourceable claims, useful organisation, and a reason somebody would select it instead of the page next door.
DeepSmith carries the packet into production through Deep IQ, which stores your company positioning, the claims to make and avoid, product facts, buyer personas, brand voice, visual guidelines, content types, and a trusted-sources list. Every writing run is grounded in that stored context, so you brief once instead of per article and voice drift stops being a monthly argument. Keep writing the page-specific evidence packet anyway. Stored context makes a page sound like you. The packet is what makes it different from its siblings.
How you know it is done: a writer can state the page's unique job in one sentence, name the evidence that makes it different, and say what it will not repeat.
Where people go wrong: they ask a model to "write 100 versions" from one outline, let the tool fill in missing evidence, or reuse a generic introduction across every spoke.
Step 7: Produce from the ledger, one page at a time
Your approved ledger is now the production queue. Produce by page identity and evidence packet, never by one prompt with a list of substitutions.
The sequence for each row:
- Select an accepted row.
- Load its evidence packet and page-specific context.
- Research and write the direct answer, the analysis, and the supporting sections.
- Apply the intended page type, metadata, source targets, and relationship targets.
- Run the answer-swap and evidence checks against the nearest sibling before it goes into review.
- Send link placement and publishing through their own steps.
Then hold the line on the count. If a subcluster yields ten strong pages, publish ten strong pages. If the next fifty rows are keyword substitutions with no new evidence, stop expanding that branch. A page target is not a reason to keep a weak candidate.
DeepSmith's Content Studio moves accepted rows through New Ideas, the Writer, and Produced Content. The Writer turns one planned idea into a finished, brand-grounded article with internal and external links, a cover image, and publish-ready metadata, because keyword coverage, heading structure, schema markup, and linking happen during creation rather than getting bolted on afterwards. Produced Content is where you review, edit, preview, revise metadata, regenerate the cover, and publish to your CMS. That is how you keep producing distinct pages at scale without the per-article overhead eating the calendar.
Once pages are live, AI Visibility reports mention rate, citation rate, share of voice, sentiment, and visibility trend, and it attributes citations to the specific pages earning them. Engine coverage depends on your plan: Pro tracks ChatGPT, Grow adds Perplexity, Scale adds Gemini, and Enterprise covers all ten tracked engines. Use it to learn which questions and pages deserve more research. No tool guarantees a citation, and any that says otherwise is selling you something.
How you know it is done: every produced page traces back to one entity, one axis value, one intent, and one evidence packet.
Where people go wrong: they measure success by URLs generated, publish everything before the audit, or mistake automatic production for automatic validation that the topic deserved a page.
Why thin pages are a risk worth taking seriously
Thin does not mean short. A long page is thin when it repeats common material, adds no original analysis, offers no useful evidence, or exists only because somebody turned a modifier into a URL. That distinction is the whole point of programmatic content without thin pages.
Google's scaled content abuse policy is about purpose, not method. It targets large amounts of unoriginal content produced mainly to manipulate rankings rather than help people, and it applies however the pages were made. Its guidance on generative AI says the same thing from the other direction: AI can help with research and structure, but generating many pages without adding value can cross that line.
For AI answers specifically, Google says there are no additional requirements, no special AI schema, and no AI-only file needed to appear. A page needs ordinary search eligibility, crawlable links, important content available as text, and structured data that matches what people see. Eligibility is not selection, though. When an engine picks a small set of supporting sources, a page that adds nothing beyond its neighbour simply gives that engine less reason to choose it. That is a selection disadvantage, and it is a good enough reason to plan carefully.
What to do next
Pick one central entity this week. Just one. Build its axis inventory, take the first twenty candidates through the six-question audit, and refuse every row that cannot name a distinct job and an evidence bundle.
You will probably cut a third of them. That is the method working, and it is what a content cluster scale strategy is for.
If you want the baseline map, the gap analysis, and the production to run in one place, DeepSmith puts Content Map, Opportunity Agents, Deep IQ, and Content Studio on the same data. Start a free DeepSmith trial and see your own coverage map before you plan the next branch.



