You can run a great internal link audit once. Running it across twelve client sites, every month, at a margin you can live with, is a different problem. That gap is where most agencies get stuck, and it is not a skill gap. It is a systems gap. This guide gives you the eight steps that turn an agency internal linking process into a service you can price, repeat, and report on across a whole roster.
You will need a crawlable site or staging environment, Search Console access where the client can grant it, a sitemap inventory, someone who can actually publish, and a shared change log. Missing one? Note it. Do not pretend it is there.
What you are actually selling
You are not selling "we will add internal links." An internal linking service is a named sequence with named outputs, and it delivers six things every cycle:
- Intake and scope record. Site, templates, URL classes, priorities, access, exclusions, release owner, reporting window.
- Baseline pack. Crawl settings, URL and sitemap inventory, a Search Console baseline where you have access, current link health.
- Prioritized backlog. One row per proposed change: source, target, target status and canonical, anchor, reason, evidence, priority, owner, state.
- Implementation batch. A bounded set of approved changes with a batch identity and a deployment date.
- QA delta. The same crawl and checks rerun after release, with before and after values plus exceptions.
- Impact report. Organic and AI visibility metrics, the change window, affected cohorts, caveats, and the next queue.
The fixed part is the workflow: the fields, definitions, states, and quality gates. The variable part is the client's priorities, architecture, CMS, page types, and release cadence.
An internal linking productized service is never defined by a link count. "Five links per page" is a quota, not a service.
Step 1: Define the service boundary and collect the intake
Define one unit of work before you crawl anything. Write down the client site, the host variants, the locale, any separate properties, which page types are in scope, and which templates are excluded. Then decide what the engagement covers: recommendations only, implementation, QA, recurring monitoring, or all four. Your intake checklist should capture:
- Primary domain plus every in-scope subdomain, locale, and property.
- CMS, framework, repository, staging environment, and deployment method.
- Sitemap files or index files, including blog, docs, product, and media sitemaps.
- Search Console and analytics properties, plus crawler access needs: authentication, robots.txt, CDN restrictions, rendering.
- Priority page classes, plus exclusions like pages tied to revenue, launches, or a migration.
- Named owners for content, technical, and release approval.
- Baseline date, reporting period, release date, and what a successful delivery looks like.
- Any other SEO work running at the same time that could confuse your measurement.
Ask for the least access that supports the work. Search Console separates owners, full users, and restricted users, and a reporting-only engagement rarely needs ownership. Search Console alone cannot publish a link either, so implementation still needs a CMS, a repository, or a client release path.
Done when: the scope has one owner, one property set, one baseline date, explicit exclusions, a tested access matrix, a release path, and a written priority order. You can say what you will deliver and what you will not touch before anyone changes a page.
Where people go wrong: starting with a crawl and asking the business questions later. You end up with a technically huge backlog nobody cares about commercially. Mixing staging, production, locales, or a second brand into one crawl corrupts the baseline too. Never promise implementation when you have no release owner.
Step 2: Set up an isolated client workspace and a source-of-truth brief
One workspace per client. Not one folder with tabs. One workspace.
Inside it, store that client's positioning, products, personas, claims to make and claims to avoid, brand voice, approved sources, and page-type definitions. Keep their data, queues, approvals, and reports separate from everyone else's. Use one naming scheme for projects, crawls, batches, and reports, so another strategist can pick up the account without reconstructing it from old messages.
This is where DeepSmith fits the agency operating model. Multi-Workspace runs multiple brands or clients from one account, each isolated with its own context, content, and plan. Deep IQ holds that account's positioning, products, personas, brand voice, and content types. When the service later produces or refreshes a page, the draft comes out in that client's voice with that client's facts. No cross-account bleed, no re-briefing.
A workspace is not a substitute for approval, access, or a crawl. It just means you stop paying the onboarding tax twice.
Done when: a new team member can open the client record and know which site is in scope, which pages matter, what language is approved, who can publish, and what is different from your default. A test recommendation comes back with the right client's facts.
Where people go wrong: one agency-wide spreadsheet or prompt holding every client's product facts. That is a leakage risk, and it gets worse the moment you add automation. Separate the data first, then automate.
Step 3: Capture a baseline crawl you can reproduce
Same basic sequence for every client, then record the exceptions.
- Collect the client's sitemap URLs. Google describes a sitemap as a list of preferred canonical URLs, not a promise that every URL gets crawled or indexed. Each file is capped at 50 MB uncompressed or 50,000 URLs, so a big inventory needs multiple files or an index.
- Crawl the public site or approved staging environment with a consistent configuration. Do not cap crawl depth when you want a complete link audit, and on a large site, segment the crawl and write down the boundaries.
- Render JavaScript when the site inserts links client-side. Important links should exist as crawlable anchor elements with a resolving href, not only as an interaction a crawler might not process.
- Merge crawl-discovered URLs with sitemap URLs, Search Console URLs, and analytics landing pages. This merge is the whole game for orphan discovery, because a page can sit in a sitemap and take traffic while nothing internal links to it.
- Freeze the baseline. Save the crawl date, configuration, URL count, status and canonical fields, rendering mode, exports, and known exclusions. Never overwrite last cycle's baseline.
Capture at least these row-level fields: URL, template, HTTP status, final destination, canonical, indexability, crawl depth, unique internal inlinks, anchor text, where the link sits (nav, template, body, rendered-only), orphan status, broken or redirected target flags, and clicks, impressions, CTR and average position where you have access.
DeepSmith's Content Map gives you the topic layer on top of that. It crawls the client's site, classifies every page onto a shared topic taxonomy and a funnel stage, maps competitors onto the same taxonomy, and rechecks sitemaps every 24 hours so new pages fold in on their own. Use it to see coverage, topic depth, and the missing pages that shape your link plan. The link-level crawl and status validation stay your own technical control.
Done when: you have one reproducible URL universe, you know why every URL entered it, you know which pages internal links cannot reach, and you can rerun the same crawl later.
Where people go wrong: treating the sitemap as the full inventory, which hides orphans and non-canonical URLs. Crawling only from the homepage, which misses URLs Search Console already knows about. Ignoring JavaScript, which invents problems that are not there. Changing crawl settings between baseline and QA, which makes the delta unreadable.
Step 4: Score every problem and opportunity in one backlog
Raw crawl data is not a deliverable. A decision queue is. Sort what you find into these issue classes:
- Orphan page. Known from a sitemap, Search Console, analytics, or a CMS export, but unreachable through the internal link crawl.
- Deep page. An important page more than three clicks from the homepage under your crawler's model, with the homepage counted as depth zero. Treat one to three as a triage target, not a Google ranking rule.
- Underlinked page. A valuable page with few unique link sources compared to similar pages in the same template or topic group.
- Bad destination. The link points at a broken URL, a redirect, a non-canonical duplicate, an excluded page, or the wrong host or locale.
- Weak anchor or context. Anchor text is empty, generic, or disconnected from both pages. Google describes good anchor text as descriptive, reasonably concise, and relevant to source and target, and warns against cramming keywords in.
- Rendering or block problem. The link only appears through a JavaScript behavior rather than a crawlable anchor, or it sits in a template or footer block so crowded that any change needs a controlled test rather than a blanket rollout.
Then prioritize with a rule you can show the client, not an opaque score. Four dimensions work: business value of the target, evidence of need, search opportunity, and risk and effort. Bands like Now, Next, and Later are fine once you have written down what puts a row in each band. Do not let raw link count decide priority. An important page with one relevant source can beat a low-value page with none.
When the target page you need does not exist yet, that is a content decision, not a link decision. DeepSmith's Content Map and Opportunity Agents surface evidence-backed topic gaps, and every idea arrives carrying the data point that justifies it. Use that to decide what to write, not as proof that a link will rank a page.
Pro tip: keep the evidence attached to the row. A strategist should be able to answer "why this source, why this target, why now, and what will we measure" by opening one line, not by digging through three exports.
Done when: every candidate has a source, target, status and canonical check, anchor proposal, placement, rationale, evidence, priority, owner, and state. Duplicates are merged, out-of-scope pages are marked, and the queue filters by client, template, batch, and owner.
Where people go wrong: the quota. "Add five links to every page" produces irrelevant links and confuses activity with value. Reusing one exact-match anchor everywhere makes the copy read badly. And treating a depth threshold or a correlation study as a guarantee will eventually cost you a client's trust.
Step 5: Get the queue approved by the client and the release owner
Approval is where internal linking for clients either becomes a service or stays a slide deck. Separate four states that usually get mashed together:
- Proposed. Backed by evidence, not yet approved.
- Approved. Client and technical owner accept source, target, placement, anchor, and timing.
- Implemented. Live in the intended environment, with a deployment record.
- Verified. Post-release crawl and QA pass, or a documented exception.
Send a decision-ready batch, not a raw spreadsheet. Each row carries the reason, the affected template or page class, the proposed anchor and copy, the expected risk, the implementation owner, the release date, and what you will check afterwards. Anything needing brand, legal, or technical review goes into an exception lane, so it does not hold up everything else.
Use the approval record to settle permissions too. Can you edit body content? Navigation? Related-content modules? Header and footer templates? Or are you recommendations-only? Links get inserted by the approved implementation step, never improvised by a writer who has not seen the queue.
Done when: the client can approve or reject each batch, the implementer knows the environment and release path, and the backlog tells recommendations apart from live changes. Rejections carry a reason, so nobody resurfaces the row next month.
Where people go wrong: sending hundreds of unranked opportunities. That is approval fatigue, and it stalls the engagement. Pushing a recommendation straight to production without checking the target's status, canonical, and release owner makes you responsible for a change you were never approved to make.
Step 6: Ship changes in controlled batches
Batching is what makes scalable internal linking possible. Group approved rows into bounded sets you can attribute and roll back: one template, one topic cluster, one page class, or one comparable test cohort. Record the batch ID, source and target URLs, final anchor, environment, implementer, deployment date, and any other change released alongside it. That last field saves you an argument three months from now.
Prefer a release path that preserves review: staging or preview, client or technical approval, production, then a post-release crawl. Where a target is only reachable through a redirect, point the link at the final destination. Your acceptance criteria are relevant, crawlable, resolving links. Not a link count.
Testing a change? Use a comparison design rather than a before-and-after story. SearchPilot's method splits statistically similar destination pages into groups, with a separate group of source pages where links change, so the difference can be attributed. Log every other release that could affect either group.
When the batch includes a new or refreshed article as the source or the target, this is where production time disappears. DeepSmith's Content Studio takes a planned idea through the Writer to a finished, brand-grounded article with internal and external links, SEO and AEO structure, a cover image, and publish-ready metadata already in place. Autowrite writes scheduled articles into Produced Content with nobody in the app, and from there you publish to WordPress, Webflow, Strapi, Sanity, Contentful, or a webhook. That is production help for approved content, not permission to edit a client's site outside their release controls.
Done when: every live change maps to an approved row and a batch ID. You can roll back or name the exact release, and your post-release crawl can isolate the changed cohort.
Where people go wrong: shipping every recommendation sitewide in one release. Attribution and rollback both die. Changing content, templates, redirects, canonicals, and links at once without logging it makes any later result impossible to explain.
Step 7: Recrawl, QA, and close the release
Rerun the baseline crawl with the same scope and settings. Same settings. This is the step people rush, and the one that protects you. Check the changed rows and the sitewide side effects:
- The source page is live and the target link is present.
- The link is a crawlable anchor with a resolving href, including after JavaScript rendering.
- The target resolves to the intended final URL, with no new broken link or avoidable redirect.
- The target's canonical and indexability are still what you intended, in the approved placement, with the approved anchor.
- Orphan status, crawl depth, and unique inlinks moved the way you expected.
- No template release quietly removed links from unrelated pages, and the crawl exposed no new broken, wrong-host, or rendering-only links.
Compare baseline and QA exports by stable URL and batch ID. Mark each row Verified, Failed, Blocked, or Accepted with Exception, write the reason for every exception, and carry it into the report. If what shipped differs from what was approved, reopen the row.
This is the point where an agency internal linking process starts to feel boring, and boring is the goal. Boring is repeatable.
Done when: the QA report has a before value, an after value, a pass or exception state, a reviewer, a date, and a batch ID for every change. The client gets a summary saying what went live, what did not, and what you will measure next.
Where people go wrong: checking a page in a browser and calling it done. That misses broken targets, canonical changes, removed links, and depth regressions. Checking only the changed pages misses template-wide side effects. A screenshot is not a recrawl.
Step 8: Report SEO and AI citation impact, then book the next cycle
Reporting is the part of internal linking for clients that decides whether the retainer renews. Build it around the change log and the cohorts, never around one hand-picked winner. Use the same date windows, page definitions, query groups, prompt set, and engine set before and after the release, and report a comparable unchanged group beside the changed one where you have one.
For the organic scorecard, define the metrics plainly. Clicks are the times a user clicked through from Google results, impressions are how often the site appeared, and CTR is clicks divided by impressions. Average position is the average position of the topmost result from the site, and it deserves that word "average" every time you say it. Newest Search Console data can be preliminary, chart and table totals can differ because the aggregation differs, and clicks and impressions attribute mostly to canonical URLs. Those caveats belong in the methodology note, not in fine print nobody will find.
Keep AI visibility on its own scorecard. A mention means the answer named the brand. A citation means the answer linked to one of the brand's pages as a source. One happens without the other all the time, so never blend them into one "AI rank" number.
This is the layer that wins renewals right now, and it is what a client is really asking when they say a competitor showed up in ChatGPT. DeepSmith's AI Visibility module lets you define the client's buyer prompts, run them on a schedule, and capture full answers. It tracks mention rate, citation rate, share of voice, sentiment, and visibility trend, with a per-platform breakdown, page-level attribution, and a competitor leaderboard showing who wins the prompts you want. Coverage follows the plan: Pro tracks ChatGPT, Grow adds Perplexity, Scale adds Gemini, and Enterprise covers all ten engines.

Map changed target pages to the prompts they support, then compare citation rate, cited pages, and competitor pages before and after the release. Report what you observed. Do not claim a link caused a citation unless your test design can isolate it. Google's guidance says its core SEO and crawlability practices still apply to generative AI features, that a page must be indexed and eligible to appear with a snippet to be eligible as a supporting link, and that meeting those requirements guarantees nothing.
Common mistake: promising "more links equals more rankings" or "internal links guarantee AI citations." Google publishes no ideal link count and no guarantee that an eligible page gets crawled, indexed, served, or cited. Say what you measured. It is a stronger sell anyway.
Done when: the cycle report can be generated from the change log, crawl exports, Search Console data, and AI visibility history without rebuilding the methodology each month. It shows what changed, what moved, what is uncertain, and what comes next.
Where people go wrong: reporting traffic without confirming the links shipped correctly. Reporting rank without clicks and impressions. Reporting mentions as citations, which overstates the result. Reporting one winning page instead of the changed cohort, which turns a service report into an anecdote.
Set the cadence, then let it run
You have the eight steps. Now make them recur. A practical default: a baseline at onboarding, a delta crawl after every material template, migration, or content release, and a recurring review on the agreed cycle. Each cycle you clone the approved crawl config, pull the newest sitemaps, record what changed, recalculate the issue classes, batch and approve, recrawl and QA, update both scorecards, and carry the open questions forward.
A few portfolio controls are what make scalable internal linking real instead of theoretical. One intake template, one row schema, one set of states, one report outline. One workspace per client, never one shared context. Versioned crawl settings and a named baseline per site. A fixed QA checklist with a second reviewer on high-risk template changes. A capacity review based on pages, templates, batches, and QA load, not just the number of logos.
Start with one client. Run the eight steps once, all the way through, and write down every place you improvised. Those notes are your template, and they are how an internal linking productized service actually gets built. The second client takes half the time. The fifth barely takes any thinking.
Want the reporting layer and the content production running off the same client context you set up in step 2? Start a free DeepSmith trial and build your first client workspace. Seven days, real data, no long-term contract.



