You spent a week adding internal links. Now someone asks if it did anything, and you don't have a clean answer. That's normal, and it's fixable. This guide shows you how to measure internal linking impact on both sides of the house: Google rankings and AI citations. By the end you'll have a baseline, a control group, a small set of internal linking KPIs, and a decision rule that says whether to scale the work, keep watching, or stop.
Here's the honest version of the question. Does internal linking work? Sometimes, on some pages, in some patterns. Your job isn't to prove links are magic. It's to prove that this change, on these pages, moved something real.
Step 1: Turn "is it working" into one question you can answer
Start by writing a hypothesis you could be wrong about. Vague goals can't fail, which means they can't succeed either.
A workable one looks like this: adding contextual links from these source pages to these destination pages will lift destination-page organic visibility, without hurting the source pages, and will make those destination pages show up more often as sources for a fixed set of AI prompts.
Then pick one primary outcome. Just one.
- Destination-page clicks or impressions in Google Search Console.
- Destination-page average position for a defined query set, treated as a supporting signal rather than an exact rank.
- Exact-page AI citation rate for a fixed prompt list.
Everything else is a supporting metric or a guardrail. That short list is your set of internal linking KPIs, and keeping it short is what makes it usable. Guardrails are what you don't want to break: donor-page traffic, 404 and 5xx responses, canonical selection, sitewide trend.
Name your pages too. Source pages are where the links go. Destination pages are what they point to. Both can move, for different reasons.
Done when: you have a written brief with the page list, the exact link change, the primary KPI, the secondary ones, your pre and post windows, the control rule, and one named owner for the change log.
The mistake almost everyone makes: picking "number of internal links added" as the success metric. That number only proves you did the work. It says nothing about crawling, indexing, rankings, clicks, or citations. Count it if you like, but never report it as a result.
Step 2: Capture a clean baseline before you touch a single link
You can't measure internal linking impact against a memory. Take the baseline first, then change the links, and export the same fields you plan to read afterward over the same length of time.
The search baseline
In the Search Console Performance report:
- Pick the property and the Web search type you'll use again later.
- Turn on clicks, impressions, CTR, and average position.
- Use the Pages dimension for your destination and source groups.
- Use Queries to record what those pages already win, with your branded versus non-branded rule written down so you can repeat it.
- Use Dates to spot trends and anomalies before you start.
- Save the date range and filters, then export the table. Don't screenshot it.
One warning on exports. Values shown as unavailable in the interface come out as zeros in the file, so check it before you calculate anything.
A 28-day baseline with an equal post-change period is a sensible start. It's a choice, not a law. Low-traffic pages and seasonal businesses need longer windows or bigger groups. Record the last complete day in your export too, because the newest Search Console data is preliminary and can still move.
The crawl and index baseline
For each target page, note the URL Inspection result: indexed status, last crawl date, selected canonical, live-test outcome, and the Page indexing status. Grab the Crawl Stats totals for the period. If you have server logs, count verified Googlebot requests to your source and destination URLs.
That canonical field matters more than it looks. Search Console usually assigns clicks, impressions, and position to the canonical URL, not the URL you edited. Record both, or you'll compare two different things later.
The AI baseline
Freeze your observation protocol before the change. For every prompt you check, store the exact wording, the engine and mode, locale, device, account state, the date, whether your brand was mentioned, whether a source link appeared, and the exact URL cited.
One manual search in ChatGPT is a spot check, not a baseline. AI answers shift with retrieval, model updates, prompt wording, and location. Repetition on a schedule you set in advance is what makes the record trustworthy.
Teams skip this part, and that's what makes AI-side results arguable later. To track internal link performance on the AI side, freeze the prompt list and re-run it the same way every time. DeepSmith's AI Visibility module keeps that record: tracked prompts, per-prompt mention and citation rates, the exact pages cited, and the answer history behind each number.
Done when: every treatment and control page has a baseline row, the period end is written down, raw exports are stored, the prompt protocol is frozen, and you've listed every other site change already in flight.
Pro tip: annotate the deployment date in your reporting sheet, then add everything else landing near it: algorithm updates, campaigns, migrations, redirects, big content edits, new backlinks. That timeline is your evidence when someone asks why traffic moved.
Step 3: Split your pages into a treatment group and a control group
Here's the single change that turns a chart story into evidence. Some pages get the links. Comparable pages don't.
If you have enough similar pages, assign them at random. Keep templates and traffic levels balanced, and spread high-traffic pages evenly instead of stacking them on one side.
Internal-link tests need more than one cohort, because the same change hits different page roles:
- Destination treatment: pages that receive the new links.
- Destination control: comparable pages that could have received links but stay untouched.
- Source treatment: pages where you insert the links.
- Source control: comparable pages where you don't.
- Holdout: eligible pages left alone on purpose, if you want to study redistribution or dilution.
SearchPilot calls one version of this "salt shaking," where you evaluate statistical features across a section to pick variant and control groups that genuinely match. You can randomize both ends, apply links to every eligible source page while randomizing destinations, or measure sources, destinations, and non-recipients separately.
Not enough comparable pages? That's common, and it's fine. Use a matched before-and-after comparison, then call the result directional. One page moving after one change is a case observation, not an experiment.
Done when: group assignment is saved before launch, roles are explicit, controls stay unchanged, and pages with unrelated edits are excluded or flagged.
Where people go wrong: putting all the important pages in treatment and all the weak ones in control. Now you've measured page quality, not links, and you won't be able to defend whatever you find.
Step 4: Ship the change and confirm Google can actually see it
Make one documented change, then check the plumbing before you look at any ranking number.
Record the deployment timestamp, source and destination URLs, old and new anchor text, where on the page the link sits, whether it's new or edited, redirect behavior, and anything else in the same release. If three things went live together, report it as a bundle. Don't hand linking the credit for a template rebuild.
The link has to be crawlable: a real HTML link with an href attribute, in the form Google follows. Descriptive anchor text helps people and search engines understand what's on the other end. Google says links help it discover pages and understand relevance. That's a mechanism, not a ranking promise.
After deployment:
- Run URL Inspection on a sample of destination and source pages. Compare indexed version, last crawl, indexing state, and selected canonical against your baseline.
- Run a live test to confirm the page can be fetched and is technically indexable. It can't predict which canonical Google will pick, so read the indexed data separately.
- Watch the Page indexing report. "Discovered, currently not indexed" means found but not yet crawled. "Crawled, currently not indexed" means crawled and left out.
- Request indexing to nudge an important page, knowing a request is not a guarantee.
- Check Crawl Stats and logs to see whether Googlebot came back and whether responses stayed healthy.
Done when: the links are live in the rendered page, sample pages show the intended canonical and HTTP behavior, target pages have been crawled since deployment, controls are untouched, and no new technical error is waiting to explain your result for you.
Common mistake: treating a fresh URL Inspection result as proof the links worked. It proves a narrow crawl and index state. Nothing more.
Step 5: Read crawl and index signals without calling them a win
Crawl data is a chain, and each link in it answers a different question. Change goes live. Page gets discovered or refreshed. Page stays indexable and indexed. Then, separately, visibility moves or it doesn't.
Crawl signals answer "did Google reach the page and could it fetch it?" Indexation answers "did Google include the page?" Neither answers "did this rank better?"
At page level, compare last-crawl dates and indexed data before and after, and check the selected canonical every time, because performance gets attributed there.
At cohort level, use the Page indexing report for trends. Two limits: it covers only URLs Google knows about, and its examples table caps at 1,000 rows. New content can take days to index, and after a fix, validation typically takes up to about two weeks. A status can stay stale until Google recrawls, so check the last-crawl date before deciding a fix failed.
At site level, Crawl Stats gives you requests, response times, host status, responses by type, file type, crawl purpose, and Googlebot type. Use the response breakdown to catch 5xx errors, surprise redirects, and a rising share of 404s. Use Discovery versus Refresh to separate first discovery from routine recrawling. Use file type so you don't credit image fetches to your destination pages.
One caution on crawl budget. Google's guidance there is aimed at very large sites. Don't manufacture a crawl-budget project out of a modest linking change. More requests is not a KPI.
With raw server logs, verify Googlebot rather than trusting the user-agent string. Google's guidance is a reverse DNS lookup on the source IP, then a forward check back to the same IP.
Then read what you find. Good supporting evidence looks like destination pages crawled after deployment, still canonical and indexed, with no new 5xx or redirect problem. Weak evidence is a rise in total crawl requests when your target pages weren't crawled. Negative evidence is target pages stuck at discovered but not crawled, crawled but not indexed, flipped to an unintended canonical, or throwing server errors.
That last group isn't a verdict on internal linking. It's a signal to fix the implementation before you judge the idea.
Step 6: Compare rankings and clicks by page, then by query
Now the part you were asked about. Keep the date lengths, search type, dimensions, and canonical page set identical to your baseline, or you're measuring your own process.
Work the Performance report in order:
- Turn on clicks, impressions, CTR, and average position.
- Start with Pages. Compare destination treatment, destination control, source treatment, and source control separately.
- Drill into Queries for those same pages, with your branded rule unchanged.
- Use Dates to mark the deployment line, and smooth noisy daily numbers with weekly or monthly views.
- Keep Web, Image, Video, and News search types apart. Never mix them between periods.
- Export for the maths. Chart totals and table totals can differ, so pick one and stay with it.
Two formulas carry most of the work. Relative change is post-period value minus pre-period value, divided by pre-period value. Treatment effect is the change in treatment minus the change in control. Same period lengths, same denominators.
Then read the combinations honestly:
- Clicks up, impressions up, position better than control: your strongest organic signal, subject to the statistical result.
- Impressions and position up, clicks flat: visibility improved, traffic didn't. Snippet, SERP features, or demand may be capping it. Don't call it a traffic win.
- Clicks up, position flat: check demand and query mix before you credit the links.
- Source pages down while destinations rise: that's a guardrail breach, not a trade you forgot to mention.
- Only the whole site rises: seasonality, an update, a campaign. A sitewide line can't isolate internal links.
- Position swings hard on a low-impression page: ignore it. Averages are unstable with little data.
Remember what average position actually is. It's the topmost position for your page averaged across impressions, not a daily rank. Of all your internal linking metrics, this is the one people over-read. Let clicks and impressions lead, and keep position in the supporting cast.
Done when: treatment and control are compared over equal windows, results carry raw counts and direction, and your language separates observed movement from proven cause. This is where you learn to track internal link performance properly, one cohort at a time.
Common mistake: reporting average position alone. A page can improve its average position because its query mix changed, while traffic sits exactly where it was.
Step 7: Check whether the changed pages earn AI citations
This is the half most internal linking KPIs still leave out. The question isn't whether AI knows your brand. It's whether your specific destination page is being linked as a source.
Keep two ideas apart. A mention names your brand. A citation links one of your pages. Only the second counts here, and only when it points at the page you changed.
Google's AI features
Google's Generative AI performance report covers impressions from generative AI features in Search, including AI Overviews and AI Mode. It gives you Pages (grouped by the final URL after redirects), Countries, Devices, and Dates, on Pacific Time. Search Labs experiments are excluded, and the report is still rolling out, so a property without enough generative-AI impressions may not have it yet.
Use the Pages view for one narrow job: did your destination page start appearing as the linked source after the change? Don't blend those impressions with prompt-level citation rates from another system. Different denominators, different questions.
Everything outside Google
For ChatGPT, Perplexity, and the rest, your unit of measurement is a fixed prompt observation. Every run stores the prompt, the engine and mode, locale, device, account state, date, the full answer, whether your brand was mentioned, whether an exact citation appeared, the cited URL and its final destination, and whether the cited page was your test page, a source page, a competitor, or nothing at all.
Then repeat it after deployment, unchanged. Change the prompt and you've changed the question. Change the locale and you've changed the retrieval conditions. Report per prompt and per engine before you aggregate, because a page cited for one prompt and not another isn't a failure. It's a normal result.
OpenAI's publisher guidance explains that public sites can appear in ChatGPT search and that crawler access affects discovery and citation. Perplexity describes its answers as backed by verifiable sources with links to the originals. That's why exact source-link observations matter. Neither is a cross-engine analytics standard, and a manual search is never a complete impression count.
Doing this by hand across engines and dates gets heavy fast, which is why it usually stops after week two. DeepSmith's AI Visibility module keeps the record in one place: mention rate, citation rate and share of voice with trends, per-prompt history, a Pages view showing which of your pages get cited and what share of your citations each one holds, and which competitor pages win the prompts you don't. Coverage rises by plan, from ChatGPT on Pro up to all ten tracked engines on Enterprise.

It gives you a stable record on both sides of the change. It can't make your link change causal, and no tool can. That's your control group's job.
Done when: the same prompts and platforms ran in both periods, exact cited pages are recorded, mentions and citations aren't conflated, and Google AI-feature impressions sit in their own column.
Pro tip: a page-level citation lift is worth more here than a brand-wide mention lift. If the brand gets named but your destination page never gets linked, the hypothesis hasn't been shown on the AI side yet.
Step 8: Attribute the movement and make an honest call
You have numbers. Before you write the word "because," walk the confounder list. This is where you either measure internal linking impact or just narrate a chart.
Check the timeline for Google ranking updates, seasonality, content edits in either group, changes to titles or schema or canonicals or robots, redirects and template releases, backlinks gained or lost, campaigns, competitor activity, AI model changes, and any outage.
Randomized page-level controls absorb a lot of that, which is the whole point of Step 3. Without randomization or a defensible matched control, use the words "associated with" or "directional." Not "caused."
On timing, SearchPilot reports that SEO tests generally reach significance in two to four weeks, with 95% confidence as the usual convention, and that real duration depends on your traffic and the size of the effect. Their suitable-test conditions assume hundreds of same-template pages and at least 30,000 organic sessions a month for the tested group. That's a condition of their method, not a minimum for your business. Smaller site? Use a matched holdout, watch longer, and write more carefully.
Their published internal-link tests give you a sense of scale. One across roughly 8,000 regional pages added links to nearby regions and reported a 7% organic-traffic uplift on receiving pages. A homepage-footer test reported 5% on destination pages. Adding more related-article links lifted the donor pages but showed no conclusive gain for recipients. Rewriting existing anchor text to be more contextual came back significantly positive for destinations. These are observations from one testing program, not a number to promise your CEO.
Then pick your decision:
- Scale the pattern: treatment beat control credibly, crawl and index health is sound, donor guardrails held, and the page-level AI signal is positive or at least unharmed.
- Keep observing: crawl and index movement is there, but ranking or AI data is still too thin.
- Call it a channel win: Google improved and AI didn't, or the reverse. Report them separately.
- Revise the implementation: pages crawled but not indexed, wrong canonical selected, redirects or server errors appeared, or the link wasn't rendered in crawlable form.
- Stop or roll back: control-adjusted performance is negative, or donor pages breached the guardrail. Recheck the design before concluding that internal linking never works.
- No conclusion: both groups moved the same direction, or prompts changed, or another release shipped the same week. Say "no conclusion" out loud. It's a real answer.
Done when: your report holds the change, the cohorts, the windows, the formulas, raw counts, the treatment-versus-control result, crawl and index checks, the AI prompt result, the confounder review, your confidence language, and the next action. That report is the real answer to "does internal linking work" for your site.

What to do next
You don't need to link everything this quarter. You need one clean measured cohort, then another.
Keep the winning pattern in a documented test library, with the design and caveats attached. Roll it out only to pages that genuinely match the ones you tested. Put a recurring date in the calendar to revisit the control group and re-run the AI prompt set, because both go stale quietly.
If the manual side is what keeps stalling the work, solve that separately. DeepSmith holds the prompt, page-citation, competitor, and trend record in one workflow, and its Content Map plus writing pipeline cut the hours of finding and placing links in the first place. The same internal linking metrics still apply to anything it produces. You can start a free trial and run your next cohort with the record already in place.
Take it one cohort at a time. The proof compounds.



