Internal linking embeddings give you a way to find related pages in a client's library that nobody would spot by scrolling through a spreadsheet of URLs. This guide is for agency strategists who look after several client sites at once and want a repeatable process for finding link opportunities across hundreds or thousands of pages. By the end you'll have a review table for each client that lists the pages worth linking, the passage where each link would sit, and a decision on every row.
The method has three parts that do different jobs. Semantic similarity finds pairs of pages that plausibly belong together. Comparing those pairs with the client's real link graph finds the ones that are missing a link. Editorial review finds the ones that are worth adding. If you keep those three jobs separate, you'll get a short, data-driven internal linking list you can defend to a client instead of a long list of URLs that only look related.
This guide covers the discovery method only. It doesn't explain what embeddings are, and it doesn't go into link architecture or anchor text, since those are covered elsewhere.
What you need before you start
Set aside a bit of time to gather four things for one client:
- A crawl or CMS export of the pages that should be eligible, with the canonical URL, page type, status, title and the extracted main text.
- A fresh export of the existing internal links, with a source URL and a destination URL on every row.
- One way of creating embeddings. That could be the Screaming Frog SEO Spider's embedding integration, an embeddings API, or a sentence-transformer model you run yourself. The API route needs a key and usage credits.
- A person who can decide whether a destination page really adds something at a certain point in a source article.
A spreadsheet is fine for the review stage. Python, or a vector search setup, helps once the libraries get large.
1. Define the client corpus and eligible destinations
Start with one client property. Export the pages you want to check for links and the pages that could receive them, and keep the canonical URL and page type next to each page's content. Then remove anything the client's publishing rules say shouldn't be involved, such as error pages, redirects, noncanonical duplicates, noindex pages, utility pages and content that can't be accessed. If a page has changed a lot since the crawl, extract it again.
You also need to decide whether blog posts, guides, documentation and product pages share one pool or get compared inside page-type groups. Either choice is fine as long as you write it down.
You're done with this step when every row points to an accessible, unique page, and you can check whether a page is allowed as a source and as a destination separately.
Common mistake: treating a sitemap as a full, current crawl. That's how redirects and duplicate URLs turn up as separate recommendations, and how one client's pages end up in another client's candidate pool. Search Console's Links report is useful context, but it isn't a complete live inventory. Google groups pages by canonical URL, combines duplicate links, can keep links it found in the past, and limits how much it shows you.
For an agency, it helps to save the crawl date, the client identifier, the eligibility rules and the model version with every run. Those records are what let you rerun a client's audit months later and explain why a page was or wasn't included. Keeping each client's inventory and credentials apart from the others is part of the same habit. If you're mapping a client's topics before you begin, DeepSmith's Content Map gives you a current view of what a site covers, and it's a handy way to check that your export matches the topics the client thinks it has.
2. Extract the main content and check its quality
Before you create any page representations, strip out the navigation, footer, sidebar and repeated template text. If those stay in, pages that share boilerplate will look related to each other when they aren't.
Screaming Frog lets you include or exclude particular tags, classes and IDs when it defines the content area. Its embedding tutorial says navigation and footer are excluded by default, and it suggests you look at View Source > Visible Content to see exactly what the embedding process receives. Oncrawl's recommender write-up makes a similar point, that scraping entire pages instead of the primary article text raises cost and lowers the quality of the recommendations.
You're done when a sample of extracted records contains real article or page text, and not mostly menus, cookie notices or related-post widgets. Flag records that are empty or very short instead of quietly scoring them.
Long pages need a decision too. Embedding services have an input limit, and if a page gets truncated, the passage that would have justified a link might be the part that got cut. If you use an API, decide up front how to treat overlong pages, and keep the excerpt or section identifier with any section-level candidate so you know what was actually compared. A score for a whole page doesn't prove that two specific passages belong together.
If you're using Screaming Frog, confirm that there is an embedding value in the AI tab and that Prompt Request Status says Success. Its tutorial describes testing a failed URL under Config > API Access > AI > AI Provider > Prompt Configuration. It also describes a Limit Page Content option that restricts what gets submitted to 5,000 characters, but that's a troubleshooting setting, not a recommended article length.
3. Generate comparable embeddings for each page
This is the step where internal linking embeddings get made, and consistency matters more than which model you pick. Use the same model and the same preprocessing for every page you're comparing in a run. Save the model identifier, the content version and the mapping from each page to its vector, because you'll need them the next time you run the audit.
With an OpenAI API route, the embeddings endpoint takes several inputs in one request. Each input has to be non-empty and inside the model's limit, and the API reference states 8,192 tokens per input and 300,000 summed input tokens per request. Tokens aren't words, and the limits depend on the model you pick and your account, so check them before you build the batching. With a local Sentence Transformers route, you encode the whole corpus the same way. Its documentation separates comparing documents to documents from searching long documents with a short query, and it suggests encode_query and encode_document for the second case.
You're done when every page you meant to compare has one usable representation, or a clear error status next to it. Test a small, varied group of pages first. Oncrawl suggests trying its pipeline on 10 queries and URLs and looking at the output before you run the full set, and that's a sensible habit whichever route you choose.
Common mistake: mixing vectors from different models or dimensions, missing failed API responses, or re-embedding unchanged pages every time. Oncrawl's pipeline caches content embeddings, so later runs only process new or changed content, which is worth copying. Model quality also varies by domain and by language, so check a multilingual library on its own content.
If you'd like to do this from a crawler, one published Screaming Frog walkthrough turns on JavaScript rendering at Configuration > Crawl Config > Rendering > JavaScript, opens Configuration > Custom > Custom JavaScript, picks the library's embedding extraction script and asks for an API key. The exported columns get renamed URL and Embeddings for the later similarity script. A Moz workflow exports those embeddings along with All Inlinks to find missing links between pages. Interface labels change between versions, so check what your copy shows before you follow any walkthrough exactly.
4. Retrieve the nearest related pages
Vector search internal linking comes down to one question here, which pages sit closest to this one. For each eligible page, rank the other eligible pages using the similarity or distance measure your model supports, and leave out the page itself. Keep the list short. The Screaming Frog and Moz examples both use five candidates per page, and that's a good place to start, though it's a workflow choice and not something proven to be the best number for SEO. Keep the raw scores so reviewers can compare candidates inside the same model and corpus.
You're done when each page has a small ranked set of different, plausible peers, with scores and stable URL identifiers next to them. Spot-check some distant topic areas and some pages that serve more than one intent.
This is where semantic similarity internal links most often go wrong, so it helps to slow down here. It's tempting to say every pair above some cosine score deserves a link, but there's no sourced universal cutoff for that. Rankings depend on the model, on which text you embedded and on the site. A very high score can mean two pages duplicate or overlap each other, which is a reason to look at them and not a reason to add a link. Screaming Frog's default 0.95 semantically similar threshold is there to flag nearly overlapping pages, and its 0.4 semantic relevance setting compares a page with a centroid of the whole site's content. Neither one was designed as a cutoff for adding internal links.
If you'd like to set a cutoff, set it from a review of your own client's pages. Look at a small sample of high, medium and lower-ranked candidates, and record the model, the page types and the corpus you used when you chose the number. Then the number belongs to that client and that run, and you won't be tempted to reuse it somewhere it doesn't fit.
Choosing how to run the search
Either way, vector search internal linking works best when the client filter and the model settings stay fixed for the whole run. How you retrieve neighbors depends on how big the library is and how often you rerun it.
- Spreadsheet or small script. Fine for a manageable library. Sentence Transformers documents
semantic_searchwith cosine similarity by default andtop_k=10by default, returning acorpus_idand ascorefor each hit. It describes direct semantic search for corpora up to about one million entries. Those are library defaults and not advice on how many links to add. - Vector index. For very large or latency-sensitive jobs, approximate nearest-neighbor search trades a little recall for speed. pgvector does exact search by default and offers HNSW and IVFFlat indexes for approximate search. A vector database is optional and isn't a prerequisite for vector search internal linking, and it isn't needed for a first run.
- Watch the operator. In pgvector,
<=>is cosine distance, not similarity, and cosine similarity is1 - cosine distance. A smaller distance means nearer, while a larger similarity means nearer, so it's easy to sort the wrong way around. - Keep the client filter in retrieval. If you filter by client after an approximate index scan, check whether you're left with too few results.
5. Subtract the links that already exist
Similarity scores alone can't tell you what's missing, so this step brings in the client's real links. Now build a set of directed pairs, each one a source canonical URL and a destination canonical URL, from your existing-links export. In Screaming Frog the documented route is Bulk Export > Links > All Inlinks, and the report includes the source URL, destination URL, anchor text, follow information and link position.
Join that set to your ranked candidate pairs. Mark an opportunity only when the source doesn't already link to the destination. A linking to B and B linking to A are two separate checks, and a reciprocal link doesn't mean both directions exist. If your link-position data allows it, separate links that sit in navigation or a template from links inside the body copy, and don't assume a template link means the body needs a second one.
Common mistake: a list of near neighbors is not a list of opportunities. Remove the directed links that are already there, and then ask whether what's left makes editorial sense. Related mistakes are joining raw URLs without resolving canonical variants, and treating something missing from Search Console as proof that no link exists.
You're done when a reviewer can see the exact missing direction on each row and can check a sample against the live source page. If the client cares about the difference, give existing body links, existing template links and no observed link their own statuses.
Moz's published example is a good picture of what you're aiming for. It joins an embeddings export with an All Inlinks export and shows each target URL with its current linking pages and its top five similar pages, with highlighted cells marking the related pages that don't already link to the target.
6. Find a passage and decide on each candidate
Open the proposed source and destination side by side. Look for a specific passage in the source where a reader would benefit from the destination's own explanation, evidence, example or next step. Record that passage and your reasoning, and mark the row accept, revise or reject.
If the whole-page similarity is broad, compare sections or look at the destination by hand instead of forcing a page-level match into a paragraph that doesn't want it.
You're done when every approved row has a real placement context and the reviewer can say why this destination helps at this point in the article. A high score by itself is never a reason to approve.
Pro tip: a hypothetical example makes the standard clear. Say a source guide about onboarding several agency clients ranks a destination about per-client brand context among its nearest neighbors, and the link is missing from the latest crawl. The reviewer finds a paragraph about re-briefing every account, and approves a link there. A second high-scoring destination mostly repeats the source article, so the reviewer rejects that pair even though the score looks good. Both decisions came from reading the pages, and neither came from the number.
Things that go wrong at this step include cross-linking near-duplicate pages, linking pages that share a subject but answer different questions, mistaking a generic template passage for editorial context, and linking to a destination that's out of date. Google's guidance on links recommends linking to other resources on your site that help readers understand a page in context, and that's a good test for each row. Screaming Frog says plainly that its semantic results need human interpretation, and Oncrawl cautions that semantic relevance on its own can overlook user behavior.
Screaming Frog's built-in views can help you look. Its Semantically Similar filter shows a closest matching URL, its similarity score and the number of similar pages, and its Content Cluster Diagram can show inlinks, outlinks and inlinks within a cluster, which makes gaps between related pages visible, including pages in different sections of a site. They help with investigating, and you still need the directed-link check and the passage check.
7. Prioritize a small queue you can defend
This is where data-driven internal linking becomes a real service and not a one-off report. Sort the approved opportunities by how important the client considers the destination, how good the source passage is, whether the destination has few internal links today, and how much editorial effort the change takes. Similarity found the candidate, but it shouldn't be the whole priority formula.
For every row, keep the client, source, destination, score, existing-link status, passage, reviewer, decision, change date and reason. Keep rejected rows as well as approved ones, so a later run doesn't recreate the same rejected task unless something about the pages has changed. Then pass the approved list to the client's normal publishing workflow.
These are the four tables that hold everything together:
| Table | Minimum fields | What it's for |
|---|---|---|
| Eligible pages | client, canonical URL, page type, main text, content-change marker, inclusion status | Defines the pools of possible sources and destinations |
| Embeddings | client, canonical URL, model and version, dimensions, representation, processing status | Prevents silent mismatches and allows incremental updates |
| Existing links | client, canonical source, canonical destination, observed link position, crawl date | Shows which directed links are already present |
| Candidate decisions | client, source, destination, score, existing-link status, passage, decision, reviewer, publication date | Supports editorial review and later verification |
You're done when you have a finite queue that you can show to a client, along with which candidates were accepted, which were rejected and which were implemented, and the reason for each.
Common mistake: promising that a score turns into rankings, impressions or AI citations. Nothing in the sources behind this guide gives a universal similarity score, a number of new links per page, a saving in review hours, a ranking improvement or a rise in citations. Screaming Frog's internal-link guide treats underlinked, high-value pages as useful priorities, but it warns that results vary by site, so don't present its illustrations as a case study.
DeepSmith fits into the wider work around this audit. Content Studio's Writer includes internal linking when it produces new articles, so pages that come out of that workflow arrive with links to related content already built in. Each client can also have its own workspace with its own brand context, which keeps one client's pages and voice from leaking into another's. The similarity audit itself is a separate process that you run alongside those.
8. Verify the change and rerun the process
After the approved changes are live, crawl the affected pages again and check that each source now points to the right destination. Compare a dated pre-change link export with a post-change one, take the finished pairs out of the queue, and refresh the representations for pages that changed or were newly published. Screaming Frog documents Crawl Analysis > Crawl Comparison for comparing a baseline crawl with a later one, and that comparison needs database storage mode.
You're done when the link shows up in the new crawl, failed placements are reopened, and the next candidate run works from current pages and current links. A client-facing record should keep four states apart: candidate found, editorially approved, published, and observed in a later crawl.
Common mistake: measuring the link you predicted instead of the link that got published, comparing different URL scopes, or giving an internal-link change credit for a visibility shift that came from something else. Google says indexing and serving aren't guaranteed. Its guidance on AI features says ordinary search fundamentals, including making content findable through internal links, are still worth doing, but it doesn't promise an AI citation from any single link. Track search and AI visibility as separate observations. DeepSmith's Pages and Prompts views show which of your pages get cited by AI and which prompts drove those citations, and they work as monitoring. They don't show that one added link caused one citation.

What to do next
Pick one client and one section of their site, and run the eight steps end to end on that. A small run of semantic similarity internal links on one section will show you where your extraction, your model choice and your review habits need work, and that's much cheaper to learn on a hundred pages than on ten thousand. Once you have a table you trust, the same steps apply to the next account. If you'd like to see how DeepSmith handles the production and monitoring around this work, you can start a free trial and try it on a real client workspace.



