DeepSmith

Sep 26 · Content Production

17 min read

How to Use TF-IDF and NLP Term Analysis to Close Content Gaps on a Page

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
Monochrome diagram of a central document card with gaps in its text lines, linked by thin lines to four peer document cards, with the cover line Find the Missing Terms.

Sooner or later a client hands you a page that covers the right subject but reads thinner than the pages ranking above it. Nobody can point to what is missing, and "add more depth" is not something a writer can act on. TF-IDF SEO work gives you a way to answer that question with evidence, by comparing your page to a small set of ranking pages, so the term frequency content gaps show up as a list of explanations that they include and yours leaves out. That is what a real content depth analysis looks like for one page, and this guide walks you through it one page at a time. It ends with a short revision brief you can hand to a writer and keep on file for the client.

You'll need Python with scikit-learn installed, a Search Console login for the client's property, and about half a day for your first page. After that it gets faster, because the worksheet and the steps stay the same for every client.

One thing to settle before you start. TF-IDF is a way of describing a set of texts that you choose. It is not Google's scoring model, and no score tells you that a page is done. The whole method is a way to produce candidates for an editor to judge, so you'll keep asking what reader question a missing term stands for.

Fix the page, the query, and the audience

Start by writing down which page you are working on, who it is for, and which market and language it targets. Then open the Search results Performance report in Google Search Console, filter to that page, and look at the Queries tab. Note the impressions, clicks, click-through rate, and position for the queries that fit the page's purpose, and record the date range you used.

Pick one query, or a small group of queries that really mean the same thing, for the comparison. If the page is an unpublished draft, pick the query you intend to target and write down that there is no baseline yet, so nobody later mistakes a hoped-for query for a measured one.

You're done with this step when your notes name one page, one search intent, the audience and market, and either a real Search Console baseline or a plain statement that none exists.

Two things tend to go wrong here. The first is mixing informational and buying-intent queries in the same comparison, which gives you a peer set that doesn't agree with itself. The second is treating every query in the report as something the page must cover. Search Console also leaves out some queries for privacy reasons, so the table is a useful sample and not a full list.

DeepSmith can help at this point as a starting signal. The Prompts view shows the buyer questions you track for a client, and the Pages view shows which of that client's pages AI engines already cite. Those signals help you choose which page and question are worth the effort. They don't replace looking at the actual search results for your chosen query, and DeepSmith does not calculate a TF-IDF comparison for a page, so that part stays with you.

Build a small set of comparable ranking pages

Search the query yourself, in the market and language your client cares about, and choose roughly five to ten pages that answer the same kind of question your page answers. That range is a practical suggestion and not a researched threshold, so pick comparability over quantity. Five pages that really match your intent are better than ten that mix definitions, product pages, and videos.

For each page you choose, record the query, market, language, collection date, page address, result type, and why you included it. Record the ones you skipped too, with a reason. If the results show several different intents, choose the set that matches your page and don't just take the first ten. Leave out duplicates, syndicated copies, and pages whose main content is a product listing or a video when your page is an explanatory article. Keep your own draft out of the peer set, because you want to compare it against the others and not blend it in.

You're done when someone else on the team could rebuild the same peer set from your notes and see why each page is there.

The usual mistakes are comparing a cost explainer against a pile of definitions, treating a brand's whole site as one document, and assuming the sample is everything Google considered. A page can rank for reasons that a text comparison cannot see, so hold your conclusions loosely.

Pull the main text from each page

Save your draft and each peer page as its own plain text file. Keep what answers the question: the headings, explanations, examples, lists, and any tables that carry meaning. Cut the navigation, cookie notices, banners, related-post blocks, comments, and footer text. Use the same inclusion rule for every page, or the comparison won't mean anything.

If you want to automate the collection, Trafilatura's Python interface has fetch_url to get a page and extract to pull out its main text, and it has options for excluding comments or tables. Be careful with the tables option. If a concept is explained inside a table, leaving tables out will hide the exact thing you're trying to find. Respect each site's access rules while you collect.

You're done when each file looks like the real page, and a spot check of a few headings and examples matches the original. If an extraction fails, write that down. An empty file will otherwise be counted as a very short article and quietly skew everything after it.

Watch for menu and footer words showing up in your results, for content that loads dynamically and never made it into the file, and for one page where you kept only the headings while the others are full articles.

Calculate peer coverage and TF-IDF candidates

Now you can run the numbers to find term frequency content gaps, and it helps to know what they mean first. Term frequency is how often a word appears in one document. Document frequency is how many of your peer pages contain the term at least once, so if four of five peers mention a concept, its document frequency is four. That's an observation about those five pages and nothing more. Inverse document frequency gives less weight to words found in many documents and more weight to words found in few. TF-IDF multiplies a term's frequency by that inverse weight, which surfaces terms that stand out inside your particular set.

In scikit-learn's default version, the inverse weight is ln[(1 + N)/(1 + DF)] + 1, where N is the number of documents and DF is how many contain the term. For a five-document set, a term found in four documents gets about 1.182, and a term found in only one gets about 2.099. That is the reason a rare one-off can outscore a concept every good page explains. A competitor's brand name or an odd anecdote can look important to the math and be useless to your reader.

Because of that, the worksheet you build should carry both signals. Make columns for the term or phrase, the number of peers containing it, the share of peers, a summary of the peer TF-IDF weight, whether your draft contains it, and the sentence around it in the draft if it does. Look first at phrases found across several peers and missing from your draft. Then look separately at high-weight terms found in only one or two peers, because those might be real differentiators, or irrelevant extras, or proper names.

Here is a script that does the calculation. Save the peer texts as UTF-8 files in a peers folder and your draft as draft.txt. It's an illustrative recipe, not a DeepSmith feature and not a scoring model that Google prescribes.

from pathlib import Path
from math import ceil
from sklearn.feature_extraction.text import TfidfVectorizer

peer_files = sorted(Path('peers').glob('*.txt'))
peer_texts = [p.read_text(encoding='utf-8') for p in peer_files]
draft = Path('draft.txt').read_text(encoding='utf-8')
assert peer_texts and all(peer_texts) and draft

vectorizer = TfidfVectorizer(
    lowercase=True,
    ngram_range=(1, 2),
    min_df=1,
    max_df=1.0,
    norm='l2'
)
peer_matrix = vectorizer.fit_transform(peer_texts)
draft_vector = vectorizer.transform([draft])
terms = vectorizer.get_feature_names_out()
peer_df = (peer_matrix > 0).sum(axis=0).A1
peer_mean = peer_matrix.mean(axis=0).A1
draft_present = (draft_vector > 0).toarray()[0]

# A review queue, not an instruction to insert every returned phrase.
review_floor = max(2, ceil(0.6 * len(peer_texts)))
rows = [
    (terms[i], int(peer_df[i]), round(float(peer_mean[i]), 4))
    for i in range(len(terms))
    if peer_df[i] >= review_floor and not draft_present[i]
]
for term, pages, mean_weight in sorted(
    rows, key=lambda row: (-row[1], -row[2], row[0])
):
    print(term, pages, mean_weight)

A few notes on how it behaves. It fits the vectorizer on the peer pages only, then transforms your draft with that same fitted vocabulary and weights. If you fit each page separately, the numbers can't be compared. It reads one-word and two-word phrases, since a bigram like advance rate keeps a meaning that the single words lose, though not every bigram it prints will be a real topic. The 60 percent review floor is just an editing filter. In a five-page set it puts anything found in at least three peers into your first review queue, and you can change it to suit your sample. Be careful with aggressive max_df filtering, because it can remove the central subject term just because every peer talks about it. Also, the default tokenizer can split technical labels that contain punctuation, so check those by hand.

You're done when you have a short list of missing candidates, each with a peer count and a note on how the peers use it.

Common mistake: Don't paste in the highest-scoring terms. A rare irrelevant phrase can have a high TF-IDF weight, while a common essential concept can have a low one, since it appears everywhere. Rank by peer coverage and relevance first, and use the weight as a second look.

Other ways this goes sideways include reporting the top TF-IDF value as the most important omission, mixing up a normalized score with a word count, aiming for a fixed keyword density, and assuming a zero means the page has no equivalent explanation in different words.

Check noun phrases and entities, then decide

The list from the last step only knows about spellings. Your page might already explain a concept using different words, and a plain word list will miss concepts that are longer phrases. This is where NLP content analysis earns its place. It's a broader set of language tools than TF-IDF, and for this job the useful output is a list of noun phrases and named entities pulled from the peers and from your draft, which you then read and judge yourself.

spaCy exposes these as Doc.noun_chunks and Doc.ents. Run both over the peers and your draft, and compare ideas and not just exact spellings. A named entity is a particular organization, place, or product. A phrase like advance rate can be an important subject-matter concept without being recognized as an entity at all, so you need the noun phrases and your own domain knowledge alongside the entity list.

Go through each shortlisted candidate and read the passages around it. Ask whether the concept is central to this reader's task, whether several appropriate peers explain it, whether your draft already explains it under another name, whether an example or named product or law is actually relevant to this client, and what evidence you could use for it.

Sort every candidate into one of five buckets: Missing explanation, Present under another name, Relevant but needs evidence, Irrelevant to page intent, and Competitor-specific or unverified. That last bucket matters for agencies in particular. If a peer names its own product, that name showing up in the entity list is not a reason to put it on your client's page, and it should never become the client's claim.

You're done when every candidate you keep points to a specific missing explanation, example, qualification, or answer, with a place on the page where it would go. Every candidate you reject should have a reason next to it.

The typical errors in NLP content analysis are calling every noun phrase an entity, treating a model's labels as fact, and adding the exact word while the reader's question stays unanswered. Entity recognition is statistical, so it can mislabel things or miss specialist terms.

Revise for information gain, not term counts

Now you can change the page, which is the part of content depth analysis where the value is created. Add only the sections, sentences, examples, or qualifications needed to resolve the gaps you accepted. Put each answer near the question it belongs to. Define specialist phrases in plain language, show how a process works, and include the conditions and exceptions a reader would want. Keep the client's verified product claims separate from general industry description, and don't copy a peer's structure or wording.

Here is a made-up example, and it is not a finding from a live results page. Say a draft about invoice-factoring costs already defines factoring. Four of five comparable cost explainers discuss an advance rate, and three discuss recourse. A real revision would explain what each concept means and how it changes a cost comparison, using verified figures or clearly framed examples if you have them. Dropping the two phrases into an existing paragraph doesn't close the gap. And if the page is really a basic definition page, a detailed costing section might belong somewhere else.

You're done when a reviewer can point to the new answer or clarification a reader now gets, and can check that it's accurate. The page should also read naturally, without the same phrase repeating in a way you can hear.

The common trap is writing to match another page's word count or term target. Google's guidance says there is no magic word count, warns against repeating phrases in ways that sound unnatural, and asks whether content adds substantial value instead of rewriting what other sources already say.

This is also a step where DeepSmith fits, as a way to produce the revised copy. The Writer turns a planned idea into a brand-grounded article with research, links, imagery, and metadata, and Deep IQ holds each client's positioning, product facts, and voice, so a revision for one client doesn't pick up another client's wording. The strategist still runs and checks the specific TF-IDF and NLP comparison from this guide, and DeepSmith doesn't claim that a revision will rank.

Rerun the comparison and measure the real outcome

Extract the revised body using the same rules as before. Then transform it with the original fitted peer vectorizer, the one you already built, and compare your shortlisted phrases and the explanations around them, before and after. Keep a decision log with the source dates, the original and revised text, and the client's approval.

After the page is published and has had enough time, go back to Search Console with the same page filter and compare a sensible date range. Look at the relevant queries, impressions, and clicks. Think about query mix, other edits to the page, and changes in the results before you draw a conclusion. If your original peer sample has aged, check the live results again.

You're done when each accepted gap is either resolved or openly declined, the client has approved the revised page, and there is a measurement plan for that page. If performance moves later, report it as something you observed and don't say the term additions caused it.

Watch out for treating a higher overlap score as proof of a ranking gain, declaring success the day you publish, or comparing mismatched time periods. The same care applies to AI citations. DeepSmith's Pages and Prompts views show which pages earn citations and which tracked prompts drove them, and they tell citations apart from plain mentions, so they work as a separate AI-visibility check. Google's guidance for AI Overviews and AI Mode doesn't set a TF-IDF threshold, and it says that a page needs to be indexed and eligible for a snippet to appear as a supporting link. A finished term audit doesn't guarantee a citation.

Hand off a revision brief

Your TF-IDF SEO worksheet turns into something a writer and a client can both read. This is one example row of a brief, and every entry here is illustrative.

FieldExample entry
TargetOne client's existing invoice-factoring cost explainer
QuestionWhat determines the cost of invoice factoring?
Comparison setFive relevant, accessible cost explainers, with date and market recorded
Candidateadvance rate
Observed peer coverageFour of five sampled pages, which is not an industry benchmark
Current draftPhrase and underlying explanation absent
Editorial actionExplain what the advance rate means and why it matters in the cost example
Evidence checkVerify the client's actual offer and any numbers before publication
Acceptance testA reader can explain the concept after reading, and the term reads naturally
MeasurementSave both versions and watch the page's filtered Search Console queries later

If you turn this into a repeatable service, use the same worksheet fields and quality checks for every client, and keep each client's query, peer set, language, product facts, and approvals separate. What you standardize is the decision process, and the copy stays different for each account.

What to do next

Pick one page from a client's list that already gets some impressions, and run steps one through five on it this week. Stop there and look at your candidate list before you write anything. If most of the kept candidates point to real missing explanations, go on to the revision. If they mostly point to synonyms and brand names, the page probably doesn't have a depth gap, and it's fine to leave it alone.

Once you have one brief done, you'll have a template for the rest of the roster. If you'd like a place to keep each client's brand context and content plan while you do this, you can start a free DeepSmith trial and see how the Writer and Deep IQ handle the revision step.

Frequently asked questions

Is TF-IDF a Google ranking factor or a score I need to hit?

No. The method here analyzes a comparison set that you choose, and its output belongs to that set. Don't present it to a client as a score Google issues or as a target that will lift rankings. Use it to find candidate explanations that an editor then reviews.

How many competitor pages and which terms should I compare?

Start with around five to ten pages that serve the same intent, language, and audience. That's a workflow suggestion and not an official threshold. Compare meaningful single words and two-word phrases along with how many peers use them and in what context, and don't add a term just because your count is lower.

What is the difference between TF-IDF and NLP entity analysis?

TF-IDF weights words by how often they appear in a document and how rare they are across your chosen set. An NLP pass can pull out noun phrases and predicted named entities. Used together they help you propose gaps, but neither one decides whether a client's page should make a particular claim.

When is a content gap actually closed?

When the revised page accurately and clearly answers a reader question that was under-explained before. A term appearing, a higher similarity score, or a longer page doesn't meet that test on its own.