DeepSmith

Jul 26 · AEO & AI Visibility

17 min read

How to Measure AI Visibility Manually Without a Paid Tool

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
Monochrome charcoal cover with the centered white line 'Measure AI Visibility by Hand' over abstract white-and-gray linework of a spreadsheet grid, connected prompt nodes, and citation marks pointing to a search-and-answer motif.

Your leadership just asked where the brand shows up in ChatGPT, and you don't have an answer yet. That's normal, and you can fix it this week. You can measure AI visibility manually with nothing more than a spreadsheet, a browser, and a couple focused hours.

Here's what you'll walk away with: a fixed set of buyer prompts, a repeatable way to run them across the AI engines, a ledger that logs every mention and citation, and real baseline numbers for your team. This is free AI visibility tracking done by hand, the honest starting point before you ever pay for a tool.

Let's build it one step at a time.

Step 1: Write down your measurement rules first

Before you open a single AI engine, decide what counts. This is the step everyone skips, and it saves you later.

Write these rules into a plain text note or the first tab of your sheet:

  • A brand is mentioned when the answer names it or clearly identifies it. Mark that 1, otherwise 0.
  • A brand is cited only when the answer links, footnotes, or attributes a claim to a page on the brand's own domain. Mark that 1.
  • If the brand is named but no first-party source appears, citation is 0, not 1.
  • Log third-party pages separately. A review site or directory that mentions you is a third-party source, not a citation to your site.
  • When an engine shows no source links at all, mark the citation field NA, never a forced 0.
  • Count each brand once per answer for mention and share-of-voice math, even if the name shows up three times.
  • Code sentiment and accuracy on their own. A glowing answer can still get your pricing wrong.

Decide your main denominator too. The cleanest choice is eligible answered runs: drop technical failures, empty answers, and surfaces that can't show citations, but report how many you dropped.

Done when: someone else could read your rules and code the same answer exactly the way you would.

Common mistake: treating every source link in an answer as a citation for you. A competitor's page or an independent review can discuss your brand without being your citation. Keeping mention rate and citation rate separate is the heart of honest manual AEO measurement.

Step 2: Map your buyer prompt set

Now build the questions. Start with 20 to 40 high-intent prompts. That's enough to cover a focused brand or product line without drowning your first audit.

Build these around how buyers actually talk, not just your SEO keywords. Pull from sales calls, support tickets, customer interviews, Search Console queries, and the objections you hear on repeat.

Then split your prompts into three buyer stages.

Awareness prompts test whether the category even knows you exist:

  • What are the best tools for [category]?
  • What's the best way to solve [problem]?
  • Which platforms help [audience] achieve [goal]?

Consideration prompts test comparisons and fit:

  • How does [Brand] compare with [Competitor] for [use case]?
  • What are the best alternatives to [Competitor]?
  • What are the strengths and weaknesses of [Brand]?

Decision prompts test purchase readiness:

  • Is [Brand] a good choice for [specific goal]?
  • Who should and should not use [Brand]?
  • Does [Brand] support [specific requirement]?

Mix generic category prompts, use-case prompts, comparison prompts, and branded validation prompts. Resist the urge to make the whole set branded. A brand shows up in a branded query because you handed the model its name, and that proves nothing about real discovery.

Give every prompt a stable ID, a buyer stage, a topic cluster, and an intent type. Then freeze the wording for the whole reporting period. If a prompt needs a fix, create a new version and log the change instead of quietly swapping it.

Done when: every prompt has an ID, a stage, a cluster, an intent, and a reason it's on the list.

Common mistake: picking only the questions where you already expect to win. That gives you a flattering baseline and a weak strategy.

Prompt discovery is the hard part, and it also automates well later. DeepSmith can generate a starter prompt set from your product, persona, and buyer-stage context, so you're not staring at a blank sheet. For your manual baseline, though, keep owning the prompt logic and labels. That's the asset you carry forward.

Step 3: Freeze your engines and test conditions

Pick your engines and lock your conditions. Consistency is what makes next month's numbers comparable to this month's.

For a full baseline, run the same prompt set across five surfaces:

  • ChatGPT
  • Perplexity
  • Gemini
  • Claude
  • Google AI Mode, including the AI search surface available in your region

Short on time? A lean first pass can start with two or three engines. Just record which ones, because you can't compare a five-engine month to a two-engine month.

Hold these conditions steady on every run:

  • Use the same account and settings each cycle where you can.
  • Keep the same language, region, device type, and browser state.
  • Record the model or surface name when it's visible.
  • Open a fresh session for every single prompt.
  • Skip follow-up questions that could reshape the answer.
  • Log the date and cycle for every run.

Do not collapse every Google AI surface into one result. Record the exact surface you used, since presentation shifts by region, account, and date.

Done when: your engine list, locale, session rules, and citation-availability rule are all written in the workbook.

Common mistake: carrying context from one prompt into the next. A previous question or a correction you typed can nudge later answers and make your results look better than reality.

Step 4: Build your tracking spreadsheet

This is the backbone of your system, and simpler than it sounds. Use separate tabs so prompt definitions, raw runs, citations, and metrics never bleed together.

Copy these headers directly.

Prompt Registry tab:

Prompt ID | Prompt text | Buyer stage | Topic cluster | Intent type | Brand | Competitor A | Competitor B | Prompt version | Active | Inclusion rationale

Run Log tab (one row per prompt, engine, and repetition):

Run ID | Run date | Cycle | Engine | Surface or model | Prompt ID | Prompt text | Buyer stage | Topic cluster | Brand mentioned | Brand cited | Cited URL(s) | Brand cited page | Competitor A mentioned | Competitor A cited | Competitor B mentioned | Competitor B cited | Mention position | Sentiment | Accuracy | Citation surface available | Raw response reference | Notes

Citation Audit tab (one row per cited source, since one answer can have several):

Run ID | Engine | Prompt ID | Citation order | Source title | Cited URL | Source owner | First-party or third-party | Source type | Brand or competitor associated | Relevant claim | Citation useful | Notes

Metrics tab (one row per engine and cycle):

Engine | Cycle | Prompt runs | Eligible citation runs | Mention rate | Citation rate | Citation rate among mentions | Mention SOV | Citation SOV | Accuracy rate | Positive rate | Top cited pages | Top competitor | Interpretation | Action owner

Change Log tab:

Date | Change type | Description | Expected impact | Person responsible

Use the Change Log for anything that could break comparability: model updates, account changes, region changes, prompt edits, brand-name changes, big website updates, and any change to your citation rule.

Use 1 and 0 for the binary fields, and a consistent NA when a surface can't show citations. Don't leave that distinction to a free-form note.

Done when: the workbook can take a raw answer without you inventing a new column every time something odd shows up.

Common mistake: keeping only the final percentage. Without the raw answer, the cited page, the date, and your coding notes, that number can't be audited or explained to anyone.

Step 5: Run your baseline the same way every time

Time to collect. Enter each prompt into each engine one at a time, in a fresh session, and save the full answer and its source list before you move on.

Follow the same seven moves on every run:

  1. Copy the exact prompt from your registry.
  2. Confirm the right engine and surface.
  3. Start a clean session.
  4. Submit the prompt with no extra context.
  5. Save the response and every visible source link.
  6. Record the model, date, and anything unusual.
  7. Code the row only after the raw answer is saved.

Here's the honest math. Twenty prompts across five engines is 100 runs. Forty prompts is 200 runs. A single 40-prompt, five-engine pass takes roughly two to four hours of manual work. If you repeat every prompt three times, you triple that.

Why repeat? A single answer is a snapshot, not a fact. AI responses are probabilistic, so the same prompt can return different recommendations and citations. Run key prompts three times, and your result becomes 0, 1, 2, or 3 of 3. Keep every repetition, not just the average. That's how you track AI citations without a tool and still trust what you see.

Put a recurring reminder on the cadence you choose. Monthly suits a general program. Weekly makes sense when you're actively publishing to earn citations. Can't run the full set weekly? Keep a fixed core set and label rotating prompts separately.

Done when: every planned prompt-engine pair has its expected runs, and any missing or failed run is marked, not silently dropped.

Common mistake: changing a prompt after you see an answer you don't like. That turns measurement into optimization and makes the whole period impossible to compare. This is the discipline that separates DIY AI visibility from wishful thinking.

When copy-and-paste across five engines starts eating your week, that's the signal to automate. DeepSmith runs scheduled collection across your configured engines, so the runs happen without anyone sitting in the app. You reach for it when manual runs cost more than they're worth, not before.

Step 6: Code each answer for mentions, citations, competitors, and quality

Now apply the rules you wrote in Step 1. This is where careful beats fast.

Mentions: mark Brand mentioned = 1 when the answer names or clearly identifies you. If the answer is a ranked list, note whether you're first, in the middle, or tucked in as an afterthought.

Citations: mark Brand cited = 1 only when a source points to a page you own, and log every cited page in the Citation Audit tab. A few examples make this concrete:

  • Brand named, no source link: mention yes, brand citation no.
  • Brand named, your product page linked: mention yes, brand citation yes.
  • Brand named, an independent review linked: mention yes, first-party citation no, third-party source logged.

Competitors: track at least two core competitors, and log each one's mention and citation separately. Keep the competitor list stable across cycles. If it changes, start a new reporting series.

Sentiment and accuracy: use a simple fixed vocabulary. Positive, neutral, negative, incorrect, or not applicable. Check accuracy across category, capabilities, audience, pricing, integrations, and limitations. When something's wrong, paste the exact inaccurate sentence into your notes so your content or product team can chase it down.

Done when: every result has a mention value, a citation value or NA, cited-source detail, competitor values, and a quality note where it matters.

Common mistake: celebrating a positive mention without checking whether the description of your product is even accurate.

Step 7: Calculate your baseline metrics

Here's where the ledger becomes a report. Calculate by engine, cycle, buyer stage, and cluster wherever your sample size holds up. Always show the numerator and denominator.

Say your Run Log puts the engine in column D, cycle in column C, brand mentioned in column J, and brand cited in column K. A basic mention-rate formula is:

=IFERROR(SUMIFS('Run Log'!$J:$J,'Run Log'!$D:$D,$A2,'Run Log'!$C:$C,$B2)/COUNTIFS('Run Log'!$D:$D,$A2,'Run Log'!$C:$C,$B2),0)

A citation-rate formula that excludes the NA rows is:

=IFERROR(SUMIFS('Run Log'!$K:$K,'Run Log'!$D:$D,$A2,'Run Log'!$C:$C,$B2)/COUNTIFS('Run Log'!$D:$D,$A2,'Run Log'!$C:$C,$B2,'Run Log'!$K:$K,">=0"),0)

For two competitors in columns N and P, a mention share-of-voice formula is:

=IFERROR(SUMIFS('Run Log'!$J:$J,'Run Log'!$D:$D,$A2,'Run Log'!$C:$C,$B2)/(SUMIFS('Run Log'!$J:$J,'Run Log'!$D:$D,$A2,'Run Log'!$C:$C,$B2)+SUMIFS('Run Log'!$N:$N,'Run Log'!$D:$D,$A2,'Run Log'!$C:$C,$B2)+SUMIFS('Run Log'!$P:$P,'Run Log'!$D:$D,$A2,'Run Log'!$C:$C,$B2)),0)

Add one column per extra competitor. If no tracked brand appears in the answer set, the share-of-voice denominator is zero. Report that as not applicable, not as zero visibility.

For prompts you repeated three times, compute mention frequency, citation frequency, and stability: is the binary result the same across all three runs, or does it flicker? A flickering result is a real finding.

Read your numbers as signals, never as traffic forecasts:

  • High mention rate with low citation rate means the model knows you but isn't using you as a source.
  • Low mention and low citation can mean weak category relevance, thin coverage, or a prompt set that doesn't match your real market.
  • High citation rate with poor accuracy means the model found your site but is reading it wrong.
  • A strong result in one engine tells you nothing about another engine.

Compare like with like. Don't hold a monthly full set against a weekly subset without labeling the difference.

Done when: the workbook shows rates by engine and cycle, includes sample sizes, and every notable percentage points to a question or an action.

Common mistake: reporting a percentage with no denominator. A 70 percent rate from 10 runs and a 70 percent rate from 200 runs do not deserve the same confidence.

Once these formulas start feeling repetitive, that's the work a platform absorbs. DeepSmith's AEO overview computes mention rate, citation rate, share of voice, trends, and per-platform breakdowns once your prompt set is configured, so you interpret instead of recalculate.

Step 8: Audit your citation gaps and set your cadence

Numbers matter only when they change what you do next. For every prompt where a competitor wins a citation, open the cited page and study it.

Record the competitor and exact page, the prompt and its buyer stage, the source type, the claim the page supports, and whether you have an equivalent page. Then name the gap: is it missing content, unclear positioning, outdated information, weak third-party coverage, or an engine-specific quirk?

Turn those findings into an action backlog with this header:

Prompt ID | Engine | Winning source | Gap type | Proposed page or update | Evidence needed | Owner | Planned date | Follow-up cycle | Result

One honest caveat: publishing a page does not guarantee a citation. This audit surfaces plausible gaps, not cause and effect. Look for repeated patterns across prompts and cycles before you rebuild anything.

Then lock your cadence:

  • Initial baseline: one setup plus three repetitions where feasible.
  • General monitoring: monthly, same prompt set, compared to baseline.
  • Active publishing: weekly on a fixed core set.
  • Big product or site change: a targeted follow-up plus the next regular cycle.
  • Prompt or engine change: a new labeled series, kept separate from old history.

Track a fixed set for at least 30 days before you draw strong trend conclusions. That's a practical floor, not a promise of statistical certainty.

Done when: each cycle produces both a metrics report and an action list tied to specific prompts and pages.

Common mistake: collecting share of voice and citation rates and then never using them to pick a page, an update, or an experiment. When competitor gap analysis becomes the slow part, DeepSmith's competitor-citation views can show which competitor pages win your prompts and which prompts drive them, which is the piece that eats the most manual hours.

Where manual measurement breaks down

You should trust this method, and you should also know its limits. Being honest about them is what keeps your manual AEO measurement credible.

Time cost. Two to four hours per full cycle adds up, and three repetitions multiply it. Weekly full coverage can quietly become a second job.

Non-determinism. The same prompt can return different answers. Repeated runs reveal stability, but never remove all the variation.

Data decay. Models, indexes, and websites shift. A report can go stale while you're still assembling it, so always stamp the collection date.

Sampling bias. Twenty prompts are not your whole market. Too many branded prompts make you look far more visible than you are.

Engine differences. Some surfaces expose citations more openly than others. Report each engine on its own, never averaged into one blurry score.

Coding errors. Manual copying drops sources and misclassifies citations. Preserve raw answers and have a second person spot-check a sample.

No business attribution. Mention and citation rates measure appearance in AI answers, not traffic, meetings, or revenue. Connect those to your analytics separately.

None of this makes the manual method wrong. It makes it a baseline you understand, which is worth far more than an opaque score you can't question.

What to do next

You now have a complete DIY AI visibility system: rules, prompts, engines, a workbook, formulas, and a cadence. Start small. Run one baseline this month, even with three engines and 20 prompts, and your free AI visibility tracking will already tell you more than most competitors know.

Manual measurement lets you track AI citations without a tool, and it works right up until scale fights back. When the same prompts have to run weekly, when three repetitions become impractical, when you need answer-level history and page-level attribution, that's when a spreadsheet stops paying for itself. The clean move is from a documented manual baseline into automation, keeping your prompt IDs and coding rules intact so the two can be compared.

That's exactly the handoff DeepSmith is built for. It runs the prompt set on a schedule across ChatGPT, Perplexity, Gemini, Claude, and Google AI Mode, computes the same metrics, and tracks which competitor pages win your citations. When you're ready to stop copying and pasting, you can start a free DeepSmith trial and carry your baseline straight in.

You've got this. One prompt set, one honest baseline, one step at a time.

Frequently asked questions

How many prompts do I need to measure AI visibility manually?

Use 20 to 40 high-intent prompts for a focused baseline. Spread them across awareness, consideration, and decision stages and your main topic clusters. Bigger markets need more, but hundreds make manual collection impractical fast.

How often should I run the audit?

Run a general baseline monthly. Move to weekly when you're actively publishing content meant to earn AI citations. Keep the prompt wording and engine settings stable so each cycle is comparable to the last.

Why does the same prompt produce different answers?

AI answers are probabilistic. They shift with model updates, retrieval changes, account context, region, and random variation. Use a fresh session for every prompt and repeat important ones three times, so your report shows frequency and stability, not one lucky answer.

What's the difference between a mention and a citation?

A mention is your brand name appearing in the answer. A citation is a link, footnote, or source attribution pointing to a page you own. If the answer names you but links only to a competitor or a review site, log a brand mention and a third-party source, not a first-party citation.