DeepSmith

Jul 26 · AEO & AI Visibility

16 min read

How to Measure How Often AI Engines Cite Your Brand (Citation Rate)

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
A monochrome, abstract-geometric cover on charcoal showing a central answer card linked to small source nodes with a fraction motif, under the white cover line 'Measure Your AI Citation Rate'.

You searched your own category in ChatGPT, saw a competitor cited instead of you, and thought: how often does AI cite my brand, really? That question deserves a number, not a gut feeling. The good news is that your AI citation rate is a metric you can define, calculate, and re-run on a schedule, and you can build the first version this week.

This guide walks you through it one step at a time. By the end you will have a fixed set of prompts, a written scoring rule, a collection sheet with real answers saved, and a baseline AI citation rate you can compare against next month. Knowing how to measure citation rate is not about clever tools. It is about the numerator, the denominator, and the discipline to keep them stable.

Let's start with the metric itself, because everything downstream depends on getting this one definition right.

Step 1: Define the citation rate metric before you count anything

To turn "how often does AI cite my brand" into a countable rule, you need one precise definition. Here is the whole thing in one line. Your brand citation rate is the share of valid prompt runs where the AI answer included at least one qualifying source from your own domain.

Written as a formula:

Brand citation rate = (valid prompt runs with at least one qualifying brand citation) ÷ (total valid prompt runs) × 100

Three words in that sentence carry all the weight, so let's slow down on each.

A prompt run is one execution of one saved prompt in one engine. Run 50 prompts once in ChatGPT and you have 50 observations. Run the same 50 across three engines and you have 150 engine-prompt observations, but each engine still has its own denominator of 50.

The numerator counts runs where the answer showed at least one source that matches your domain rules. It counts runs, not links. If a single answer cites three of your pages, that is still one positive run. Save the three URLs for page-level analysis, but the binary metric moves by one.

The denominator is valid runs only. A valid run produced a real, scorable answer. Timeouts, collection failures, and empty responses are invalid. Do not silently delete them, and do not count them as zero-citation runs. Invalid is its own bucket.

So a run has three possible outcomes: positive (a qualifying brand source appeared), zero (a valid answer with no qualifying brand source), or invalid (no scorable answer at all). That three-way split is the backbone of the whole measurement.

A quick worked example. If 14 of 50 valid runs cited at least one of your pages, your rate is 14 ÷ 50 × 100, or 28 percent. That is an illustration, not a benchmark. There is no universally "good" citation rate, so resist the urge to compare your first number to anyone else's.

This is the citation rate metric AEO reporting is built on, and it is deliberately narrow. It measures whether AI pointed a reader to your site as a source, not whether AI happened to say your name.

Step 2: Write down exactly what counts as a citation

A citation is a visible source attribution, not a name-drop. Your brand can appear in the body of an answer with no link at all, and under this metric that does not count. Keeping mentions and citations separate is what makes the number trustworthy.

Write a rule a second person could apply without asking you to explain it. Something like this: a run counts as a brand citation when the answer attributes information through a clickable source link, a source card, a footnote, or a numbered reference, and that source URL resolves to a domain your rules allow.

Counts as a citation:

  • A source card or inline link pointing to a page on your allowed domain.
  • A footnote or numbered source that resolves to your allowed URL.
  • A redirected or canonicalized URL that lands on your domain.

Does not count:

  • Your brand name in the answer text with no source attached.
  • "Company X is one example" with no link or reference.
  • A third-party review, directory, or marketplace page that mentions you, when your rule only allows your first-party domain.
  • A bare URL sitting in the body that was not presented as a source.

Now lock your domain scope, in writing, before you collect anything. Decide whether your subdomains count: docs, help center, blog, support. Decide about country domains, product-specific domains, and pages you host on a publisher or marketplace. Decide how you will normalize tracking parameters, redirects, capitalization, and trailing slashes so the same page is never counted as two.

Common mistake: expanding the allowed-domain list after you see disappointing results. Do not do that. If your rules change, you start a new baseline or you annotate the change clearly. Quietly widening the net breaks every comparison you make later.

Step 3: Build a fixed set of buyer prompts

Your prompts are your denominator, so they matter more than almost anything else. Build them from real buyer questions, not a keyword list you exported once.

Cover the journey. You want informational questions, pain-point and use-case questions, research and comparison questions, "alternatives to" questions, and vendor-evaluation questions people ask right before they decide. Reaching for only your brand name is the classic trap, because a brand-name query can only tell you how AI behaves when someone already knows you exist. The prompts buyers actually ask rarely name you at all.

For each prompt, save six things: a stable ID, the exact text, the intent or buyer stage, the topic, a branded or unbranded tag, and a one-line reason it belongs in the set.

How many prompts? There is no universal answer, and anyone who gives you one is guessing. The operational ranges people use run from roughly 7 to 10 per topic area, to a small starter baseline of 10 to 15 queries, to 100 to 200 seed prompts for broad programs. A practical starting point for a focused brand is around 50 stable prompts, but only if your team can genuinely re-run them every month. A maintainable set you actually repeat beats a huge set you abandon.

How you know this step is done: every prompt has its ID, its exact text, its intent, its topic, and its tag. A teammate could open the sheet and run the exact same set without asking you a single question.

One more rule that saves you pain later. When you think of new prompts, put them in a separate discovery set first. Let them build history there. Adding prompts straight into your core set changes the denominator and quietly ruins your trend line.

Step 4: Lock your engines and collection conditions

The same prompt can return different sources depending on where and how you ask it. So before you run anything, decide and record the conditions, and hold them steady every time.

For each collection, write down the engine and the specific product experience, the model or version if it is visible, whether web browsing or search was on or off, the country and language, the date and time, whether you were logged in, and any personalization or memory settings. Use a clean or isolated session so your own history does not color the results.

That last point matters more than it sounds. An answer shaped by your logged-in history is not comparable to a clean-session answer, and mixing the two is how people accidentally "measure" their own browsing habits instead of the engine.

Pick your engine coverage based on where your buyers actually are. To measure AI citations across ChatGPT, Perplexity, Gemini, Claude, and Google AI Mode is ideal, but a focused two-engine baseline you maintain beats a five-engine one you run once and drop.

Common mistake: running ChatGPT with browsing on one week and off the next, then putting both in the same trend. Those are different measurement surfaces. Keep them apart, or you will read a settings change as a performance change.

How you know this step is done: your protocol names the exact engine mode being tested, and another analyst could reproduce one run from your notes alone.

Step 5: Run every prompt and save the raw answers

Now you collect. The unit here is one prompt-engine run per row, not one brand or one page or one day. That row-per-run discipline is what lets you audit the number later.

For each run, record the run status, the full answer text, every visible source, the normalized source URLs, whether a qualifying brand citation was present, an invalid reason if it failed, and the model, mode, time, and region. Yes, save the whole answer. A percentage with no underlying evidence cannot be checked after an interface or model update, and those updates happen constantly.

For a small pilot, a spreadsheet is genuinely fine. It is transparent, it costs nothing, and it forces you to look at real answers while you validate your scoring rule. Start there with no guilt.

The strain shows up as you scale. Manual collection tends to get unreliable somewhere around 30 queries, especially once you add engines, repeats, and a monthly cadence. Saving every raw answer, normalizing URLs by hand, and never dropping a failed run becomes a second job.

That is the point where a tracking platform earns its place. A tool like DeepSmith runs your saved prompts on a schedule, keeps the full answer history, normalizes source links, and splits results by prompt, engine, and topic, so the raw evidence is captured for you instead of by you. It does not replace your definition or your domain rules. It operationalizes them. You still decide what counts; the platform just does the repetitive collection without missing a run.

How you know this step is done: every expected prompt-engine run is either a scorable answer or an invalid status with a stated reason. Nothing is missing and nothing is quietly blank.

Step 6: Score each answer with one binary rule

Scoring is where consistency is won or lost, so keep it boring. For every valid run, mark a single field, brand_citation, as Yes or No. Mark failed runs as Invalid, in their own column, never as No.

Alongside that binary field, keep a column of normalized source URLs and an optional citation_count for the number of distinct qualifying brand URLs. The count is a useful secondary field for page-level work. It is never a substitute for the binary numerator.

A few columns make the whole sheet auditable: collection date and time, engine and mode, prompt ID and exact text, topic and intent, branded or unbranded tag, run or repeat ID, answer status, the Yes/No citation flag, the URL list, the canonical domain, and a notes field for odd interfaces or ambiguous sources.

Pro tip: the real test of your scoring rule is agreement. Have a second person spot-check a sample of runs. If they land on the same Yes, No, or Invalid as you did, your rule is solid. If they hesitate, your rule is too vague, and you tighten it before you trust any number.

Watch for the scoring traps here. Counting a third-party review's link as a first-party citation. Counting the same URL twice in the binary metric because it appeared twice in one answer. Treating a name mention as a source link. Using different allowed-domain rules for different engines. Each of these quietly inflates or deflates your rate, and each is avoidable with one written rule applied everywhere.

Step 7: Calculate your baseline and set a re-test cadence

Now the number falls out on its own. For each engine and each prompt group:

Brand citation rate = positive valid runs ÷ total valid runs × 100

Report it in a small table, per engine, so nothing hides in an average:

EnginePositive runsValid runsCitation rate
Engine A145028%
Engine B95018%
Engine C65012%

Then repeat the table by intent group, so informational, comparison, and evaluation prompts each get their own line. A brand can look invisible on high-intent buyer questions while looking fine on soft informational ones, and a blended average would hide exactly the gap you care about.

Show your working next to every percentage: the numerator, the denominator, the date range, the engine, the prompt-set version, the domain rule, and your invalid policy. This transparency is what separates a real citation rate metric AEO teams can defend from a screenshot someone will question. Thirty percent from 10 prompts and 30 percent from 100 prompts are not equal evidence, and an honest report makes that visible.

Resist blending engines into one headline number too early. If you must produce an overall figure, be explicit about how: either pool all positive runs over all valid runs, or average the per-engine percentages. Never add percentages together, and never swap a link count in for the binary numerator.

Then set your cadence. Monthly re-testing is the simple, defensible baseline rhythm. Check volatile or high-value prompts more often if you like, but keep the prompt set, engine modes, region, domain rules, and repeat policy identical each time. Keep a change log of anything you alter. If an engine changes how it shows sources, mark a break in your trend line instead of blaming or crediting your own content. This is how to measure citation rate over time without fooling yourself, and it turns a one-off audit into a living dashboard.

How you know this step is done: your latest report lines up cleanly against the last one, and anything that is not comparable is flagged rather than buried.

Measure each AI engine on its own

One habit will keep your numbers honest: never assume two engines are directly comparable. Each one retrieves and displays sources differently, so report them side by side, not merged.

ChatGPT citation behavior depends heavily on whether web search fired. Answers with no browsing should not sit in the same pool as web-connected ones. Record the browsing state every time. One study observed that only a fraction of ChatGPT conversations triggered a web search at all, which is a useful reminder that "no source shown" is often a mode difference, not a real zero.

Perplexity puts sources front and center, but the sources it reads are not always the ones it visibly cites. Score only what is visible and passes your domain rule. A citation-heavy interface does not automatically mean a higher brand citation rate for you.

Gemini shows citations through its own interface and retrieval path. Report it on its own and do not assume its percentage maps onto ChatGPT's.

Google AI Overviews and Google AI Mode are two separate surfaces, even when the answers read alike. They can return different URLs for the same question, so measure them independently rather than treating them as one Google number.

Claude should be scored with the same domain rule when the experience you test surfaces source links. If the mode you are using does not retrieve or show sources, record that accurately instead of logging it as a zero.

Read your multi-engine results as a set of baselines that are each comparable within their own engine, not as one global score. That framing keeps you from drawing conclusions the data cannot support.

What to do next

Take a breath. You now have the whole method: a locked definition, a written citation rule, a fixed prompt set, controlled conditions, saved raw answers, consistent scoring, and a per-engine baseline with a cadence. That is more rigor than most brands ever apply to AI search.

Your next move is small. Save the prompt set, the protocol, the scoring rule, and the raw answers somewhere your team can find them, then schedule the next collection. Do not jump to optimizing pages yet. First you need two data points to compare, because a single reading is a snapshot, not a trend.

If the manual version starts eating your week, that is your signal to let a platform carry the collection. DeepSmith can run your prompts on a schedule, keep the full answer history, report citation rate by platform and by prompt, and attribute citations down to the exact page, so you spend your time reading the trend instead of assembling it. You can start with a free trial, point it at your real prompts, and see real answers before you commit.

You do not need a perfect measurement system. You need a repeatable one, and you just built it.

Frequently asked questions

Is citation rate the same as how often AI mentions my brand?

No, and the difference is the whole point. Citation rate requires a visible source attribution, usually a link to a page on your allowed domain. A mention is just your name appearing in the answer text, with no source behind it. You can be mentioned constantly and cited rarely, which is why the two are tracked as separate metrics.

One answer cited three of my pages. How many citations is that?

One. In the binary brand citation rate, a run counts as a single positive whether it cites one of your pages or five. Save all three URLs for page-level attribution and your optional citation count, but the headline metric moves by one run so that answers with more links do not get unfair weight.

How many prompts do I need to measure AI citations reliably?

There is no universal minimum. Your set should cover the main buyer questions and topic groups and stay stable over time. A focused brand can begin with a maintainable core of stable prompts and grow a larger seed pool for broad categories. Whatever size you choose, always report the prompt count, the groups, and the numerator and denominator so the strength of the evidence is clear.

How often should I re-measure?

Monthly is a sensible baseline rhythm for a trend you can trust. You can check important or volatile prompts more often, as long as you hold the prompt set, engine mode, region, domain rules, and repeat policy constant. If your prompts, domain rules, or an engine's behavior change meaningfully, mark a break in the trend or start a fresh baseline rather than pretending the old and new numbers are comparable.