Your boss just asked what your share of voice is inside ChatGPT, and you froze. That is normal. Almost nobody has a clean answer yet, because the metric is new and most guides make it sound harder than it is. Here is the good news: ai share of voice is a percentage, and by the end of this guide you will know exactly how to calculate share of voice AI answers give your brand, defend it, and explain it without hand-waving.
You do not need a data science team. You need one formula, a fixed set of questions, a clear list of competitors, and a little patience with messy answers. Let's build it together, one step at a time.
Start with the one formula that does the work
Before the steps, get the core idea in your head, because everything else is just feeding this line.
AI share of voice is the percentage of brand mention or citation events your brand gets out of all such events for a defined set of brands, measured against a fixed set of prompts run on one or more AI engines. That is the whole definition. Sit with it for a second.
Now the aeo share of voice formula, in two flavors:
- Mention-based: (your brand's mention events divided by the total mention events for your comparison set) times 100.
- Citation-based: (your brand's citation events divided by the total citation events for your comparison set) times 100.
A mention event is one appearance of your brand in a single answer, no matter how many times the name repeats in that answer. A citation event is one time the engine links to a page you own as a source. They measure different things. Mention share asks how often the engine brings you into the conversation. Citation share asks how often it trusts your pages enough to point people at them. Keep them separate. Never average the two.
Here is the part people trip over, so read it twice. Share of voice is not the same as your mention rate. Mention rate divides your mentions by the number of prompts you asked. If five brands show up in one answer, five mention rates can add up to way more than 100 percent. Share of voice divides by the total for your comparison set instead, so the whole set sums to exactly 100 percent. That is why it is a share, not a rate. This one distinction is the difference between a number you can trust and a number that quietly lies to you.
That sums-to-100 property is also your simplest sanity check. If you calculate share of voice in ai answers for every brand in a closed set and the percentages do not add up to 100, you counted something wrong.
Feeling clearer? Good. Now let's turn that formula into a repeatable process you can run every quarter.
Step 1: Pick your variant and counting rule
Two small decisions here, and both need to be written down in one sentence a colleague could follow.
First, decide whether you are scoring mentions or citations. If you are just starting, use mentions. They are easier to collect and easier to explain.
Second, decide how you count. Answer-level counting means a brand scores one event per answer, no matter how often it is named. Literal counting means you count every single time the name appears in the text. Answer-level is the calmer, more stable default, so start there.
You will know this step is done when you can hand someone a single sentence like this: "We score answer-level mention events across our tracked prompts." That is it.
Where people go wrong: they quietly switch rules when a number looks bad. Count Brand A one way and Brand B another, and your whole comparison falls apart. Pick one rule and apply it to every brand, every time.
Step 2: Build and freeze your prompt set
Your prompt set is the sample. It is the list of questions you will run against the engines, and it decides which brands even have a chance to show up. Change it later and you have thrown away your ability to compare over time.
A defensible set mixes three kinds of questions the way your buyers actually ask them: informational ("what is X"), comparison ("X vs Y"), and recommendation ("best X for Y"). Recommendation prompts pull the most brand names into an answer. Comparison prompts pull a medium amount. Informational prompts pull the fewest. So the mix you choose quietly shapes your result, which is exactly why you fix it and leave it alone.
How many prompts? There is no official standard, and anyone who tells you there is one is selling you their number. Some measurement vendors call 50 a practical floor and 100 to 200 a comfortable range; one large methodology uses more than 5,000. Pick a set big enough to represent how buyers really research your category, then stop fiddling with it.
You will know this step is done when your set is saved, labeled with a version like v1.0, and locked. Any change starts a fresh baseline.
Where people go wrong: they stack the set with prompts where they already win. That inflates the score and fools nobody for long. This guide assumes you already have a prompt set built; if you do not, build that first and come back.
Step 3: Declare your competitor set
Your competitor set is your denominator, and the denominator is a choice, not a fact. Choose it on purpose.
Write down every brand that counts, including yours, with all its aliases: product names, the parent company, domain variations, common misspellings. Then decide between a closed set and an open set. A closed set is a fixed list you declared in advance, and it ignores any brand the engine names that is not on the list. An open set includes every relevant brand the engine actually surfaces. Closed sets are reproducible and stable, which makes them easy to trust over time. Open sets catch brands you did not think of, but the denominator can drift as new names appear. Both are fine. Pick one, write it down, and keep it consistent.
You will know this step is done when a colleague could take your brand list and produce the exact same totals from the same set of answers.
Where people go wrong: they leave a strong emerging competitor off the list because including it drops their share. That is not measurement, that is wishful thinking. If the engine keeps naming a brand, it belongs in your set.
Choosing a fair competitor set is also where a lot of teams lean on a platform. DeepSmith lets you define exactly what counts as your brand and who your competitors are, then holds that list steady across every collection run, so your brand share of voice ChatGPT reports this quarter compares cleanly to last quarter.
Step 4: Collect answers the same way every time
Now you actually run the prompts. The rule here is boring on purpose: keep everything identical except the question.
Run every prompt through every engine under the same conditions. Same locale. Same language. Same logged-in or logged-out state. Same date window. Same wording. Then save the full response text and label each row with the engine it came from.
There is one more wrinkle, and it matters. AI answers are noisy. Ask the same engine the same question twice and you can get different brands in a different order. So do not trust a single run. Run each prompt several times and aggregate; a practical starting point some vendors use is 5 to 10 runs per prompt per engine. Your total sample is prompts times engines times runs.
You will know this step is done when every prompt, engine, and run sits as its own row in a dataset, with the full answer text intact and the date recorded.
Pro tip: this is the step that eats the most human hours if you do it by hand, and it is the step most worth automating. Running fifty prompts across three engines, five times each, on a schedule, is 750 collections you do not want to do in a browser tab. This is where a tool like DeepSmith earns its place: it runs your tracked questions on a schedule, keeps the full answer history, and stores it so this quarter compares to last quarter without you copying anything into a spreadsheet.
Step 5: Count events and normalize brand names
You have a pile of answers. Now you turn them into numbers.
Apply your Step 1 counting rule to every brand in your set, for every answer. Before you start, write a short brand-match list that resolves the fuzzy cases: does "Acme," "Acme Inc," "acme.com," and "the Acme platform" all count as one brand? Does a parent company's mention count for the subsidiary? Does a product line count for the brand? Decide these in writing, once, and apply them the same way everywhere.
You can note whether each mention was positive, neutral, or negative while you are in there. Just do not let sentiment secretly change the count, because that is a different metric for a different day.
You will know this step is done when you have exactly one total per brand per engine, and every total traces back to specific answer rows you could point to.
Where people go wrong: they let gut feel decide the edge cases in the moment. Two people counting the same answers then get two different totals. A written brand-match list is what makes your number reproducible instead of personal.
Step 6: Calculate per engine, then aggregate
Here is where the formula finally pays off. Do it per engine first, always.
For each engine, take your brand's events and divide by the total events for your whole comparison set, then multiply by 100. That gives you a clean per-engine share. Your brand share of voice ChatGPT returns and your share of voice in AI answers on Perplexity are separate numbers, and they will often disagree. A brand that ranks first on one engine can sit fifth on another for the very same prompts. That gap is a finding, not an error.
Only after you have the per-engine numbers do you roll them into one. Choose your aggregation rule and say it out loud:
- Pooled events: add your events across all engines, divide by the summed comparison-set events across all engines. This weights engines by how much data you collected.
- Equal-engine average: compute share per engine, then average the percentages. This treats every engine as equally important.
- Pre-declared weights: decide each engine's weight before you collect (say, based on where your buyers actually search), then take the weighted average.
You will know this step is done when you have one per-engine table, one aggregate number, and the aggregation rule named right next to it.
Where people go wrong: they show the aggregate alone and bury the per-engine breakdown. That hides exactly where you are winning and losing, which is the most useful thing the number can tell you. Always publish both.
Step 7: Report the number with its inputs
A share of voice number with no context is unverifiable, which means it is unusable the moment someone questions it. So report it with its receipts, every time.
Alongside the percentage, always include the variant (mention or citation), the counting rule, the full comparison set, the prompt set version and count, the engines and locales, the collection window, the number of runs per prompt, the aggregation rule, and the raw totals per brand. It sounds like a lot. Written once as a template, it takes a sentence or two.
Here is a template you can adapt: "Our AI share of voice across ChatGPT and Perplexity for Q3 was X percent, measured as answer-level mention events against a closed set of four brands. The prompt set was v1.0 with 100 prompts across informational, comparison, and recommendation types. We collected 5 runs per prompt. Per-engine shares were listed below; the aggregate used pooled events."
You will know this step is done when someone who was not in the room could reproduce your number from the report alone.
Where people go wrong: they publish one confident percentage with nothing behind it. The first smart question from leadership then sinks it. Show your inputs and the number becomes something you can stand behind.
Walk through a full worked example
Numbers make this concrete, so let's run one all the way through. This is exactly how to calculate share of voice AI answers give your brand.
Say you fix a 100-prompt set, a closed set of four brands, and 5 runs per prompt, all on a single engine. That is 500 answer runs collected. Using answer-level counting, your mention events land like this:
- Brand A (you): 142
- Brand B: 198
- Brand C: 110
- Brand D: 50
- Total comparison-set events: 500
Your mention share of voice is (142 divided by 500) times 100, which is 28.4 percent. Brand B lands at 39.6 percent, Brand C at 22.0 percent, and Brand D at 10.0 percent.
Now the sanity check: 28.4 plus 39.6 plus 22.0 plus 10.0 equals exactly 100. Because all four brands share one denominator, a valid closed-set share always adds to 100. If yours does not, go find the counting error before you report anything.
Want to see how mention and citation differ? Take the same setup but count citation events (links to a brand-owned page) instead:
- Brand A (you): 38
- Brand B: 70
- Brand C: 45
- Brand D: 27
- Total: 180
Your citation share of voice is (38 divided by 180) times 100, roughly 21.1 percent. Notice it sits below your 28.4 percent mention share. That gap is a real signal: engines name you fairly often but cite your pages less. Both numbers live side by side. Neither replaces the other.
The mistakes that quietly break your number
Most bad share of voice numbers come from a short list of repeatable errors. Scan these before you publish anything.
- Dividing by prompts instead of by comparison-set events. That gives you a mention rate, not a share. Use the comparison set's total as your denominator.
- Switching counting rules between brands. One rule for everyone, always.
- Pooling engines without labeling. Stacking ChatGPT and Perplexity into one bucket with no rule hides the differences that matter most.
- Reporting a bare percentage. No inputs means no trust. Show the receipts.
- Treating vendor benchmark ranges as law. You will see ranges like "under 15 percent is a gap, 25 to 40 percent is competitive, above 40 percent is strong." Those come from individual vendor methodologies. Use them as loose orientation, not as a pass-fail grade.
- Calling share of voice a business guarantee. It measures relative visibility in a sample of answers. It does not measure clicks, conversions, or revenue, so do not let anyone frame it that way.
- Ignoring no-answer states. Some engines do not show an AI answer for every query. Record those separately. Do not silently treat them as a loss.
One more honest note: AI answers move. In one published multi-engine experiment, the brands listed, their order, and even how many were recommended shifted noticeably from run to run. So treat any single share of voice as an estimate from a sample, not a permanent ranking. Track the trend across quarters, not the wobble between two snapshots.
Metrics that sit next to share of voice
Share of voice is one number, not the whole picture. A few neighbors are worth a glance once your share calculation is running: citation rate (how often your pages get cited as sources), prominence or position scoring (weighting mentions by where they appear, which is contested because it amplifies the noise), sentiment (whether a mention was positive or negative), and full competitor benchmarking (share plus citation rate plus sentiment plus trend together). You do not need to build these today. Just know they exist, so your share of voice number stays honest about what it does and does not cover.
What to do next
Take a breath. You now have everything you need: one formula, seven steps, and a worked example you can copy. If you only do one thing this week, run the calculation once, by hand, for a single engine. Rough is fine. A real 28.4 percent you can defend beats a perfect number you never start.
Then decide what to keep doing by hand and what to hand off. The formula stays simple, but collecting hundreds of answers on a schedule, holding your competitor set steady, and building the per-engine table every quarter is the part that quietly eats your week. That is the part DeepSmith was built to run for you: it tracks your questions across engines, computes mention rate, citation rate, and share of voice with per-platform breakdowns and a competitor leaderboard, and keeps the answer history so each period compares cleanly to the last. It tracks and reports what the engines show; it will not promise you a ranking, and you should be wary of anything that does.
You can start a free DeepSmith trial and see your own real numbers before you pay. Either way, you have the method now. Go run it.



