You searched your category in ChatGPT, and there you were. Cited. You exhaled.
But cited where? At the top, where the answer says "start here," or in a collapsed source list seven names down, where almost no one looks? The real question is not whether you appear, it is how prominent is my citation when I do, and that gap is what ai citation prominence measures. Most teams never check it. This guide shows you how to score whether you lead the answer or sit buried at the bottom, and how to track that as its own number over time.
Here's the good news: you already have the hardest part done. You know you show up. Now let's find out how much that appearance is actually worth.
What "cited or buried" really means
Being present and being prominent are two different things. A yes or no citation count tells you that you exist somewhere in the answer. It says nothing about whether a buyer will ever see you.
Think of three questions, not one.
- Citation rate asks: are we present at all?
- Share of voice asks: how much of the conversation do we own versus competitors?
- Prominence asks: when we are cited, are we leading the answer or lost in the footer?
Prominence is the third question, and it is the one most dashboards skip. A brand can be cited on 90% of tracked prompts and still lose every buyer, because every appearance sits at position seven of twelve in a source list nobody expands. Another brand cited only 20% of the time, but leading four out of five of those answers, will win far more attention. Frequency without position is a vanity number.
So prominence does not replace citation rate or share of voice. It sits beside them. All three belong on the same dashboard, and prominence is the one that explains why your share of voice is moving the way it is.
Why buried costs more than it used to
Before you spend an hour measuring anything, it helps to know why this one dimension is worth the effort. The short version: the answer is now the destination, and only the top of it gets seen.
Look at what happens to clicks. Only around 1% of AI Overview searches end in a click through to a cited source. The answer itself is where the reader stops. When almost no one clicks past the answer, your only real estate is the answer, and your only leverage is how high inside it you sit.
The traffic math backs that up. For queries where an AI Overview shows up, organic click-through has dropped sharply, from roughly 1.76% to about 0.61%, a fall near 61% in one widely cited 2025 analysis. The clicks did not spread evenly, either. A lead position inside an AI Overview is worth something like a traditional organic position six in traffic terms, and it carries a stronger trust signal because the AI is vouching for you.
Here is the part that frees you: being cited or buried in ai answers is not locked to your search ranking. Close to 60% of AI Overview citations come from URLs that never appear in the top 20 organic results. Your placement inside the answer is a separate game from your placement in the blue links, which means you can move your prominence without first winning the whole SERP. That is exactly why it deserves its own metric, and its own measurement.
Score every appearance on a simple three-tier scale
You cannot measure what you cannot name. So before you count anything, give position a scale. Three tiers are enough, and they are simple enough to apply by eye.
Tier 1, the lead mention. You are named first, named in the opening sentence, or offered as the primary recommendation. In an AI Overview, you are the single source at the top of the panel. In a numbered list, you are number one. To the reader, this feels like the AI pointing straight at you.
Tier 2, the in-body citation. You are woven into the narrative of the answer, something like "according to your brand," or you sit second or third in an inline shortlist. You are shaping what the reader actually reads, just not leading it.
Tier 3, the source-list appearance, the buried case. You show up only in the reference cluster, sidebar, or footnote list at the bottom. In most chat interfaces that list is collapsed or cut off. The reader has to hunt to find you. This is what "buried" looks like.
Here is the scale on one worked answer. Say the prompt is "what are the best AI search analytics platforms," and the answer body names one platform as the leader, then two more in a sentence, then drops seven names in a source list at the bottom. The leader scores Tier 1. The two named in the body score Tier 2. Everyone in the source list, no matter how many, scores Tier 3.
That is the whole language you need. Lead, in-body, or buried. Once every appearance has a tier, you can finally measure ai answer prominence instead of guessing at it.
Build a fixed prompt test set
Now let's gather the raw material. Prominence is scored one answer at a time, so you need a stable set of answers to score.
Write a fixed list of 30 to 100 buyer questions you want to be present for. Keep it fixed so you are comparing the same prompts week over week, not chasing a moving target. Cover three buckets:
- Informational: "what is X," "how does Y work."
- Comparative: "X versus Y," "best tools for Z."
- Decision: "which X should I buy," "recommend a Y for this use case."
You will know this step is done when your list reflects how real buyers actually ask, across all three stages, not just your favorite branded keyword.
Where people go wrong: they run one flat, generic prompt and call it coverage. AI engines shift who they put forward based on the intent they sense. A persona-blind keyword hides that completely. Write the questions a buyer would type, in their words, at each stage.
Run each prompt across the engines, on repeat
One screenshot is not a measurement. It is a single frame of a moving picture.
Send every prompt to every engine you care about. At a minimum, cover ChatGPT, Perplexity, and Google AI Overviews. Add Gemini and Claude if your buyers use them. For each run, capture the full answer text and any citation data the engine returns, and stamp it with a timestamp so you can compare over time.
Then do it again on a schedule. Position in ai answers behaves more like a probability than a fixed rank. A prompt that returns you as the lead six times out of ten is a very different reality from one that leads nine times out of ten, and you only see that by running the same prompt repeatedly. Weekly is the floor for an active program. Monthly is the low end. Daily is overkill for almost everyone.
This is the step that quietly eats your week if you do it by hand, and it is where a platform earns its place. DeepSmith runs your tracked prompts across ChatGPT, Perplexity, Gemini, Claude, and Google AI Mode on a set schedule and keeps the full answer history, so you are scoring stored results instead of re-running searches every Monday. You define the questions once; it checks them and reports back.
You will know this step is done when you have a timestamped record for each prompt and engine, not a folder of one-off screenshots.
Turn the tiers into one prominence score
Now the fun part, where a pile of answers becomes a single number you can watch.
For every prompt, on every engine, mark your appearance as Tier 1, Tier 2, Tier 3, or absent. Then weight the tiers. A clean, reasonable weighting looks like this:
- Tier 1: 1.0
- Tier 2: 0.7
- Tier 3: 0.3
- Absent: 0
Average that weighted score across your whole prompt set, and you have a brand-level prominence score between 0 and 1. Multiply by 100 if you prefer a 0 to 100 scale. That single number is your answer to the question buyers really care about: how prominent is my citation, on average, when the AI talks about my category?
Small sets you can score by hand. Large sets you can hand to an LLM to classify, using the same three-tier rule you just defined. Either way, the scale is what makes it possible.
A common mistake to avoid here: counting a Tier 3 source-list appearance as a win. It inflates your citation count without adding any real influence, and it hides the fact that you are being buried. The point of a prominence score is to filter those low-value appearances out, not to celebrate them.
Refine the score with the signals position alone misses
Your weighted score is already useful. If you want it sharper, four secondary signals turn a rough tier into a real ranking.
- List rank within a tier. Second in an inline shortlist beats fifth. Weight by where you land, not just which tier you land in.
- Recommendation strength. "Best overall" carries more weight than "an option worth considering," which carries more than "may suit some teams."
- Citation proximity. Is your link sitting right next to your name, or detached in a general pile at the bottom? Adjacent counts for more.
- Consistency. Does your position hold across repeated runs, or swing between lead and buried? A stable Tier 1 is worth more than a flickering one.
You do not need all four on day one. Start with the three tiers, get a baseline, then layer these in as your program matures. Momentum matters more than a perfect model.
Where people go wrong: they treat position as static, as if one good run locks it in. It does not. Score the pattern across runs, not the single best screenshot you happened to catch.
Know what your tools can and cannot show you
One honest caveat before you buy anything. Most AI visibility tools on the market today report citation rate and share of voice cleanly, and treat position in ai answers as a secondary signal at best. Very few expose the lead, in-body, or source-list tier as a named, first-class number you can trend.
That is not a reason to skip tracking. It is a reason to know what you are getting. A tool that surfaces per-prompt position data and per-page citation attribution gives you most of what you need; you supply the tier judgment on top. When you compare options, ask two plain questions: does it show where you land inside each answer, not just whether you land, and does it hold the full answer history so you can trend ai citation prominence rather than react to one run? If the answer to both is yes, you can build a reliable prominence view around it.
Roll it up, trend it, and act on what it tells you
A score you take once is a snapshot. A score you take every week is a story. This is where prominence starts paying you back.
Watch the two lines together, citation rate and prominence, because their relationship is the real signal.
- Citation rate flat, prominence rising: your content is earning better placement. Keep doing what you changed.
- Citation rate flat, prominence falling: you are being pushed down the answer. Something is climbing over you.
- Prominence up on one engine, flat on another: the engines are reading your content differently, and one needs attention.
That last point matters more than it looks. The engines genuinely differ. In one large study of citation patterns, ChatGPT leaned heavily on Wikipedia, Perplexity leaned on Reddit, and Claude rewarded technical precision. One engine's lead mention is another engine's buried footnote, so a single-engine view will lie to you.
Reading the trend is the measurement. Acting on it is the payoff, and it needs two things your score alone will not give you: which of your pages is winning the citations, and which competitor page is beating you when you lose. DeepSmith's Pages view shows exactly which URLs on your site are earning citations and for which prompts, and the Competitor Citations leaderboard shows who is taking the top positions you want, on which exact pages. That is the difference between knowing your prominence slipped and knowing which page to fix and which rival to displace.
You will know this step is working when a falling score sends you to a specific page, not to a vague worry.
The mistakes that quietly ruin a prominence score
Even a good method fails if you fall into these. Watch for them.
- Confusing mentions with citations. Being named in prose is a mention. A linked attribution is a citation. They behave differently and need separate counters.
- Trusting a single screenshot. Volatility is the rule, not the exception. One Tuesday capture is not a measurement.
- Celebrating the footer. A Tier 3 appearance is presence, not prominence. Do not let it inflate your numbers.
- Tunnel vision on one engine. ChatGPT, Perplexity, Gemini, and Google AI Overviews each produce different patterns. Measure across all the ones your buyers use.
- Persona-blind prompts. Flat keywords hide how engines shift by intent. Test with real buyer-stage questions.
None of these are hard to avoid once you can see them. If you only fix one this month, fix the single-screenshot habit. Track over time, and the rest gets easier.
What to do next
You do not need a bigger team to start this. You need a fixed prompt set, a three-tier scale, and a standing weekly slot to run and score it. That is the whole engine. Everything after is refinement.
Start small. Pick ten prompts, run them across two engines, score whether you are cited or buried in ai answers, lead, in-body, or buried, and write down the number. Next week, run the same ten and watch the line move. That first trend is the moment prominence stops being a feeling and becomes a metric you manage.
If the weekly running and scoring is the part that keeps slipping, that is exactly the manual work worth handing off. DeepSmith tracks citation rate, share of voice, and visibility trend across five engines, and shows you the pages and competitors behind the movement, so your time goes to acting on the score instead of assembling it. You can see real data on your own prompts during a free trial before you pay for anything.
You are closer than you think. You already know you show up. Now go find out whether you lead or hide.



