You already know AI engines mention your brand sometimes. What you don't know is whether they're saying anything good about you when they do. This guide walks marketing leads through how to track sentiment AI mentions carry: a rubric for labeling each answer, a baseline to compare against, and a way to watch the trend across engines over time. By the end you'll have a repeatable scorecard, not a gut feeling from reading a handful of ChatGPT answers.
Know the difference between a mention, a citation, and sentiment
Before you build anything, get three terms straight, because mixing them up is the most common way teams end up with a scorecard that doesn't mean anything. If you haven't looked at your AI presence at all yet, run a one-time audit first and come back to this once you have a starting picture.
A mention is just the AI naming your brand somewhere in its answer. A citation is the AI linking to one of your pages as a source. Sentiment is a separate question from both of those, about whether the AI frames you well, badly, or somewhere in between when it does mention you. An answer can mention your brand favorably without citing your site at all, and it can cite your page while still describing your brand in lukewarm terms. Treating these three as interchangeable is how a rising mention count gets mistaken for improving brand sentiment when the tone hasn't actually moved.
This guide covers brand sentiment AI search metric work specifically: the classification and trend tracking. It doesn't cover setting up the full monitoring program that generates those answers in the first place, and it doesn't cover what to do once you find a negative pattern. Those are separate problems worth their own process.
Step 1: Write your sentiment rubric
What to do: Before you read a single answer, write down what positive, neutral, and negative mean for your brand, in one sentence each. Use a three-class scheme:
- Positive: the answer presents your brand favorably or recommends it for a relevant use case. Praise, a stated preference, or inclusion among the best options all count.
- Neutral: the answer mentions or describes your brand without a clear favorable or unfavorable judgment. A factual description or a feature listing with no evaluative language lands here.
- Negative: the answer presents your brand unfavorably, warns against it, or points to a real shortcoming. Criticism, a warning, or an unfavorable comparison all count.
Add rules for the cases that will actually come up. If an answer's overall recommendation is favorable, it's positive, even if it mentions a caveat. If the overall conclusion is unfavorable, it's negative, even if one line praises a feature. If the answer is balanced or mostly factual with no clear direction, it's neutral. Label the answer's overall framing, not one isolated adjective buried in the middle of it.
How to tell it's done: Hand the rubric to a second person and have them label a sample answer. If they can do it without asking you what the categories mean, the rubric is specific enough to use.
Common mistake: Treating every mention as positive just because your brand showed up. A factual, no-opinion mention of your brand is neutral, not positive, even if the surrounding paragraph criticizes a competitor.
Pro tip: Require a short written reason next to every label, even a single clause. It's the difference between a rubric someone can check later and a rubric that only makes sense in the moment you applied it. Borderline calls get reviewable, and inconsistent interpretation gets exposed fast.
Step 2: Fix your prompt set and engine list
What to do: Pick the buyer-relevant prompts you'll track and the AI engines you'll track them on, then use that same set every single period. Keep the wording, the prompt order where it's practical, and the language consistent. Write down which version of the prompt set you're using so you can point back to it later.
This step isn't about building a complete AI visibility monitoring program from scratch, that's a bigger project with its own process. What matters here is comparability: if you change the questions you're asking, you've changed the population you're measuring, and any sentiment shift you see afterward might just be an artifact of a different prompt mix.
DeepSmith runs a set of tracked prompts across AI engines on a recurring schedule as part of its AI Visibility metrics, which gives you a consistent measurement set without having to rebuild your prompt list by hand every period. That keeps the comparison frame stable, though it's not a promise that any individual answer will stay the same from one collection to the next.
How to tell it's done: Someone on your team can reproduce last period's exact prompt set and name which engines were included, without digging through old notes.
Common mistake: Building your baseline from broad category questions, then later comparing it against a period built mostly from branded questions. That mismatch alone can look like a sentiment swing when nothing about how AI talks about you actually changed.
Step 3: Keep the answer-level record, not just the final number
What to do: For every response that mentions your brand, save the date, the engine, the prompt, the full answer or a reviewable version of it, the sentiment label, and the short reason for that label. Keep the raw count alongside whatever percentage you end up reporting.
How to tell it's done: You can trace any positive, neutral, or negative number in your scorecard back to the exact prompt and answer that produced it.
Common mistake: Keeping only the final percentage and throwing away the underlying answers. Without them, you can't explain a sudden shift or tell a real pattern apart from a fluke caused by a small sample.
DeepSmith's AI Visibility module keeps recurring tracked-prompt data and answer-level visibility metrics in one place, so the underlying answers stay available for review instead of living in a spreadsheet someone forgot to update. The tool helps you monitor that data; the classification judgment in the next step is still yours to make.

Step 4: Label the AI framing in each answer
What to do: Read the language around your brand in each saved answer and apply the rubric from step 1. You're labeling the overall framing, not just whether the brand appeared. Use neutral for anything factual, or anything genuinely balanced with no clear conclusion either way.
How to tell it's done: Every brand-mentioned answer has exactly one headline label, and every borderline call has a short explanation attached to it.
Common mistake: Confusing rank with sentiment. A brand listed second in an answer can still be described warmly, and a brand listed first can still come with real warnings attached. Judge the actual wording and conclusion, not where your brand landed in the list.
Common mistake: Confusing the source with the sentiment. An AI answer citing your own website as a source doesn't automatically mean the answer is positive about you.
Step 5: Calculate your baseline
What to do: Once you have a full, consistently collected period behind you, use it as your baseline. Calculate the positive, neutral, and negative rates for that period, save the counts behind each rate, and do this separately for every engine you track before you calculate anything blended across engines.
A useful primary number is the positive sentiment rate:
positive sentiment rate = positive brand-mentioned answers divided by total brand-mentioned answers, times 100
The same formula, swapping the numerator, gives you a neutral rate and a negative rate. If every brand-mentioned answer gets exactly one label, the three rates should add up to 100 percent.
How to tell it's done: Your baseline table has the period, the prompt-set version, the engines covered, the raw counts, the rates, and the no-mention volume, all in one place.
Common mistake: Setting a target of 100 percent positive sentiment. Most brand sentiment in the real world lands somewhere in the neutral middle. A perfect positive rate isn't a realistic default or a fair goal to hold your team to.
Pro tip: Treat the baseline as a starting reference, not a fixed target. Once you understand the normal spread and the volume of answers behind it, you can set a more sensible goal than "everything positive."
Step 6: Track the trend by period and by engine
What to do: Repeat the exact same calculation on a schedule you can stick to. Weekly, monthly, and year-over-year comparisons all work, as long as the cadence stays consistent inside one trend report. For each new period, calculate the change in positive, neutral, and negative rate from your baseline, the change from the immediately preceding period, the change in no-mention rate, the change by engine, and the change in your underlying answer count.
Use percentage points when you talk about these changes. If your positive rate moves from 24 percent to 31 percent, that's a 7 percentage-point increase, not a vague "sentiment got better." A simple table or line chart works well here, with separate lines for positive, neutral, and negative so a reader isn't stuck inferring the trend from one blended number.
DeepSmith treats sentiment and visibility trend as part of the same AI search analytics it already tracks, so the scorecard sits alongside your other visibility metrics instead of living in a separate tool. It's a place to watch the numbers move period over period, not a guarantee that the numbers will move in your favor.
How to tell it's done: Your report shows a dated sequence of comparable numbers, and you can tell whether a change happened everywhere or on just one engine.
Common mistake: Calling one unusual answer a trend. A single snapshot tells you where you stand right now. You need repeated collection under a comparable setup before you can call anything a trend.
Step 7: Decide how you'll combine sentiment across engines
What to do: Once you have rates for each engine, you'll want one overall number for the top of your report. There are two defensible ways to get there, and they don't produce the same answer.
A pooled rate combines every brand-mentioned answer across all engines into one bucket and calculates a single rate from that. Engines with more observations naturally carry more weight in the result. A macro-average instead calculates a rate for each engine separately, then averages those engine rates together, so every engine counts equally regardless of how many answers it produced.
Pick one method for your headline number and stick with it period over period. Don't switch between the two depending on which one looks better that month. The clearest way to present this to a marketing team is usually a per-engine table alongside the single overall number, with the aggregation method named right next to it, so nobody assumes a blended figure means something it doesn't.
You can also compress the three-class distribution into one directional number if you want a compact line to track: subtract your negative rate from your positive rate to get a net sentiment score, which runs from negative 100 to positive 100. Treat this as an optional secondary view. It should never replace the full positive, neutral, and negative breakdown, because a net score of zero can mean "everything's neutral" or "half positive, half negative," and those are very different situations to be in.
How to tell it's done: Anyone reading your report can see which aggregation method produced the headline number, and the per-engine breakdown is sitting right there if a number looks off.
Common mistake: Hiding a big gap between engines inside one blended figure. If ChatGPT sentiment is strongly positive and Perplexity is trending negative, a single combined number buries exactly the thing you need to see.
Step 8: Read the change without overclaiming what it means
What to do: Look at the trend alongside the actual answer text, the prompt mix, the engine, how many observations you're working with, and any known model or product changes. Figure out whether a shift is showing up broadly or on just one platform.
Use these patterns as a starting point for how to read a move:
- Positive rate up and negative rate down, with a comparable prompt and engine mix, is a real directional improvement.
- Positive up and neutral down while negative holds steady means more answers are expressing a preference, but whatever's driving the negative answers hasn't been resolved.
- Neutral rate rising while positive and negative stay flat usually means the AI is describing you more often without taking much of a position either way.
- A negative rate that rises on one engine only is worth investigating on that engine specifically before you call it a company-wide problem.
- If every rate moves at once because you changed the prompt mix, don't attribute it to sentiment. Re-establish your baseline against the new prompt set instead.
How to tell it's done: Your report can answer three plain questions: what changed, where it changed, and how much evidence actually backs the change.
Common mistake: Treating a positive sentiment bump as proof a specific campaign caused it. A sentiment trend is something to monitor, not evidence of cause on its own. It's also a mistake to treat one engine's result as your brand's whole AI reputation, since different engines can frame the same facts differently depending on what they retrieve.
If a negative label turns out to be based on something the AI got factually wrong about your brand, fixing that is a separate response process, not part of the measurement work covered here.
A few things worth knowing about AI sentiment specifically
AI-answer sentiment isn't the same thing as traditional brand sentiment pulled from reviews, social posts, or press coverage. The two measure different conversations and can move in different directions, so don't present them as one number.
Small changes in how you word a prompt can shift how an AI classifies the answer, and the same prompt can produce a slightly different answer on a different run. Model versions change over time too, which can affect how comparable your numbers stay across a long trend. The practical response is the one running through every step above: keep your prompts and your labeling rules stable, keep the raw answers, note any known model changes, and don't read a single result as a definitive verdict on how AI perceives your brand.
A percentage on its own can also mislead you. A positive rate built from a dozen answers looks identical on a chart to one built from a few hundred, but the second one carries a lot less uncertainty. Always show the volume next to the rate.
What to do next
Run through steps 1 and 2 once to set your rubric and lock your prompt set, then let steps 3 through 6 run on whatever cadence you picked, weekly or monthly usually works well for most teams. After two or three periods you'll have a real trend line instead of a single snapshot, the kind of scorecard worth having ready the next time stakeholders ask how AI search is going for your brand. If you want the tracked prompts, the answer-level data, and the sentiment trend built for you instead of a spreadsheet you're maintaining by hand, start a free trial and set it up against your own brand.



