DeepSmith

Sep 26 · AEO & AI Visibility

16 min read

How to Monitor Your Brand's Reputation Across ChatGPT, Gemini, and Perplexity

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
Abstract illustration of three dark rounded nodes representing separate AI engines, each connected by thin lines to a central stack of log pages and small chart marks, with the white text Track Your Brand Across AI Engines centered on a charcoal background.

Checking your brand's name in ChatGPT once and feeling good about the answer does not tell you much. Ask the same question tomorrow, or ask it in Gemini instead, and you can get a completely different result. If you want to monitor brand reputation AI assistants are building for you, you need something steadier than an occasional search: a fixed set of questions, a consistent way of asking them, and a place to keep what came back. This guide walks through building that system, engine by engine, so you can track brand mentions ChatGPT and the others generate without guessing whether a change is real or just noise.

By the end you will have a prompt library, a run routine for each engine, a cadence to run it on, and a simple way to log and compare what you find over time. This is continuous AI brand tracking, not a one-time check. If you just need a single snapshot of where you stand today, that is a different job, this is about keeping watch, and it is how you monitor brand across Perplexity Gemini and ChatGPT without treating three different products as one.

Define what monitoring means for your team

Before you collect a single answer, write down what you actually mean by monitoring. It sounds obvious until two people on your team classify the same ChatGPT answer two different ways.

Start with a plain definition: you are watching whether AI assistants mention, recommend, cite, or leave out your brand when they answer the questions your buyers actually ask, then comparing that across engines and over time. That is what it means to monitor brand reputation AI assistants are shaping for you, and it only works if the rules are settled before you start. From there, settle a short list of decisions:

  • Which three engines you are watching (ChatGPT, Gemini, and Perplexity, for this guide) and which mode of each one counts as your main run.
  • Your primary market and language, since answers shift with both.
  • The buyer stage each prompt represents, so you are not lumping awareness questions in with decision-stage ones.
  • Your competitor set, named and fixed, not added on the fly when someone new comes up.
  • How you treat a failed or missing run. It should never be silently counted as a non-mention.

Write this down as a short measurement policy: what counts as your brand, what counts as an owned citation, which brand aliases you accept, and which competitor names are in scope.

How to tell it is done: hand the policy to a second person and have them classify one answer using only what you wrote. If they land somewhere different than you would, the policy still has gaps.

Where people go wrong: they change the rules after seeing a bad answer. One week a mention of the product name counts, the next week someone decides it does not because the context felt off. Lock the definitions before you start collecting anything you plan to compare over time.

Build a fixed buyer-prompt library

A prompt library is the backbone of AI brand monitoring tools and manual tracking alike. Without one, you are just poking around, and poking around does not produce a trend. This is also the step where you decide how you monitor brand across Perplexity Gemini and ChatGPT with the same set of questions, instead of three unrelated lists that never let you compare an engine to another.

Build prompts the way a real buyer would ask them, not the way your team would. A useful starting point is 24 prompts split into six groups of four:

  • Brand discovery: "What is [Brand]?", "Who is [Brand] best for?"
  • Category discovery: "What are the best [category] tools for [audience]?"
  • Comparison and shortlist: "[Brand] versus [Competitor]: which is better for [use case]?"
  • Trust and evaluation: "What are the strengths and weaknesses of [Brand]?"
  • Use-case and buyer-stage: "What is the best [category] platform for [specific workflow]?"
  • Source and evidence: "Which sources explain [Brand]'s product or category?"

Swap in your own category, use cases, audience, and named competitors, and keep the wording exact once you settle on it. Do not rewrite a prompt just because one answer disappointed you.

DeepSmith's Prompts area is built around this exact idea: a tracked set of buyer questions with per-prompt mention and citation rates and a full answer history. Its Discover Prompts feature can generate a starter set from your product, persona, and buyer-stage context, which saves you the blank-page problem. Still, review and edit what it suggests rather than accepting every line, since the point is real customer language, not a generic list.

How to tell it is done: every prompt has a stable ID, exact wording, a group, a buyer stage, and the engines it runs against, and you could paste any of them unchanged into a new conversation right now.

Where people go wrong: they only write branded prompts like "What is [Brand]?" That tells you whether the engine recognizes your name, not whether it recommends you over a competitor in a category or comparison question, which is usually where the real gap shows up. A related mistake is writing dozens of near-duplicate prompts that pad the count without adding a new question a buyer would actually ask.

Pro tip: split your library into a core set that never changes and an experimental set for testing new markets or emerging buyer language. The core set is what gives you a real trend line. The experimental set lets you explore without breaking it.

Standardize how you run each engine

This is the step most manual monitoring skips, and it is why so many teams end up comparing answers that were never comparable in the first place. Each engine needs its own run card: a documented routine that anyone on your team can follow and get the same kind of answer.

ChatGPT. Search can trigger automatically or be turned on manually, so decide which one you mean to test and do it the same way every time. Start each prompt in a fresh conversation rather than a running thread, since ChatGPT can carry earlier context into follow-up answers and that stops being an independent measurement. Note whether memory is on, since it can influence how ChatGPT rewrites your query, and note your location or VPN setting for anything location-sensitive. Save the full Sources panel and every cited link, not just the answer text.

Gemini. Keep the surface consistent: a consumer Gemini Apps answer and an API-grounded response are not the same product and should never sit in the same comparison. Start a new conversation for each prompt, record whether the answer showed sources, and check whether any linked source came from a connected file rather than the open web. A missing source is a separate observation from a missing brand mention, so track them separately.

Perplexity. Perplexity offers standard Search, Pro Search, and Deep Research, and these are not interchangeable. Choose one mode for your main tracking series and stay with it; use the others only as a clearly labeled side series if you want them. Record the model when the interface lets you pick one, start a new thread per prompt, and save the full numbered citation list mapped back to its sources.

How to tell it is done: a teammate who was not in the room can rerun any prompt and land on the same button clicks, the same account type, and the same starting conditions you used.

Where people go wrong: comparing a ChatGPT answer with Search on to a Gemini answer with no visible sources, or setting a Perplexity Deep Research report next to a plain Search answer, then reading the difference as a brand signal when it is really just a mode difference.

Pick a cadence and stick to it

A prompt library only earns its keep if you run it on a schedule you actually keep. Weekly is the right default for most teams: run the full core set across all three engines on the same day, in the same order, and note the timestamp. If your category moves fast, or you are in a launch window, add a smaller daily watchlist of your highest-value prompts. Once a month, step back and look at the whole trend: prompt coverage, competitor changes, and whether your library still needs upkeep.

Treat a rebrand, a pricing change, or a major site update as its own labeled diagnostic run, separate from the normal weekly series, so you are not merging a one-off spike into your regular trend.

Make failures visible instead of invisible. A prompt can come back as a successful mention, a successful non-mention, a citation, a non-citation, a platform error, a login problem, a missing source panel, or simply not run. Each of these is a different thing, and folding them all into a silent zero will quietly corrupt your numbers.

DeepSmith's AEO settings let you define your brand, your competitors, and how often collection runs, then execute the prompt set against live engines on that schedule and record whether the brand was named, linked, or missing. That is the piece to reach for once weekly manual runs across three engines start slipping, which for most teams happens faster than they expect.

How to tell it is done: you have several comparable collection periods behind you, and every gap in the data has a documented reason instead of a silent zero standing in for it.

Where people go wrong: running the full set only when someone remembers, then calling the result a trend. A second common mistake is changing the cadence during a launch and comparing those numbers to an ordinary week without flagging the difference.

Log the full answer, not just yes or no

If you want to track brand mentions ChatGPT and the other two engines generate, a spreadsheet cell that says "mentioned: yes" tells you almost nothing six weeks from now. Keep one row per prompt, per engine, per run, and record both the structured fields and the raw evidence behind them.

At minimum, capture:

  • Run identity: run ID, collection date and time, engine, interface, search mode, model if visible, account or workspace, market and language, prompt ID and version, and run status.
  • Answer evidence: the full answer text, whether the brand was mentioned, the exact wording used, any alias detected, the brand's position in a list if there is one, competitors named, and whether the answer recommended, compared, or simply mentioned your brand.
  • Citation evidence: whether an owned page was cited, which page, any third-party source cited, its title and URL, and its citation number where the engine provides one.

A screenshot alone will not cut it. Without the prompt version, the run conditions, and the full text, you cannot tell later whether a change came from the engine, the prompt, your location, the source that got cited, or a simple classification error on your end.

Note whether the language the assistant used was positive, neutral, or negative as raw evidence in that same row; scoring sentiment systematically is a separate discipline worth its own process once mentions and citations are solid.

How to tell it is done: a reviewer can open one row and see the prompt, the conditions, the answer, and the classification, without rerunning anything to make sense of it.

Where people go wrong: saving only a screenshot or a one-word verdict. When something shifts later, there is nothing to investigate, and you are left guessing.

Compare mention, citation, and share of voice by engine

Build your comparison in layers, and always look at the per-engine numbers before you touch an aggregate.

Coverage first. How many planned runs actually completed for each engine and period. A low mention rate that is really a low completion rate is an operations problem, not a reputation one.

Brand visibility. For each engine, your mention rate (mentions divided by eligible runs), your owned citation rate (runs with at least one owned citation divided by eligible runs), and how often a competitor shows up where you do not.

Competitive visibility. Which competitors appear most often on the same prompts, which one gets recommended first, and which competitor pages keep getting cited in place of yours.

Source and answer change. New or disappearing owned citations, sources that show up across more than one engine, and prompts where the wording used to describe you has shifted even though the prompt itself has not changed.

Resist the pull to average everything into one number right away. A brand can be strong in Perplexity and nearly invisible in Gemini, and an aggregate score will smooth that gap out of view exactly when you need to see it.

This is the layer where DeepSmith's AI Visibility overview earns its place: mention rate, citation rate, and share of voice with per-platform trends, a competitor leaderboard, and the sources cited most often. Its Pages view shows which of your own URLs are actually getting cited, and Competitor Citations shows which competitor pages are winning the same prompts you are tracking. It saves you from stitching three engines' worth of spreadsheet rows together by hand, but it will not replace understanding what the underlying numbers mean, so keep your measurement policy even after you automate the collection.

DeepSmith's AI Visibility overview dashboard showing mention rate, citation rate, and share of voice as separate top-line metrics, a bar chart breaking mention and citation rate out by ChatGPT, Perplexity, and Gemini, and a competitor leaderboard ranking tracked rivals against your own position.

How to tell it is done: for ChatGPT, Gemini, and Perplexity separately, you can say how often your brand appeared, how often it got an owned citation, which competitors showed up instead, and whether that changed from the last comparable period.

Where people go wrong: blending the three engines into one score from the start. That hides a strength in one engine and a total absence in another, which is exactly the kind of thing you need to catch.

Review changes and keep the system honest

After every collection period, spend a little time on the numbers before you trust them. Check which runs completed, look for any big swings in mention or citation rate, and open the actual answers behind those swings rather than reading the dashboard number on its own. Ask whether the prompt, the mode, the model, the account, the location, or the cited sources changed, since any one of those can move a number without your brand's actual standing changing at all.

Record confirmed changes in a short log so the team has a memory longer than one person's. Review your experimental prompts monthly, and touch the core set less often, only when a prompt stops representing a real buyer question or a genuine market shift shows up.

If one answer looks wrong, do not rewrite the prompt on the spot. Rerun it exactly as written under the same conditions first, and compare the full evidence before you decide anything changed.

How to tell it is done: there is a named owner for this review, a regular deadline, a change log with entries in it, and a clear rule for when a prompt gets added, edited, or retired.

Where people go wrong: treating the dashboard as self-explanatory and skipping the raw-answer review. That produces a confident-sounding explanation for a number that nobody actually checked.

A flow diagram showing a single prompt library branching out to ChatGPT, Gemini, and Perplexity, all three converging into one shared evidence log, which feeds a trend comparison, with a labeled return line carrying the process back to review and refine the prompt library.

What to do next

Once the manual framework above makes sense to you and your team, the next decision is whether to keep running it by hand or hand the recurring collection to a tool. A spreadsheet can carry a small pilot just fine, as long as you keep the exact prompts, the full answers, the citations, the timestamps, and the run conditions in it. It gets harder to sustain once you are tracking three engines, a real competitor set, and enough prompts to see a pattern.

When manual collection starts slipping, or you want continuous AI brand tracking with the mention rate, citation rate, share of voice, and competitor citation reporting done for you on a schedule, DeepSmith covers ChatGPT, Perplexity, and Gemini together on the Scale plan, at $399 a month or $299 a month billed annually, with 200 tracked prompts. Pro tracks ChatGPT alone, and Grow adds Perplexity but not Gemini, so if all three of the engines in this guide matter to you, Scale is where that coverage starts. A 7-day free trial is available, with no long-term contract. Once your visibility data shows a gap, the same platform can turn that gap into a content backlog, so the monitoring work feeds directly into what you write next instead of sitting in a report nobody opens again.

Frequently asked questions

How often should I monitor my brand in ChatGPT, Gemini, and Perplexity?

Run your core prompt set weekly for most teams. Add a smaller daily watchlist if your category moves fast, you are in a launch window, or leadership wants near-real-time visibility. Review the full trend monthly, and keep any special diagnostic runs clearly separate from your regular series.

Should I use the exact same prompt in every AI engine?

Yes, for your core comparison set. Identical wording is what makes the answers comparable across engines. Record each engine's mode and settings separately, since ChatGPT, Gemini, and Perplexity retrieve and present information in different ways even when asked the same question.

Why did my brand show up yesterday but not today?

A number of things can cause that: a retrieval change on the engine's side, a different model or mode, a location or personalization difference, a source that changed, leftover conversation context, or simply a failed run. Rerun the exact prompt under the same conditions and read the full answer and citations before treating it as a real reputation signal.

Do I need a tool to monitor brand mentions across AI engines?

Not necessarily to start. A spreadsheet works for a small pilot as long as you preserve exact prompts, full answers, citations, timestamps, and run conditions. AI brand monitoring tools become worth it once manual collection gets inconsistent, or once you need scheduled runs, prompt-level history, page-level citations, and competitor comparisons across engines without doing it all by hand.