A buyer asks ChatGPT what your product does, or Gemini which tool fits their use case, and the answer that comes back might be missing, out of date, or just wrong. This is what AI brand reputation management is for: an ongoing way to find those answers, fix what's broken in them, and give AI systems better information to work from. If you run marketing or content for a growing company, this playbook walks you through the whole loop, from picking the right questions to watch through verifying that a fix actually held.
What you need before you start: a short list of the AI assistants your buyers actually use (ChatGPT and Gemini cover most B2B audiences), and someone on your team who can own corrections when they come up. Whether the immediate job is to manage brand reputation ChatGPT surfaces or fix something Gemini keeps getting wrong, the underlying process is the same.
Define the AI answer questions that matter to your brand
Start by writing down the questions a buyer, a reporter, or a partner might actually type into an AI assistant. Not just "what is [your brand]," but the full range: who you're for, what you do differently, how you compare to the two or three tools people mention in the same breath, what you cost, and what your limits are. Include a few that are likely to surface complaints or outdated claims too, because those are the ones that actually threaten your reputation.
Write the questions in plain language, the way a person would ask them, not as search keywords. "What are the drawbacks of [brand]?" and "Is [brand] trustworthy?" belong on the list next to "What does [brand] do?" It also matters whether you're asking a branded vs unbranded prompt: a prompt that names your company will usually describe and cite you, but a category question forces you to compete just to show up, and that's where most of the real risk and opportunity sits.
Keep a stable core set of questions you check on a schedule, so you can compare answers over time, and a smaller rotating set for testing new ideas or new competitors. For each one, note the category (comparison, pricing, complaint-prone, factual) and the buyer stage it maps to. That's what turns "we should check this sometime" into an actual system.
You'll know this step is done when every important question has an owner and a reason it matters, and the list includes some neutral and negative-leaning prompts alongside the flattering ones. The common mistake here is only tracking your own brand name. If you never ask the category question, "which tool should a marketing team use for X," you won't see the gap where a competitor is winning ground you didn't know you were losing.
DeepSmith's AI Visibility area is built for this step: you define the prompts you want tracked, and Discover Prompts can generate a starter set from your product, persona, and buyer-stage context. Treat the suggestions as a starting point. The questions that matter to your reputation are the ones you and your team decide carry real business risk, not whatever a tool proposes by default.
Build a monitoring baseline across ChatGPT, Gemini, and Perplexity
Once you have your question set, run it across the assistants your buyers actually use, and save the full answer every time, not just a pass or fail score. For each result, record the engine, the date, whether your brand was mentioned, whether it was recommended or just listed, whether the answer cited one of your own pages, who else got named, and whether the tone read as positive, neutral, or negative.
This is where a lot of teams get their numbers wrong: mention rate, citation rate, share of voice, and sentiment are five different measurements that answer five different questions, and none of them can stand in for the others. A brand can be named constantly and still have a low citation rate, which means AI knows who you are but isn't using your own pages as evidence. It can also be cited a lot in answers that lean negative. If you only look at one of these AI visibility metrics, you'll misread what's actually happening.
Checking manually works for a spot check, but it breaks down fast as a system. One person typing a question into ChatGPT once a week can't separate a real trend from a one-off answer, because results shift with personalization, timing, and how the question is worded. That's the real difference between LLM reputation management done on a schedule and someone glancing at a chat window when they remember to. DeepSmith queries your tracked engines automatically and keeps the full answer, the citations, and the sentiment reading in one place, so you have a dated record to look back on instead of a memory of what you saw last month.
You'll know the baseline is solid when you can answer where you're visible, where you're absent, which of your pages get cited, and which competitors are showing up in your place. The mistake to avoid is reporting one aggregate number with no prompt, engine, or date attached to it. A score with nothing behind it can't be checked, and it can't be improved.

Audit what each answer actually says
Every flagged answer deserves a closer look before you act on it. Pull the exact wording, and sort it into a category: an outright wrong fact, missing context, outdated information, fair criticism, or just low visibility with nothing actually wrong. Those need completely different responses, so getting the category right matters more than reacting fast.
Once you know what kind of problem you're looking at, check the source behind it. Open whatever page the AI cited and see whether it actually backs up the claim, and check when it was last updated. Then compare that against your own pages, and against any independent sources covering the same topic. If your own About page says one thing and your pricing page says another, that inconsistency is often exactly what's confusing the answer. Google itself says that generative AI search features rely on pages that are crawlable, indexed, and eligible to appear in ordinary search results, so a page that isn't showing up normally has no path into an AI answer either. This part of the work overlaps closely with a standard AI citation audit, and if you haven't run one recently, a broader AEO audit is worth doing alongside this step so you're not fixing symptoms while the underlying access and structure problems stay in place.
Give each issue a severity: something touching pricing, safety, or a core capability claim deserves urgent attention, while a slightly stale description or an unflattering but fair line is lower priority. Assign an owner and a next action, whether that's a source correction, a report to the platform, a customer-facing response, or just watching it for now.
Common mistake: editing a page before you've figured out where the wrong information actually came from. If a stale fact is sitting on three different pages, fixing one of them and republishing does nothing for the other two, and the AI system may keep pulling from whichever one it already indexed.
You'll know this step is working when every flagged answer has a category, a source trail, a severity, and an owner attached to it, and your team can tell the difference between something that's actually false and something that's just an opinion they don't love.
Correct the facts at the source
Fix the underlying information before you try to influence what the AI repeats. Start with your strongest first-party page, usually an About page, a product page, or your documentation, and state the correct fact plainly. Don't bury it in marketing language, because a vague paragraph is harder for a retrieval system (and a skimming reader) to pull the right answer from.
Check that your other pages agree with the fix. Contradicting pricing or product details across your own site is one of the most common causes of an AI system repeating something wrong, since it has no way to know which of your pages is the current one. Where a third-party listing or profile is carrying the outdated claim, request a correction there too if the platform allows it.
Most AI products also give you a way to flag a bad answer directly: ChatGPT's search feature has a feedback control on individual responses and warns that results can be incomplete or outdated, Perplexity has a report option, and Google's AI Overviews and Knowledge Panels have their own feedback and edit-suggestion paths. Use them, but treat them as a signal you're sending, not a fix that takes effect immediately. Platforms don't retrain on your feedback in real time, and the correction that actually holds is the one on your own page.
Keep a simple change log: the old claim, the corrected one, which page changed, and the date. Then give it time and rerun the affected prompt rather than assuming the job is done the moment you hit publish. The biggest risk to a brand's positioning in AI answers usually isn't one dramatic error, it's small inconsistencies like this compounding quietly across pages nobody rechecks.
Respond to negative sentiment without hiding the truth
Not every negative-sounding answer is a problem to fix, and treating all of them the same way is a mistake. Sort what you're seeing into four buckets: a false or outdated claim, which you correct and report; valid criticism, which you address by actually fixing the underlying issue; missing context, where a fuller explanation would help; and competitive framing, where a rival simply has stronger proof on that specific point.
For valid criticism, the instinct to publish a rebuttal is usually the wrong one. If a customer complaint is fair, the fix is fixing the thing they complained about and then documenting that you did, not writing around it. AI systems, like readers, tend to notice when a response dodges the actual issue. A negative mention isn't automatically a loss either, since the framing around a mention matters more than whether your name showed up at all, but a mention that repeats something false or unfair does need a real response.
Sentiment itself needs care as a measurement. Score each answer as positive, neutral, negative, or mixed based on the language actually used, not just whether you appeared. Keep the original wording behind every score, and have a second person check the classification on anything that matters, since sentiment reads as a judgment call more often than teams expect.
Pro tip: keep a stable core prompt set for tracking sentiment over time, and a smaller rotating set for new competitors or buyer language you're starting to see. That gives you a comparable baseline instead of a program that resets every time something new comes up.
A dedicated AI brand sentiment tool can flag which answers to look at first, but the label is a prioritization signal, not a verdict. Always read the actual answer before deciding what it means for your reputation, and never let a dashboard score stand in for the words an AI system actually used.
Reinforce the narrative you want AI to repeat
Once you've corrected what's wrong, the next job is making sure there's good material for AI systems to find when they go looking for the rest of the story: what you're for, what you do differently, where you fit, and where you honestly don't. This is the part of managing brand reputation in AI answers that's about building, not just fixing.
Match what you publish to the actual gaps your audit turned up, not to what a competitor happens to have written. If you're absent from a category question, write a genuinely useful explanation of that category and where you fit in it. If you're mentioned but rarely cited, strengthen the specific page that should be the evidence for that claim. If a competitor owns a comparison prompt, an honest comparison with real evaluation criteria does more work than a page that just asserts you're better. Google's own guidance on helpful, people-first content backs this up: original information, real expertise, and a clear point of view outrank pages written mainly to repeat a phrase.
The content that earns citations tends to share a few traits: the direct answer near the top, clear headings, specific claims instead of vague ones, and enough original detail that it isn't interchangeable with ten other pages saying the same thing. This lines up with how answer engines judge trust and authority: they're looking for the same signals of real expertise and original information that search engines always have, applied to whether a source is safe to cite. Thin pages written to repeat a keyword, or content copied from a competitor's structure, tend to do the opposite of what you want here. One frequently cited generative engine optimization study found that tactics like citing sources and adding statistics improved specific visibility metrics in its own evaluation, evidence that concrete, well-sourced content tends to perform better, not a guarantee of the same lift for every brand.
This is also where a production system pays off, since publishing four or five evidence pages a quarter by hand is a slow way to close real gaps. DeepSmith's Content Map shows coverage gaps and untapped topics against your competitors, Opportunity Agents turn those gaps into ideas with the specific data point attached, and Content Studio can take an idea from a backlog entry to a published, on-brand article without you rebuilding the brief from scratch each time. The output still needs your judgment on what's worth writing and why, but the manual assembly work goes away.
Verify the change and keep the loop running
A correction, a customer fix, or a new page doesn't count as done the moment it's live. Give the change some time (a source has to be recrawled, and a platform's index has to catch up), then rerun the exact prompt that originally flagged the problem, on the same engines. Compare the new answer to your baseline: has the fact changed, has the sentiment shifted, is the citation showing up where you wanted it. If a fix didn't hold, that's useful information, not a failure, and it usually means the underlying source still needs more work.
Feed whatever you learn back into your prompt set. New competitors show up, buyer language shifts, and a question that didn't matter last quarter sometimes matters a lot this one. A workable rhythm looks like this: check your core prompts on a regular schedule, do a weekly pass on anything new or negative, review trends monthly across engines and topics, and once a quarter, retire stale prompts and add fresh ones. This is what separates real AI visibility measurement from a one-time audit that gets screenshotted once and forgotten.
You'll know the loop is working when you can point to a specific issue, the action you took, the answer that came back afterward, and the date you checked it. That record is what makes the whole program defensible, both to your own leadership and to yourself the next time someone asks whether any of this actually moved the needle.
Where to go from here
You can't control brand narrative AI systems produce with a single fix, but you can steer it over time. None of these steps work in isolation for long. A prompt set that never gets monitored goes stale, a monitoring habit with no audit step just produces anxiety, and an audit with no correction step is a list nobody acts on. Treat this as one connected loop: define, monitor, audit, correct, respond, reinforce, verify, and back to the top. Pick one step you're currently skipping and build it into next week's schedule before adding anything else.

If you want to see this working end to end on your own brand, a DeepSmith free trial gets you real tracked-prompt data and a first look at where you're actually showing up in AI answers, before you commit to anything.



