If you have ever typed your own company name into ChatGPT just to see what comes back, you already know the feeling. Sometimes the answer is close. Sometimes it is missing the one thing you actually do best. This guide walks you through a structured, point-in-time audit AI brand description check: a way to check what AI says about my brand, capture the exact language several engines use about your company, compare it against the facts you stand behind, and end up with a short list of what to fix first. By the end you will have a repeatable method, not a one-off screenshot.
Step 1: Define the audit question and lock the conditions
Before you open a single chat window, write down the question this audit is trying to answer. It should be one sentence, something like "how does this engine describe our company to a buyer looking for a solution in our category" or "where does the engine's answer disagree with our approved facts." That sentence keeps you from drifting into a general fishing expedition once the answers start coming in.
Then set the conditions and write them down too: the date, your time zone, the market and language you are testing, and the list of engines you plan to check. Decide whether you will run each prompt in a signed-in account or a logged-out session, and use a fresh conversation for every prompt unless you are specifically testing how the engine handles follow-up questions. Do not rewrite a prompt after you see the answer it produced. If the wording felt wrong, that is data, not a mistake to correct.
It helps to remember what this audit is not. AI answers shift with the engine, the mode, whether search is turned on, your location, your account state, and which sources happen to be available that day. You are not trying to produce a lab result that anyone could reproduce forever, and this is not the same as an ongoing AI brand perception audit that tracks change over months. You are creating an honest, dated record of what a buyer would have seen if they asked the same thing.
Common mistake: running one casual query about your company in a personal ChatGPT account and treating the answer as your brand's complete AI reputation. A single branded query tells you whether the engine recognizes your name. It tells you almost nothing about how you show up for the actual problems your buyers are trying to solve.
Done when: you have a written audit sheet with a fixed date, time zone, market, language, the exact engine and mode list, your session condition, and a named person who can judge whether a claim is accurate.
Step 2: Map the buyer prompts before you query anything
The prompts you choose decide what the audit can and cannot tell you, so build them before you touch an engine. Start from real buyer questions, not from your company name. A reasonable first pass covers 20 to 30 prompts across your main use cases, audiences, and question types. Some teams go further and work with 30 to 50 prompts when they are testing several models at once. Treat both numbers as planning benchmarks, not a rule you have to hit exactly.
Build the set across a few categories so you are not just testing recognition:
- Unprompted category discovery: "what are the best tools for [category]"
- Problem and use case: "what should a [role] use to solve [problem]"
- Comparison: "compare [brand] with [competitor] for [use case]"
- Branded understanding: "how ChatGPT describes my company" or "what is [brand], who is it for, and what does it do"
- Capability and limitation: "what can [brand] do, and what can it not do"
- Decision-stage evaluation: "is [brand] a good fit for [specific situation]"
- Misconception check: "is [brand] a [category or attribute]"
- Alternatives and switching: "what are the best alternatives to [brand]"
Mix branded and non-branded prompts on purpose. Branded prompts tell you if the engine knows who you are. Non-branded prompts tell you whether a buyer who has never heard of you would ever land on your name at all, which is usually the more useful signal. Comparison and alternative prompts are where competitive framing shows up most clearly.
Once your list exists, tag each prompt with the buyer intent behind it and how much it matters commercially. Prioritize the prompts tied to revenue-bearing use cases, the segments you care about strategically, and anywhere leadership has already noticed a competitor showing up instead of you.
If building that first list from scratch feels slow, DeepSmith's Discover Prompts feature can generate a starter set of candidate questions straight from your product, persona, and buyer-stage context, so you are not staring at a blank spreadsheet. Treat the output as a draft: read every suggested prompt, edit the wording, and drop anything that does not reflect how your buyers actually talk.
Pro tip: lock the exact wording of each prompt for the whole audit and write it down character for character, including punctuation. Adding a qualifier like "best," "for startups," or "in the United States" can change which competitors show up and how the whole answer is framed, so a small edit mid-audit can quietly invalidate your comparison.
Done when: every prompt has a category, a buyer intent, a target audience or situation, and a reason it made the list. The set is broad enough to expose both what the engine already knows and what it is missing.
Step 3: Pick the engines and modes your buyers actually use
Choose engines based on where your audience actually looks, not on which ones are easiest to test. A practical starting set for most companies is ChatGPT, Google AI Mode or AI Overviews, Gemini, and Perplexity, with others added if your market leans on them.
These engines do not all behave the same way, and that matters for how you audit them. ChatGPT Search can pull in timely answers with links out to sources on the web. Google AI Mode generates a conversational answer with links for further reading and supports follow-up questions in the same thread. Perplexity builds its whole answer around cited sources and positions itself as a real-time answer engine. Gemini can ground its response in Google Search and cite verifiable pages when it does. A source might show up as an inline citation in one engine, a source card in another, and a linked phrase somewhere else, so you cannot treat every "citation" the same way across tools.
Write down exactly which product and mode you used for every result. "I checked ChatGPT" is not specific enough once you are comparing a search-enabled session against a plain conversation, or a standard Google result against Google AI Mode.
Common mistake: lumping ChatGPT Search, a regular ChatGPT conversation, Google AI Mode, and an ordinary Google search result into one bucket called "AI results." They pull from different sources and cite differently, so mixing them together hides real gaps instead of revealing them.
Done when: you have a declared engine and mode list with a reason for including each one, and every result you capture states exactly which surface it came from.
Step 4: Run every prompt and capture the answer word for word
With your prompt list and engine list fixed, run each prompt on each engine and record what comes back in full, not a summary of it. Open a new conversation for each independent test unless you are deliberately checking follow-up behavior.
For every result, save:
- The prompt ID and the exact prompt text
- The engine and mode
- The date, time, and time zone
- Whether the brand was named, and the exact phrases used to describe it
- Claims about category, audience, product, capability, limitation, pricing, geography, and differentiators
- Any competitors named or recommended, and where they appeared relative to you
- Every source or citation shown, including whether one of your own pages appeared
- Follow-up answers, kept clearly separate from the first response
Take a screenshot or export alongside the written record. A screenshot preserves things plain text loses, like where a source sat on the page, whether you were the first name mentioned or the third, and how much visual space the answer gave you compared to a competitor. Resist the urge to clean anything up as you go. Do not fix a misspelling, drop a hedge the engine included, or merge two separate source links into one, because those details are exactly what you are trying to measure.
This is also where DeepSmith's AI Visibility view earns a mention, because it already captures full answers and separates mentions from citations automatically, so a marketer running this process regularly does not have to rebuild the capture step by hand every time. The underlying answer text is what actually matters for a snapshot like this one; a metric on its own cannot tell you whether a description is fair.

Common mistake: recording only "mentioned" or "not mentioned" for each result. That single word throws away the language you actually need to check for accuracy, missing information, and framing. You are trying to catch a specific wrong sentence, not produce a checkbox.
Done when: someone who was not in the room could read your audit sheet and reconstruct exactly what was asked, where, when, and what came back, without reopening a single engine.
Step 5: Check each claim against an approved fact sheet
Before you judge any answer, put together a short reference sheet that your company is willing to stand behind: what you are, who you serve, your products and use cases, the segments or locations you support, your real capabilities, any limitations you would rather not gloss over, your approved differentiators, and claims you should never make. This sheet is your ruler. Without it, "wrong" just becomes a feeling.
Break each engine's answer into individual claims and check every one against the sheet. Mark each claim as:
- Accurate: matches the approved facts
- Partly accurate: broadly right, but incomplete, overstated, outdated, or missing a needed qualifier
- Incorrect: contradicts an approved fact
- Unverifiable: plausible, but not something your fact sheet actually confirms
Keep the quality of a citation separate from the accuracy of the claim next to it. An engine can cite a genuinely reputable source and still summarize your company wrong, and it can make an accurate claim with no citation attached at all. A confident tone and a linked source are not the same thing as being correct, so do not let either one talk you out of checking the underlying fact.
Done when: every claim that matters has a status next to it, and where something is wrong or missing, you have written down which approved fact it conflicts with.
Step 6: Score completeness and framing separately from accuracy
Once you know what is factually right or wrong, look at what a buyer would actually walk away understanding. This is a different question from accuracy, and it deserves its own pass.
For completeness, check whether the answer names the right category, describes the right audience, explains your main offering clearly, includes the capability that would actually drive a decision, and mentions a real limitation if one matters. An answer can be completely true and still leave out the one fact that would have changed a buyer's mind.
For framing, look at how you are positioned rather than what is technically stated. Are you presented as a serious option or an afterthought? Is a competitor given more space or stronger language than you are? Does the wording imply a price point, a maturity level, or a limitation your fact sheet does not actually support? Framing is not the same as tracking sentiment over months. This is a single snapshot, so record the wording and the emphasis you see today; you are not trying to establish a trend from one reading.
Use plain status labels instead of inventing a numeric score on the spot: complete or incomplete, fair or distorted, correct or incorrect. If your team later decides a numeric scale is worth adopting, define it before you review any more answers and apply it the same way every time, rather than adjusting it once you see which results you like.
Done when: your notes clearly separate "wrong," "missing," and "framed unfairly" as three different problems, because each one needs a different fix later.

Step 7: Turn your observations into a prioritized gap list
A brand AI answer audit is only worth the time if it ends in something usable, so the output here should be a short, specific backlog, not a vague impression. Create one row for every material issue you found, with fields like these:
| Field | What to record |
|---|---|
| Prompt ID | The buyer question that produced the issue |
| Engine and mode | Where it showed up |
| Exact wording | The relevant phrase, copied verbatim |
| Gap type | Accuracy, completeness, framing, mention, citation, or competitor displacement |
| Evidence | The fact-sheet line or source that supports your judgment |
| Buyer consequence | What a buyer could misunderstand or miss because of this |
| Priority | High, medium, or low |
Prioritize the rows that are flatly wrong, that showed up on a high-intent prompt, that repeated across more than one engine, or that put a competitor in the recommendation slot where you should have been. Resist the urge to solve anything here. This step produces a diagnostic backlog for someone to work from next, not a content plan or a set of fixes to ship today.
This is where DeepSmith's Pages and Competitor Citations views can help you organize what you already found: Pages shows which of your pages are actually earning citations, and Competitor Citations shows which competitor pages are winning the same prompts you tested, broken out by platform. They help you locate and prioritize evidence faster. They do not replace the judgment call you just made about whether a sentence was fair.
Done when: you have a short list of concrete, sourced observations, each one specific enough that a teammate could act on it without asking you what you meant.
What to do next
Treat this snapshot as the input to a longer conversation, not the end of one. If your only goal was to check what AI says about my brand once and move on, the gap list is your finish line: it tells you what is wrong right now, in one sitting, under conditions you wrote down. It does not tell you whether that description is getting better or worse, and it should not be dressed up as an AI brand perception audit that tracks change over time when it is really a single dated check. If a gap seems worth watching, that is the signal to move into a repeatable monitoring process rather than repeating this manual exercise every few weeks by hand.
If you would rather have the prompt, answer, citation, page, and competitor evidence organized in one place going forward, DeepSmith's AI Visibility workflow is built for exactly that. You can start a free trial and see your own answers captured the same way, without rebuilding the spreadsheet every time.



