You keep asking yourself the same question: does ChatGPT cite my content, or am I invisible in there? Take a breath, because you can get a real answer today. This is a hands-on guide on how to test AI citations for specific pages, using the right prompts, the right engine settings, and enough repeat runs to tell a real citation from a lucky fluke. By the end, you will know how to check AI citations for any page you care about, and you will have a small spreadsheet with a defensible read of reality, not a vibe.
One thing first, and it is the thing most people get wrong. One run of one prompt is not data. Both engines are non-deterministic, they personalize answers, and they refresh their indexes on their own schedule. So the goal here is not a single dramatic screenshot. The goal is a controlled little experiment you can trust. Let's build it together, one step at a time.
Know the difference between a mention and a citation
Before you type a single prompt, get clear on what you are actually looking for. This is the step that saves you from fooling yourself later.
A citation only appears when the model decides it needs outside evidence to answer well. That decision has three quiet layers. First, will the model search at all? Both engines search only when the question needs fresh, comparative, or product-specific information, so an evergreen definition often triggers no search and no link. Second, once it searches, it reads a small handful of pages and picks which sentences are worth attributing. Third, it decides which of those sentences get a citation marker and which do not.
That last layer is where two very different outcomes hide. So write down three definitions now and hold them all the way through:
- Mention. Your brand or domain is named somewhere in the answer text, but it is not linked as a source.
- Citation. Your brand is linked as a source. In ChatGPT that means an inline footnote or an entry in the Sources panel. In Perplexity that means an entry in the numbered source list at the bottom, whether or not the little inline number shows.
- No appearance. Your brand is simply not in the answer.
These are not the same thing, and you should never blur them. A mention without a link is not a citation, even when the model clearly leaned on your page. Any serious tool tracks mention rate and citation rate as separate numbers, and so should you. You can only check AI citations honestly once these three buckets are crisp in your head, because every run you do from here gets sorted into one of them.
While you are here, fix one more thing: your threshold. Decide up front what counts as "present." A common call is one or more citations across your runs. Writing that down now stops you from quietly moving the goalposts once you see results you like.
Here is the good news. Once you can tell a citation from a mention, most of the confusion is already behind you.
See how each engine shows its sources
ChatGPT and Perplexity display citations differently, and if you do not know where to look, you will miscount. Two minutes of orientation here prevents that.
ChatGPT behaves like a hybrid. It decides whether to search, runs a query if it does, reads the results, and writes an answer that mixes conversational prose with citation markers. You will see superscript numbers attached to specific claims. When those inline markers are missing, look at the footer for a Sources button. Click it, and a panel lists every URL the model consulted. OpenAI's own help docs say exactly this: if inline citations do not appear, open the Sources panel beneath the response. That panel is your ground truth for what the model actually read. The visible markers are only the highlights.
One more ChatGPT detail matters. There is a small globe icon in the composer that forces web search on or off. If web search is off, the model answers from training data and never links out. So a blank result may not mean you were ignored. It may mean the model felt confident enough to skip searching entirely. Large-scale analysis of hundreds of thousands of real ChatGPT conversations found only about 18 percent trigger a web search at all. Record which mode each prompt ran in.
Perplexity is the opposite personality. It is built around citation. Almost every claim carries a numbered footnote, and the source list is always visible at the bottom. Two toggles change what you see. Pro Search runs multiple sub-queries and pulls more pages, so it produces more citations than the default mode. And the Choose Sources control (which replaced the older Focus selector) lets you pick Web, Org Files, or both. If you leave it on anything but Web, you silently change which pages are even eligible to be cited. Reset it to Web before every run.
You will know this step is done when you can open either engine and point to exactly where a citation would show up. Where people go wrong is trusting the prose and never opening the Sources panel, so they undercount.
Build your prompt set from real buyer language
Now for the prompts. This is where your test is won or lost, so slow down here.
The strongest prompts are the questions your buyers actually ask, phrased the way they actually phrase them. Not keyword-shaped strings like "best CRM SaaS," but conversational ones like "what's the best CRM for a 10-person SaaS sales team." Mine them from real sources: sales-call transcripts, support tickets, win/loss notes, review sites, and the People Also Ask box. Those are the words people use when they are searching without already knowing the answer.
Aim for 20 to 30 prompts, balanced across the funnel. A set clustered only at the bottom gives you a flattering, dishonest read. Spread them:
- Problem-aware: "how do I reduce churn in a subscription product"
- Comparison: "best CRM for SaaS sales teams under 10 people"
- Decision, brand-specific: "HubSpot vs Salesforce for early-stage SaaS"
- How-to: "how to set up lead scoring in HubSpot"
- Switching: "alternatives to Mixpanel for product analytics"
Two rules keep the set clean. Exclude any prompt that names your own brand, because those are biased toward you and tell you nothing about discovery. And drop any pure definitional prompt that returns the same answer no matter who is asking, because it cannot separate you from anyone.
Why 20 to 30 and not 200? Because prompt selection, not volume, is the real bottleneck. A well-worn practitioner rule holds that roughly 30 carefully chosen, strategically important prompts beat 200 random ones. Below 10 prompts you cannot see a pattern. Above 50 will not fit one focused afternoon. When you are done, you should have a spreadsheet with columns for prompt text, funnel stage, topic cluster, the competitors you expect to see, and a timestamp. That sheet is the backbone of everything that follows.
Lock your test environment before you run anything
Here is where a little discipline pays off enormously. If you change conditions between runs, you cannot tell whether a different result came from your page or from your own sloppiness. So hold everything constant.
Pin one ChatGPT model for the whole session and write down which one. Choose one Perplexity mode, Pro Search on or off, and stick with it. Set Choose Sources to Web. Use one account per engine, and if you can, a clean account with Memory off and Custom Instructions cleared, since ChatGPT quietly conditions answers on stored facts about you. Work in an incognito window so prior chats do not personalize the results. And run the whole set in one sitting, ideally the same day, logging your start and end time.
Anything you cannot hold constant is not a dealbreaker. It is just a variable you have to flag in your notes so future-you knows to distrust it. That honesty is the whole point of the exercise.
You will know this step is done when you could hand your settings list to a colleague and they could reproduce your exact setup. Where people go wrong is treating "I used ChatGPT" as a controlled environment. It is not. The model, the mode, the memory, and the browser all move the answer.
Run each prompt at least five times
This is the heart of the method, and it is also the step everyone is tempted to skip. Please do not skip it.
For each prompt in your set, run it five times in ChatGPT and five times in Perplexity, each in a fresh chat. Vary the wording slightly across runs, same intent, different phrasing. That tests whether the engine associates your brand with the topic itself, not with one exact string. Screenshot every result and save the response text, because these interfaces change their layout almost weekly and the screenshot is your durable record.
For every single run, capture the same fields: the prompt as you typed it, the timestamp, the engine and mode, whether there was a mention (yes or no), whether there was a citation (yes or no), the URLs cited, where the citation sat (opening sentence, body, or sources-only), and a one-line quote of the sentence doing the citing. It feels tedious. It is also the difference between evidence and anecdote.
A quick word on the traps that quietly corrupt results. These are the mistakes that make smart people draw wrong conclusions:
- Memory or Custom Instructions left on. ChatGPT conditions on stored facts about you. Turn them off, or your test measures your profile, not your page.
- Testing inside a Perplexity Space. Spaces lean on uploaded files and a custom system prompt, so they measure how well those files answer, not your web visibility. Use the main search box.
- The wrong Perplexity sources filter. Academic or Social restricts which pages can be cited. Reset to Web.
- Switching models mid-test. A run split across two ChatGPT models is not comparable. Pin one.
- Testing a page the day you publish it. Retrieval indexes lag, often by hours or days. Wait 48 to 72 hours before you test a brand-new URL.
- Trusting the Sources panel without clicking through. An academic audit of citation reliability found Perplexity's cited links were wrong or unsupported roughly a third of the time. Open a sample of the URLs and confirm they actually say what the model claims.
If five runs is genuinely too much right now, three is the bare floor to beat a coin flip. Anything fewer is a story, not a measurement.
Score and read your results
You have a full spreadsheet. Now let's turn it into three honest numbers per prompt, and three across the whole set.
- Mention rate = runs where your brand is named, divided by total runs.
- Citation rate = runs where your brand is linked as a source, divided by total runs.
- Position rate = runs where your brand shows up in the opening sentence or the sources panel, divided by total runs.
Average these across your prompts for the headline view, but always keep the per-prompt breakdown too. One prompt hitting 80 percent can hide ten prompts sitting at zero, and the average would lie to you.
Now read the numbers with a steady hand:
- A zero percent citation rate across five runs is a real signal that the engine is not finding or not trusting your page on that topic. A zero on a single run means nothing.
- One hit in five runs is noise. Treat anything under 20 percent as inconclusive and re-test later.
- When a competitor keeps showing up on prompts where you do not, your best diagnostic is their cited page. Open it. Note the structure: definitions near the top, FAQ blocks, comparison tables, hard numbers. That page is your baseline to beat.
- When your citation rate is high but you only ever appear in the sources panel, never in the prose, your page is being used as a reference but not quoted as an authority. That is a different problem from not being cited at all, and it needs a different fix.
Do not read "no citation" as defeat. If the model answered confidently without searching, it simply felt no need to. Re-test with a phrasing that forces retrieval: name a competitor, a specific price, a recent date, or a numeric comparison. You will often watch the citations appear.
Keep the emotional temperature low while you read all this. The point of learning how to test AI citations this carefully is not to grade yourself, it is to find the two or three topics where a small fix would move you from mention to citation. A rate of zero on a hard prompt is information, not a verdict. Circle those prompts, note which competitor owns them, and you have your next move already written for you.
Decide whether to make this a habit
Take a moment here, because you have earned it. You now know whether ChatGPT and Perplexity cite your pages, where your links land, and which topics leave you invisible. That is a genuine read of reality, and most teams never get this far.
So the real question shifts. Do you want to run this by hand every month, re-doing the runs, the screenshots, and the scoring? Or do you want it to run on a schedule while you spend your time on the pages that need fixing? This spot-check is the diagnostic. A tool is what the diagnostic looks like when it runs itself. DeepSmith is one option here: its AI Visibility module tracks mention rate, citation rate, share of voice, and visibility trend per engine, using the same prompts and the same metric definitions you just used by hand, collected repeatedly and charted for you. Coverage layers by plan, with ChatGPT on the entry tier and Perplexity, Gemini, and more added as you move up. Same test, no manual afternoon.
Whichever way you go, you have already done the hard part: you know what to measure and why.
What to do next
Start small. Pick five prompts this week, run them five times each on both engines, and fill in your sheet. That single afternoon tells you more than a month of guessing. The searches that brought you here, quiet worries like "does ChatGPT cite my content," "am I cited in ChatGPT at all," and "test if Perplexity cites me for our best page," all get answered by that one sheet. From there, decide which zero-citation prompts are worth chasing, read the competitor pages winning those answers, and rebuild your own around the structures the engines clearly prefer.
If you would rather not run this by hand every month, you can start a free DeepSmith trial and let the tracking run on a schedule while you focus on closing the gaps it surfaces. Either path works. The only wrong move is going back to guessing.



