DeepSmith

Sep 26 · AEO & AI Visibility

16 min read

Are AI Visibility Scores Reliable? What Day-to-Day Noise in AI Answers Means for Your Tracking

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
Abstract illustration of scattered gray data points converging into a smooth white trend line on a dark background, with the text AI Visibility: Signal or Noise, representing how noisy repeated measurements resolve into a clearer signal over time.

You check your AI visibility score on a Monday and your brand shows up in six out of ten answers. You check the same prompts on Wednesday and it's four out of ten, with nothing else different: no new content, no site change, nothing you did. Here's the verdict up front: AI visibility scores are useful as repeated, transparent estimates of how often your brand shows up across a defined set of prompts and engines, but a single score or a single ranking snapshot is not a reliable measure of your real-world AI visibility reliability. The evidence on this is mixed in a specific way. It strongly supports that day-to-day AI ranking volatility is real and comes from how these systems generate and retrieve answers, not from bugs in your tracker. It does not support any one fixed rule for how many prompts, checks, or runs make a score dependable. This piece walks through why the same prompt gets a different answer, what the direct research on AI visibility score accuracy actually found, and how to read your own numbers so ordinary LLM output variance doesn't send you chasing a problem that was never there.

What an AI visibility score is actually measuring

An AI visibility score isn't a census of every question a real person might ask an AI engine about your category. It's an estimate built from running a defined set of prompts through one or more AI engines and recording what happened: did the engine name your brand, did it cite your pages, where did you land relative to competitors.

Each prompt run is really one observation. For a given prompt, on a given engine, on a given day, you get a yes or no on whether you were mentioned, a yes or no on whether you were cited, a position if you showed up at all, and a read on who else was in the answer. A score is the aggregate across many of those observations, which makes it closer to an estimated probability than a fixed rank. If your mention rate is 40 percent, that means your brand turned up in 40 percent of the observations your tracker recorded. It doesn't mean 40 percent of everyone who asks that question in the real world will see you, and it doesn't mean you hold a permanent fourth spot anywhere.

A few metrics get lumped together under "AI visibility" and it's worth keeping them separate, because each behaves differently:

  • Mention rate, how often the engine names your brand at all.
  • Citation rate, how often the engine links to or identifies your pages as a source.
  • Share of voice, your visibility relative to competitors under the tracker's own method.
  • Rank position, where you land within a particular answer or list.
  • Visibility trend, how an aggregate metric moves over time.
  • Sentiment, how the answer talks about you, which usually runs through a separate classifier and can add its own error on top of whatever the model itself is doing.

None of these numbers means much without its measurement frame attached. Before you trust a score, it's worth asking which prompts fed it, how many, on which engines, how often, whether the same prompts were reused each time, and how mentions and citations got detected in the first place. That measurement frame is the whole basis for AI visibility reliability: the same raw number can be a defensible estimate or a coin flip depending on what sits behind it.

Why the same prompt can get you a different answer

The most direct explanation for AI ranking volatility sits at the generation layer. An AI answer gets built one token at a time, and temperature controls how much randomness goes into picking each one. Google's own documentation on content generation parameters says that even at temperature zero, the model still selects the highest-probability token at each step but some variation can still occur. Anthropic's platform glossary is more direct about it: identical inputs can produce different outputs across API calls even with temperature set to zero. So temperature zero is better described as lower-variance, not as a guarantee of an identical answer every time. OpenAI's own documentation on reproducible outputs reaches a similar place. Its API is non-deterministic by default, and while a fixed seed plus unchanged parameters can get you outputs that are mostly identical, OpenAI is explicit that determinism still isn't guaranteed.

That matters for tracking because a small difference early in generation can change what shows up at the end: a different brand in the list, a different order, a citation that appears in one run and not the next. This kind of LLM output variance sits at the generation layer itself, before retrieval or ranking even enter the picture.

Temperature is only one lever, and it's not even the main one on a lot of these swings. Research on randomness in large language models points to several other sources sitting underneath it: deliberate sampling choices, silent model updates the provider ships without announcing, server-side batching, hardware differences between requests, model splitting or routing across a mixture-of-experts system, and small numerical rounding differences depending on the order operations run in. A tracker that holds the visible prompt constant is not holding the whole measurement process constant, because plenty of what shapes the answer sits behind the API and isn't exposed to whoever is running the check.

There's a related wrinkle worth flagging: a stable model name doesn't guarantee a stable system behind it. Providers can and do change what's running behind an unchanged public label, and a dated model identifier still won't tell you about a changed system message, a changed filter, or a changed routing rule. Anthropic's own deprecation notice, published February 25, 2026, is a concrete example: Claude Opus 3 was retired as of January 5, 2026. If your tracking crossed that boundary, you crossed a measurement-regime change, not necessarily a change in how your brand is doing. A responsible read of your own numbers means recording the model or engine version wherever a tracker exposes it and flagging the point where a version changed, rather than reading every dip after that point as a content problem.

Retrieval adds another layer entirely, and it happens before generation even finishes. Many AI search products decide, per query, whether to search the live web at all, and if they do, which documents to pull in and in what order. That retrieved set can shift in composition, order, and freshness from one run to the next, and the model then writes its answer from whatever context it got handed. This is not a hypothetical: a 2026 AI-search stability study found that identical prompts submitted at the same time still produced meaningfully different cited sources and brand mentions, which rules out "the content changed between checks" as the explanation. Google's own framing of Search results as dynamic, changing with the open web, evolving ranking systems, and several broad core updates a year, is a useful adjacent point here too. It doesn't prove every AI-answer change traces back to a Google-style update, but it does establish that retrieval systems are moving targets, not fixed databases you're reading from.

What the direct research says about AI visibility score accuracy

SparkToro published research on January 27, 2026 that ran 2,961 total prompt runs across ChatGPT, Claude, and Google's AI Overview or AI Mode, using 12 prompts and roughly 600 volunteers, with about 60 to 100 runs collected per prompt. The finding: ChatGPT and Google's AI systems had less than a 1-in-100 chance of returning the exact same brand list across repeated runs on the same prompt. Claude repeated its list slightly more often but was less likely to keep the same order. Getting the same list in the same order twice was roughly a 1-in-1,000 event across the tools tested. SparkToro's own conclusion wasn't that AI visibility measurement is pointless, it was that a precise rank position is a weak metric on its own, while a visibility percentage measured across dozens or hundreds of prompt runs is a more defensible directional read. Worth noting: this study didn't test whether the same instability shows up in the consumer interfaces people actually use day to day versus the API calls researchers ran, and its authors themselves called for larger prompt panels and more rigorous statistical treatment before anyone treats the numbers as settled.

A separate study, with data collected January 24 through March 20, 2026 across four verticals (telecommunications, real estate, sporting goods, and consumer electronics), took a more granular statistical approach. It ran eight prompts per vertical across ChatGPT, Gemini, Google AI Mode, and Perplexity, using both daily observations over roughly six weeks and simultaneous repeated prompts designed specifically to separate random noise from a genuine trend over time. Its two stability measures, Jaccard similarity for source overlap and Rank-Biased Overlap for order sensitivity, found that consecutive-day cited-source overlap averaged only 0.34 to 0.42 across the four verticals, meaning roughly 60 to 65 percent of cited sources turned over between one day's check and the next. The rank-sensitive measure came in even lower, at 0.21 to 0.26, meaning the order of what did show up also moved around. Brand-set overlap between consecutive checks ran 45 to 59 percent. The study modeled how much of that settles down as you add repeated runs: a single run carried a standard error of 0.370, which fell to 0.081 at seven runs and 0.062 at eight, at which point the reported confidence interval tightened to roughly plus or minus 0.12. Its authors recommend at least seven runs per prompt per day for brand-visibility monitoring and at least eight for source-level coverage, but that's a recommendation sized to this study's own four verticals and four engines, not a universal number for every category. The study also found ChatGPT skipped web search entirely on 57.8 percent of its runs, meaning a chunk of "no citation" results in any dataset like this can reflect the engine's own decision not to search that query, not a judgment about the brand.

Horizontal bar chart showing the standard error of a visibility estimate falling from 0.370 with one run per prompt to 0.081 at seven runs and 0.062 at eight runs, illustrating how repeating a prompt several times shrinks the noise around the estimate far more than one additional run does.

A 2026 production dataset from Ranqo, a company that builds AI visibility tracking, reported figures at a much larger scale: 102 brands, over 3,500 tracking runs, and more than 100,000 prompt responses across five engines. That's useful for showing what repeated measurement looks like at production volume, but it should be read as vendor-produced evidence, not independent validation. The paper's author has an equity stake in Ranqo, and the pipeline and dataset are the company's own production systems. Its own stated limitations include no visibility into how the underlying engines actually retrieve or retain information, no test of whether its findings hold across different model configurations, and sentiment results that mix genuine answer variation with error from the separate classifier used to score sentiment.

A separate line of research, from the biomedical AI literature, draws a distinction worth borrowing here: repeatability, meaning agreement across runs under the exact same conditions, versus reproducibility, meaning agreement across conditions that were deliberately changed. That work also found, in its own domain, that a more consistent answer wasn't automatically a more accurate one. The transferable point for AI visibility tracking is the same: a stable score tells you the measurement is settled, not that the measurement is correct.

Telling ordinary noise from a real signal

None of this means every score movement is meaningless, and it doesn't mean you should ignore your dashboard either. It means treating a single number differently depending on where it came from.

A movement is more likely to be ordinary AI search measurement noise when it shows up in just one check, comes from a single prompt or a very small panel, affects one engine while comparable engines stay flat, or sits inside the normal run-to-run range your own tracker has shown you before. It's also a signal to discount when the panel size, detection method, or sample composition changed at the same time as the score, or when the "drop" is really a rank change caused by list order shifting or one extra competitor getting pulled into that particular answer.

A movement earns more trust as a real signal when it holds up across several scheduled checks using the same prompt panel, shows up across multiple related prompts rather than just one, survives a resampling check with a slightly different prompt set, and appears across the engines that actually matter to your audience rather than just one. It matters more when it's larger than the normal variation you've already seen for that metric, and when your tracker gives you enough raw observations to actually look at what changed rather than just a single number.

Even a persistent movement doesn't automatically tell you the cause. A model update, a retrieval change, or a change in how the tracker detects your brand can all produce a real, lasting break in the data without any actual change to your brand or your content. The safest way to read your own numbers is to treat the score as a trend estimate rather than a live truth reading, avoid reacting to any single number in isolation, compare like with like across checks, and look at the prompt-level detail behind the aggregate whenever your tool lets you. Most of what looks like a sudden crisis in your dashboard is ordinary AI search measurement noise working its way through a small sample, not a real loss of ground.

How to read a tracker's numbers without overreacting

A tracker worth trusting will tell you what's actually behind its number: the size of its prompt panel, how those prompts were chosen, whether it repeats the same prompts or rotates them, how often it samples, which engines and answer modes it queries, what counts as a mention versus a citation, and whether you can get to the raw prompt-level results rather than just the rolled-up score. Guidance published July 24, 2026 by Clique Studios makes a related practical point: 25 prompts sampled once will read noisier than 500 prompts sampled weekly, which is really just the sample-size math showing up in a real dashboard. The same guidance recommends monthly measurement as a reasonable default for most teams, because it smooths out run-to-run noise while still fitting a normal reporting cycle, with weekly checks reserved for periods right after a known model change or during an active content push, and quarter-over-quarter comparisons doing the heavy lifting for any real strategic call.

This is where a platform's design choices either help you or work against you. DeepSmith's AI Visibility module tracks a defined panel of prompts across up to ten AI engines depending on plan, and reports mention rate, citation rate, share of voice, and trend over time, broken out per platform, with a full answer history behind each prompt rather than just a rolled-up score. The point of exposing that detail isn't to promise a number that never moves. It's that a score you can actually inspect, prompt by prompt and check by check, is the only kind you can tell apart from a bad day.

There's no universal number that tells you your setup is reliable. There's no study that establishes one fixed noise percentage, one fixed prompt count, or one vendor's score as the definitive measure of AI visibility. Studies to date haven't fully confirmed that what happens through an API matches what a person sees typing the same question into a consumer app, and a stable answer still isn't the same thing as a correct one. A rank change on its own doesn't prove a competitor got better or that you got worse. Treat all of that as the honest state of the evidence rather than a gap in your own understanding.

What would change this verdict

The verdict here holds until the research catches up in a few specific ways: a shared, cross-vendor standard for prompt count and run frequency, direct comparisons between API-level and consumer-interface variance, and studies that separate genuine sentiment shifts from classifier noise rather than reporting them together. Until then, the honest position is that AI visibility scores are directionally useful and single snapshots are not, and that gap is exactly what a repeated, transparent measurement design is for.

If you want a place to start, pull up your own tracker's methodology page before you look at this week's number. Check how many prompts feed your score, whether they're the same prompts each time, and whether you can see the raw runs behind the aggregate. That fifteen minutes will tell you more about whether Wednesday's dip is real than staring at the number itself ever will.

Frequently asked questions

Why does my AI visibility score change every time I check it?

Because the AI engine may generate a different sampled response, pull in a different set of sources, run on a changed backend or model version, or simply expose your tracker to a small sample whose composition shifted between checks. A change in the number doesn't automatically mean your brand or your content changed.

Are AI search rankings reliable?

A single answer's ranking position isn't reliable enough to act on by itself. A repeated visibility estimate, built from a stable prompt panel run multiple times with transparent methodology, is far more useful for spotting a real trend.

Does setting temperature to zero stop AI ranking volatility?

No. Anthropic's and Google's own documentation both say temperature zero doesn't guarantee an identical output every time. Temperature is one source of variation among several, and hosted systems can still change models, routing, retrieval, or answer mode underneath it.

How many prompts do I need for a reliable AI visibility score?

There's no universal number. More observations generally reduce sampling variance, but how many you need depends on how volatile the topic and engine are, how diverse your prompt set is, how many engines you're combining, and how precise you need the answer to be. One 2026 study recommended at least seven repeated runs per prompt for its own brand-monitoring design, but that number was sized to that study, not handed down as a rule for every category.