DeepSmith

Jul 26 · AEO & AI Visibility

17 min read

How to Measure Whether Your AEO Efforts Are Actually Working

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
A monochrome charcoal cover with the white headline "Is Your AEO Working?" surrounded by abstract trend lines, small bar-chart fragments, and connected citation nodes.

You shipped the schema fixes. You rewrote the answer sections. You published the comparison page. And now you are staring at a dashboard asking the only question that matters to your boss: is my AEO working, or did the numbers just wander on their own?

That question is fair, and it is hard, because AI answers move even when you do nothing. A brand gets named in one run and skipped in the next. A page gets cited this week and forgotten the next. So the honest answer to "did my work move anything" is almost never a screenshot. It is a method.

Here is the good news. You do not need a data-science team to prove AEO impact. You need a fixed set of prompts, a patient baseline, a change log, and a few control checks that separate your effect from the noise. This guide walks you through nine steps to measure AEO results you can actually defend to leadership. Take it one step at a time. By the end you will know how to tell a real lift from a lucky reading.

First, understand why AI visibility is so noisy

Search ranking is a position. AI visibility is a probability. Ask the same engine the same question twice and it might name your brand, cite your page, drop you entirely, or reword the whole answer. That is not a bug in your tracking. That is the medium.

Three kinds of noise are worth naming, so you stop mistaking them for progress.

The first is week-to-week churn that is mostly nothing. BrightEdge found that 96.8% of cited domains showed no week-over-week change at all. Of the small slice that did move, most moved down, not up. So most weeks, a flat reading is the expected reading.

The second is longer-term turnover in who gets cited. Profound reported large month-to-month rotation in citation domains: 54.1% on ChatGPT, 40.5% on Perplexity, and 59.3% on Google AI Overviews in one mid-2025 comparison across roughly 80,000 prompts per platform. The citation pool changes a lot, even without your help.

The third is run-to-run variance on the exact same prompt. Practitioner repeat-tests have reported answers swinging by roughly 10% to 34% on identical inputs. Treat that as a working warning, not a law, and let it lower your trust in any single reading.

There is a decay layer too. Scrunch, studying about 3.5 million citation events, estimated an average citation half-life of roughly 4.5 weeks, closer to 3.4 weeks on ChatGPT and 5.8 weeks on Perplexity. So a proof method has to outlast the decay, not race it.

If that feels like a lot, breathe. All of it points to one discipline that protects your read on AEO effectiveness: never trust a single reading, always compare the same prompts over time, and always ask what else could have moved them. The nine steps below are that discipline, made concrete.

Step 1: Build a fixed buyer prompt set

Everything downstream depends on this. If your prompt set drifts, your comparison is meaningless, because you are measuring different questions before and after. So build the set once, carefully, and freeze the core.

Aim for a balanced library of roughly 80 to 150 prompts. Cover five dimensions: persona, industry or vertical, use case or product fit, compliance or framework, and competitor comparison. Tag every prompt with a funnel stage (awareness, consideration, decision, expansion) and an intent (educational, pain point, pricing, comparison, evaluation). Include at least three head-to-head prompts against named competitors, because those are the ones leadership cares about most.

Where do good prompts come from? Not from an AI brainstorm alone. Pull them in this order: sales-call and interview language first, then Search Console queries with impressions but weak clicks, then validated keyword research, then Reddit and support tickets, and only last, AI-generated starters as a supplement. Copy the words your buyers actually use.

How to tell it is done: every prompt has a stable ID, every important persona and revenue vertical is represented, find, compare, and buy intents are all present, and you could rerun the exact set months from now without editing it.

Common mistake: freezing the list forever, or generating the whole thing in one AI session. Recheck your prompt strategy around every 90 days and consider a fuller refresh near 180. When you add prompts, add them as a separate cohort so your original measured core stays comparable. The core is your ruler, so do not stretch it.

Step 2: Set your baseline before you change anything

You cannot show a "before and after AEO measurement" if you never captured the before. This is the step teams skip when they are excited, and it is the one that later makes their whole case fall apart.

Run the complete prompt set on every target engine, on the cadence you plan to keep. For each prompt, record whether the brand was mentioned, whether it was cited, the cited URL and page type, any competitor mentions, the description's accuracy, and the engine, date, and prompt version. Then hold that baseline for at least four weeks. Six is more defensible when you can afford it, because the most volatile engine estimates decay in roughly three to four weeks, and a baseline is not one collection. It is repeated observations that show your normal range before you touch anything.

This is one place a platform earns its keep. DeepSmith's AI visibility tracking runs a defined prompt set on a schedule and reports mention rate, citation rate, share of voice, per-platform breakdowns, competitor citations, and full prompt-level answer history. Its Discover Prompts feature can even generate a starter library from your product, persona, and buyer-stage context, so you are not building the set from a blank page. The point is not the tool. The point is that a real baseline needs repeated, structured collection, and doing that by hand across engines gets old fast.

How to tell it is done: every prompt has a four-to-six-week history, every engine has comparable coverage, and you can state your baseline mention rate, citation rate, share of voice, and competitor visibility as ranges, not single points.

Common mistake: calling one week's readings a baseline. Normal churn alone can create a fake "movement" before you have changed a single thing.

Step 3: Pick a cadence that samples the signal

How often should you read? Often enough to catch the trend, on the same schedule across engines so your comparisons stay fair.

Daily collection is enough for most cases. Weekly is the practical floor. Three times a week is a strong middle that catches shifts faster without draining your budget. A useful rule of thumb is three to seven reads per citation half-life window, which for a fast-decaying engine lands around one read every few days. That heuristic is a way to reason about frequency, not an industry standard, so adjust it to your budget.

One firm rule: use the same cadence on every engine you compare. Reading ChatGPT daily and Perplexity weekly quietly biases the comparison, and you will not notice until someone smart asks about it.

How to tell it is done: the schedule is live, identical prompt versions run across your chosen engines, collection has already covered at least one baseline half-life, and you know what the cadence costs.

Common mistake: cranking up frequency to rescue a messy design. More readings do not fix a stacked change or a drifting prompt set. Frequency samples the signal, it does not create one.

Step 4: Change one thing at a time

This is where AEO effectiveness is won or lost, and it is almost boring how simple the rule is: isolate one intervention against one prompt cluster, then leave everything else alone.

Pick a single meaningful change. Restructure one answer section. Add or correct schema on a page type. Clarify product terminology. Strengthen author and organization details. Publish one targeted page. Then match that change to the prompts it could plausibly move. Schema belongs against prompts where structured understanding matters. Product-clarity edits belong against product-fit and comparison prompts.

For every intervention, write down the hypothesis, the exact change, the affected URLs, the target prompt cluster, the engines you expect to respond, the start date, the control prompts you left untouched, and any other release, migration, or campaign happening at the same time. That last line matters. A site migration in the same window can masquerade as an AEO win.

How to tell it is done: your change log can answer four questions cleanly. What changed, when, what it was meant to affect, and what stayed the same. If you cannot answer those, you do not have a result. You have a hypothesis.

Pro tip: add the change log before you add another dashboard. A dated record of what shipped, which prompts it targeted, and which you deliberately left alone is the foundation of attribution. Almost everyone reaches for more charts first. Reach for the log.

Common mistake: stacking five changes in one week. If the title, schema, internal links, and three new pages all move together, you have a bundle, and you can only honestly report it as a bundle.

Step 5: Run the before and after AEO measurement

Now the comparison you have been building toward. Take the same prompt IDs, the same engines, the same cadence, the same competitor definitions, and the same rules, and compare the window before your change to the window after. Then wait. Give it at least one citation half-life before you judge the result, and a second window if you can, because a one-day spike against a six-week baseline is not evidence.

Report it in a way a skeptic can check. You might frame brand presence as, for example, 10% of prompts before versus 40% after. You might frame competitive wins as two of twelve core prompts before versus eight of twelve after. You might frame accuracy as a share of correct descriptions before versus after. Those are illustration formats, not promised outcomes, and your real numbers are whatever your data says. Always show the sample size, the date ranges, the engine, the raw counts, and the percentages.

One number trap catches almost everyone. A move from 10% to 40% is a 30 percentage-point increase and a 300% relative increase. Both are true, and they sound wildly different. Report both when it helps, and never quietly swap one for the other to make a slide look better.

DeepSmith's trend reporting supports these period-over-period comparisons, its prompt history shows the individual answers that changed, and its Pages view ties citations back to the exact URLs and prompts that earned them. That page-level view stops you from crediting an aggregate lift to the wrong page.

How to tell it is done: your report names the fixed prompt set, the baseline window, the change date, the waiting period, the post-change window, before and after counts and percentages, the engines, and a per-cluster breakdown.

Common mistake: switching prompt sets between periods, reporting only the flashy relative number, or counting more total citations as a win when you quietly added prompts or engines.

Step 6: Add holdout and control logic

Here is the uncomfortable truth about a clean before-and-after: it shows movement, but it does not prove your change caused it. Engine refreshes, competitor moves, seasonality, and broad market interest can all lift your numbers while you take the credit. Controls rule those out. They turn "we think" into "we can show."

You have a few options, and you can stack them:

A prompt-matched holdout is the workhorse. Pair your targeted prompts with structurally similar prompts you deliberately did not optimize, matched on stage, intent, topic, competitor presence, and baseline visibility. If the treated prompts move and the holdout stays flat, your case gets much stronger.

An engine-matched check asks whether the same topic moved on other relevant engines. Agreement helps, but be careful: if every engine jumps at once, that can signal a market-wide shift rather than your edit.

A competitor control tracks rivals on the same prompts. If every brand rises together, you are probably watching the tide, not your work.

For broader campaigns, geo holdouts and synthetic-control methods like GeoLift borrow from marketing incrementality testing: withhold the treatment from a comparable control group and compare outcomes. AI engines give you no native AEO holdout, so you build the control yourself from matched prompts and pages. Around 30 days is a sensible floor, though decay may push you longer.

How to tell it is done: for every candidate lift, you can point to the untreated matched prompts, say which engines moved, show what competitors did, and state whether the control held. A lift with no control evidence is an observed association, not proof, and you should label it that way.

Common mistake: picking a control that was always going to move on its own, or waving off controls because your tool lacks a holdout button. The button not existing does not make the logic optional.

Step 7: Apply decision rules for noisy results

You will not always get a clean answer, so decide your rules before you look at the data. Then you are grading the result, not rationalizing it.

Keep a short rulebook:

  • No conclusion from a single weekly reading.
  • Require a rolling window that covers at least one relevant citation half-life.
  • Confirm the move on at least two buyer-relevant engines when you can.
  • Require the lift to survive into a second observation window.
  • Analyze the targeted cluster on its own, not buried in aggregate visibility.
  • Record both raw counts and percentages, every time.

Watch the multiple-comparisons trap too. Test five metrics and celebrate whichever one crosses a line, and you have basically guaranteed a false positive somewhere. Hold a stricter bar, or correct for how many things you tested, and do not crown a winner because one metric twitched.

A simple grading scale keeps everyone honest. Call a result Observed when it shows in the raw data but lacks duration or controls. Promising when it clears normal noise and lasts one window but the control evidence is thin. Supported when it survives a second window, appears on the right surfaces, and the controls back the story. Inconclusive when the data is sparse, the prompts changed, the control moved, or the change was stacked. Most real results start as Observed. That is fine. The label is the honesty.

Common mistake: treating a vendor dashboard as ground truth without asking how it got its number. Is it observed from repeated live prompt runs, or inferred from something else? Ask before you cite it upstairs.

Step 8: Package the proof for leadership

Your CFO does not want your methodology. They want to know what moved, why it is not luck, and what you will do next. So report in four short blocks.

Block one, what changed and when: the intervention, the URLs, the prompt cluster, the baseline dates, the post-change dates, and the engines. Block two, what moved: mention rate, citation rate, share of voice, accuracy, with raw counts, percentages, and percentage-point changes. Block three, why it is not coincidence: your matched controls, competitor results, cross-engine confirmation, the elapsed half-life, and second-window survival. Block four, what happens next: continue, expand, stop, or run a larger test, plus the downstream business signal you expect and its likely lag.

For senior readers, lead with outcomes, not signals. Start with pipeline, bookings, or qualified leads where you have them. Then show the visibility movement that could explain them. Then show the clusters and pages where the work happened. Visibility is a leading indicator, not a revenue receipt, so connect the dots without claiming a causal chain you have not tested.

How to tell it is done: someone outside your team can read it in two minutes and come away knowing the change, the movement, the strength of your controls, and the next decision.

Common mistake: presenting percentage-point changes with no controls, or implying that a citation bump already became traffic and revenue. Say what you proved. Flag what you did not.

Step 9: Turn measurement into a recurring system

One proven result is a nice slide. A system is what compounds. So set three rhythms and let them run.

Collection: run the stable prompt set on your chosen schedule. Review: inspect cluster movement weekly against your change log. Reporting: ship a monthly or per-intervention proof package.

Underneath those rhythms sits one loop: find a visibility gap, create or improve the content that closes it, measure the same prompts, read the result, then decide to continue, stop, or adjust. This is where the whole platform idea pays off. DeepSmith closes that loop in one place, from spotting where you are invisible and which competitor pages win the citation, to producing the on-brand article that targets the gap, to remeasuring the same prompts afterward. That is a workflow that shortens the distance between a change and a measured result. It is not a promise that any single edit will win the citation, because no honest tool can promise that.

How to tell it is done: measurement is no longer a fire drill for when leadership asks. It is a standing rhythm that produces the same defensible proof every month.

What to do next

You do not need all nine steps live by Friday. You need the first two. Freeze a prompt set this week, start a baseline, and add the change log before you touch a single page. That alone is enough to measure AEO results you can stand behind, and to finally answer "is my AEO working" instead of guessing. Momentum matters more than perfection, and a modest, controlled measurement beats an elaborate one you never finish.

If you would rather not stitch the baseline, the prompt history, and the before-and-after comparison together by hand, that is exactly the loop DeepSmith was built to close. You can start a free trial and see real prompt data and real drafts before you pay. Either way, the discipline is the same, and you already understand it now.

Frequently asked questions

How long before I know whether my AEO work is working?

It depends on the engine and your site. On established, well-structured domains, first mentions can show up within days to a few weeks, faster on Perplexity, slower on Google surfaces. Stable citation patterns usually take longer, often eight to twelve weeks, and newer or technically weaker sites take longer still. Use the first month to observe and debug, and save strong conclusions for a window that covers at least one citation half-life, ideally two.

Is one week's reading enough to prove AEO impact?

No. Repeated prompts vary run to run, and citations decay over several weeks, so a single reading tells you almost nothing on its own. Use a fixed prompt set, repeated observations, and a rolling window that covers at least one relevant half-life before you call anything a result.

How is this different from just asking ChatGPT about my brand?

Manual checks are great for a gut-feel look, but they cannot set a baseline, keep prompts consistent, compare competitors, measure a trend, or tie movement to a dated change. Real measurement needs repeatable prompts, dates, engines, controls, and a change log. That is the difference between a screenshot and a case.

Do I still need holdouts if the before-and-after lift is obvious?

If you want proof rather than a hunch, yes. A clean before-and-after shows movement, but matched holdout prompts, competitor controls, a cross-engine check, and a second post-change window are what separate your intervention from a normal engine or market swing.