One client asks whether they show up in ChatGPT. The next week, another asks about Perplexity. Soon you are running spot checks in five browser tabs and losing two days a month to hand-built reports. This guide gives you an agency AEO workflow that runs the same way on every account, so you can manage AEO for clients as a real service line instead of a scramble.
If that feels like a lot right now, that's normal. Almost all of the work here is standardizing once. Per client AEO tracking gets cheap after that.
Step 1: Build one prompt framework you reuse on every client
Prompts are the unit of measurement in AI search. Build the framework once, then fill it per account. The rest of the agency AEO workflow leans on this step, so it is worth an afternoon.
Use five categories for every client:
- Informational: "What is X," "how does X work." Awareness.
- Commercial investigation: "Best X tools," "X vs Y," "alternatives to X." Comparison.
- Transactional and branded: "Best X for manufacturing," "X pricing." High intent.
- Topical and educational: niche how-to questions inside the client's category.
- Product-led: feature, integration, and use-case questions.
Then run the same four moves per client. Pull seed topics from their product, personas, and buyer stages. Generate fifty to a hundred candidate prompts. Validate them against real buyer language from sales calls, support tickets, and Search Console queries. Weight them by how often people actually ask, so a question asked tens of thousands of times outranks one asked fifty times.
Lock somewhere between thirty and a hundred final prompts per client. Thirty per quarter is a sensible floor if you want numbers you can defend in a report.
You know this step is done when every client has a locked, categorized prompt set covering at least two categories, with the weight sitting where their buyers actually are.
Where agencies go wrong: tracking vanity prompts. A prompt your client loves but nobody asks will quietly inflate their share of voice and hide the gap that matters. Compare each client's prompt mix against at least one competitor's to find what you missed.
Step 2: Lock the engine set before you promise anything
Here is the thing that catches most agencies off guard. Engines do not cite the same web.
ChatGPT leans on encyclopedic, high-authority sources, with Wikipedia and Reddit carrying a large share of what it cites. Perplexity's most-cited domains skew toward Reddit, LinkedIn, and institutional sites, with dense citations and ranked snippets. Gemini and Google AI Overviews favor Google's own properties alongside Wikipedia and LinkedIn. Claude cites fewer external sources and leans long-form. Google AI Mode pulls heavily from product schema and forums.
Barely any of the domains ChatGPT cites overlap with the ones Perplexity cites. Track one engine and you are reporting on a fraction of the answer-engine real estate while telling your client it is the whole picture.
Set a floor and put it in the contract. Three engines for smaller accounts, usually ChatGPT, Perplexity, and Google AI Overviews or AI Mode. Five for enterprise work.
Then confirm your tooling actually covers what you sold. Vendors count "engines" differently, and coverage often rises with the plan tier rather than being flat.
DeepSmith tracks ChatGPT, Gemini, Perplexity, Claude, and Google AI Mode, and coverage scales the same way your roster does. Pro is $99 a month and tracks ChatGPT. Grow is $199 and adds Perplexity. Scale is $399 and adds Gemini. Enterprise covers all five with custom limits. That maps cleanly onto client tiers: your SMB accounts do not need five engines, and your enterprise accounts will not accept fewer.
You know this step is done when every client's engine set is written down, matched to their tier, and confirmed against what your tool really tracks.
Where agencies go wrong: treating one engine as a proxy for all of them. It is the single most expensive assumption in this workflow.
Step 3: Baseline every client the same way
You cannot show progress without a starting line. Run the full prompt set across the full engine set and freeze the result.
Capture the same nine things for each account:
| Metric | What it tells the client |
|---|---|
| Mention rate | How often an engine names them at all |
| Citation rate | How often an engine links to their pages |
| Share of voice | Their slice of the category against a named competitor set |
| Visibility trend | Period-over-period direction |
| Sentiment | Whether the mention helps or hurts |
| Average position | How prominent they are inside the answer |
| Source diversity | How many distinct domains the engine pulls from |
| Page-level attribution | Which of their URLs get cited |
| Prompt coverage | The share of prompts where they appear at all |
Give the numbers a frame, because a mention rate on its own means nothing to a client. A share of voice down in the low teens usually signals a real citation gap. Somewhere in the mid twenties to forties is competitive. Above that, they are leading the category.
Record it per engine, not just as a blended total. A client cited well in ChatGPT and invisible in Perplexity has a very different problem than one who is weak everywhere, and the fix is different too.
Use the same fields, the same formulas, and the same competitor logic on every account. Consistent per client AEO tracking is what makes month three comparable to month one.
You know this step is done when you have a dated snapshot per client, per engine, that you can point back to in month three.
Where agencies go wrong: reporting raw metrics with no baseline and no category context. The client cannot tell whether 18 percent is good news, so they assume it is not.
Step 4: Give every client its own isolated workspace
This is the step that decides whether your roster can grow.
One workspace per client. Never a shared setup. No shared prompt sets, no shared reports, no shared context. The moment Client A's positioning leaks into Client B's draft, you have lost the thing they hired you for.
Capture structured brand context once per account:
- About the company: positioning, differentiators, claims to make and claims to avoid.
- Products and services: a profile per product with features, value props, use cases, and their real competitor list.
- Buyer personas: goals, triggers, requirements, challenges.
- Brand voice: tone and texture settings, so output reads like them and not like you.
- Visual guidelines: palette, typography, illustration style.
- Content types and trusted sources: the formats you produce and the domains they are willing to cite.
- Their sitemap, imported and classified, so internal links and coverage signals work from day one.
DeepSmith is built on this shape. Multi-Workspace runs each brand fully isolated with its own context, content, and plan, and Deep IQ holds that client's positioning, products, personas, and voice as structured context every other module reads from. Onboarding pulls the brief, competitors, starter prompts, and first ideas from the client's own website, so you are producing on-brand work in the first week instead of the second month.
Common mistake: treating brand context as a briefing document instead of stored data. If the client's voice lives in a strategist's head or a Google Doc nobody opens, every draft costs you an editing pass, and the account stalls when that strategist takes a holiday.
Add two guardrails on top. Keep a claims-to-make and claims-to-avoid register per client, and run a weekly cross-client spot check to catch voice bleed before the client does.
You know this step is done when a new writer can produce a passable draft for any account on their first day, without a briefing call.
Step 5: Turn per-engine gaps into a prioritized backlog
Now the baseline earns its keep. Diagnose each prompt and each page, per engine.
Every tracked prompt falls into one of four buckets:
- The client is missing entirely.
- They are mentioned but not cited.
- They are cited, but buried low in the answer.
- They are winning, and you should protect it.
Then look at the pages. Which owned URLs pick up citations, and which get zero? Map who wins the citations they lose, on which exact pages, and which third-party domains dominate their prompt set. That last one often matters most. If Reddit threads and one industry publication own the answers, no amount of blog posting will fix it alone.
Turn all of that into a ranked action backlog with four kinds of work: content to write, schema to add, third-party placement to earn, and technical fixes to ship.
Check the technical side early, because it is cheap and it silently caps everything else. A robots.txt that blocks GPTBot, ClaudeBot, or their peers will suppress citations, and plenty of sites block them by accident through a CDN or firewall rule. Treat llms.txt as hygiene rather than a lever. It is new, and it is not a confirmed ranking factor.
You know this step is done when each client has a ranked backlog where every item names the prompt or page it is meant to move.
Where agencies go wrong: shipping work without recording which action targeted which metric. Six weeks later you cannot tell the client why anything changed, and "trust us" does not renew a retainer.
Step 6: Produce the content that closes the gaps
Diagnosis is the easy half. Production is where agency margin lives or dies.
Answer engines reward a specific shape of page, and it is consistent enough to standardize:
- A short, declarative lead paragraph of roughly forty to sixty words that answers the question outright.
- Definitions in the first sentence of any concept a buyer might ask about.
- Headings phrased as the question the buyer actually asks.
- Lists, tables, and comparison matrices, which engines lift disproportionately often.
- Original statistics and quotable lines, which travel with attribution.
- Author bylines and dates for freshness and expertise signals.
- Entity-rich copy that names products, categories, and competitors explicitly.
- Internal links to related pages, plus outbound citations to authoritative sources.
- An FAQ block at the end, marked up with FAQPage schema.
On schema, prioritize rather than boiling the ocean. FAQPage and Organization first. Product and Offer on commercial pages. HowTo on tutorials. Article for bylines and dates. BreadcrumbList site-wide. Review and AggregateRating where trust drives the decision.
This is also where a track-and-produce loop pays off, because the gap data and the draft come from the same place. DeepSmith's Content Studio turns a tracked gap into a finished, brand-grounded article with keyword coverage, heading structure, schema, internal linking, and metadata built into the pipeline rather than bolted on afterward. Autowrite runs it on a schedule per client, so pieces land in Produced Content without anyone opening the app, and your strategist reviews for judgment instead of mechanics. Every finished article arrives with social posts ready to copy, which turns one deliverable into a channel plan.
You know this step is done when your default article structure is the citation-ready one, and nobody has to remember it.
Where agencies go wrong: letting output sound the same across every client. Volume without voice isolation costs you the positioning your client is paying you to build.
Step 7: Template your multi-engine AEO reporting once
Where do most agencies lose their evenings? Right here. Build the report structure a single time, then fill it per account. Multi-engine AEO reporting is the deliverable your client actually sees, so it deserves more thought than the average monthly deck gets.
A monthly client report needs ten blocks:
- Executive summary: wins, declines, and what you are doing next, in one paragraph.
- Per-engine mention rate, citation rate, and trend.
- Share of voice against the competitor set, over time.
- Top-cited owned pages, with URL, citation count, and engine.
- Top-cited external sources dominating their prompts.
- Prompt-category breakdown, so they see performance by intent.
- Sentiment overview.
- The action backlog, prioritized.
- Period-over-period deltas.
- A trend narrative that tells the direction-of-travel story in plain words.
Deliver it branded. A white-labeled PDF with your logo and colors, an optional client login for the live dashboard, a scheduled monthly email, and Slack or Teams alerts when visibility drops. Alerts matter more than they sound. Being the one who tells the client about a dip is very different from being asked about it.
You know this step is done when a report takes minutes to assemble instead of half a day, and any team member can produce one.
Where agencies go wrong: building each report by hand. For teams that manage AEO for clients at any scale, it is the most reliable way to cap how many accounts one strategist can carry.
Step 8: Run a cadence that matches how fast AI search moves
AI search changes weekly. Your cadence has to respect that without eating your team.
- Daily: automated visibility checks across all tracked prompts and engines, with alerts on sharp drops and competitor overtakes.
- Weekly: an internal review of trend charts and a cross-client QA spot check. Fifteen minutes per account.
- Monthly: the client-facing report and a refreshed action backlog.
- Quarterly: a strategy memo, a prompt-set refresh, a prompt-category mix review, and a competitive benchmark reset.
The quarterly refresh is the one people skip. Buyer language moves, models update, and competitors enter. A prompt set that was right in January is quietly measuring the wrong thing by June.
When a model update knocks citations down across a client base, having per client AEO tracking on a daily schedule is what lets you tell the difference between one client's problem and an industry-wide shift. That distinction is worth a lot in a renewal conversation.
You know this step is done when the cadence lives on a calendar with owners, not in one person's memory.
Where agencies go wrong: reporting quarterly on a discipline that moves weekly. By the time the client sees the drop, it is a quarter old.
Step 9: Package it as a priced service line
An agency AI visibility service you can sell beats a bundle of favors you keep absorbing.
Package it in tiers that mirror the engine sets you locked in Step 2. Smaller local and e-commerce accounts sit at the entry tier with three engines and a smaller prompt set. Mid-market SaaS and professional services sit in the middle with more prompts and deeper competitive work. Enterprise and multi-brand clients get the full engine set, more prompts, and a quarterly strategy layer. Industry retainers run from the low thousands per month at the small end into the tens of thousands for enterprise and multi-region work, so there is real room to price by scope rather than by hours.
Three engagement models work: a straight monthly retainer, a project-based setup fee followed by a retainer, or a retainer with a performance bonus layered on top.
Name what each tier includes: engines tracked, prompts tracked, articles produced, multi-engine AEO reporting cadence, and whether the client gets live dashboard access.
You know this step is done when you can quote a new logo from a rate card instead of improvising a scope every time.
Where agencies go wrong: selling AEO as a one-off audit. The audit is the easy part, and it does not compound. The tracking, production, and reporting loop is what renews.
What to do next
Do not roll this out across the whole roster at once. Pick your most engaged client and run Steps 1 through 3 this week. You will have a locked prompt set, a defined engine set, and a dated baseline in a few days.
Then use what you learn to harden the template, and onboard the rest of the roster into the same agency AEO workflow. Momentum matters more than perfection here. You are closer to a productized offer than you think.
If you want the tracking and the production running off the same client context instead of stitching four tools together, start a free DeepSmith trial and set up one client workspace. Seven days is enough to see real prompt data and real drafts for that account before you commit.



