Your CEO asks where the company shows up in AI answers. You have a dozen brands, a page inventory nobody has read end to end, and teams in markets you rarely talk to. So where do you start?
Right here. This guide gives you a repeatable large site AEO audit that measures every brand and every market on its own, then rolls the results up into one view leadership can actually read. You will finish with a baseline, a named cause behind every gap, and a schedule that keeps the numbers comparable from one month to the next.
Step 1: Define your audit dimensions before you collect anything
Decide the shape of the report first. Tools come later.
Write down the dimensions every result will be filed under. At minimum: portfolio, brand, product or business unit, market, language, domain and subdomain, page, topic, funnel stage, buyer intent, competitor set, engine, collection period, and prompt version.
Then build two small reference records.
For each brand, record the canonical name, aliases, former names, product names, common misspellings, the parent-company relationship, the main domain, regional domains, social profiles, and the names third parties use for you. Without this, a system treats a product, a parent company, and a regional brand as three unrelated things. Or worse, as one.
For each market, record the country, the language, the currency, local product naming, the target domain or directory, and one hard question: is this a real local offer, or a translated version of a global page? The answer changes how you read every number that follows.
Now separate three roll-up levels. The answer level is one prompt, one engine, one collection event. The operational level groups by prompt, topic, page, market, language, competitor, and engine. The executive level groups by brand, region, portfolio, period, and business priority.
Done when: every result you collect can be assigned to exactly one brand, market, language, engine, prompt version, and period. Anything you cannot place gets labeled unknown, not quietly dropped into a default market.
Common mistake: leading with one blended enterprise score. A strong global brand hides a weak regional one. A well-performing corporate domain hides a product that engines never surface. The AI visibility multiple brands problem is a reporting problem before it is a content problem, and a single average is how you lose it.
Step 2: Build one canonical inventory of pages, markets, and languages
Start with the pages, not the dashboard. You cannot audit what you have not listed.
Pull from every source you have. XML sitemaps and sitemap indexes. CMS exports. Regional domain and subdirectory lists. Product and documentation inventories. Canonical URL data, hreflang relationships, and Search Console page data. Then the important pages that never made it into a sitemap at all, which is usually where the surprises live.
Then classify each URL by brand, market, language, page type, topic, funnel stage, product, and canonical status. Include the pages people forget: documentation, pricing, product, newsroom, investor, research, support, and policy. Answer engines cite those constantly.
For a very large estate, split the crawl by sitemap and by business unit. Google's own sitemap limits are a handy way to cut it up: a single sitemap file is capped at 50,000 URLs or 50 MB uncompressed, a sitemap index tops out at 50,000 entries, and Search Console accepts up to 500 sitemap index files per site. Use fully qualified URLs and list the preferred URL rather than every duplicate.
One caution. Sitemap inclusion is inventory evidence. It is not proof that a page will be crawled, indexed, or cited. Submitting a sitemap is a hint to search engines, nothing more.
Done when: you can produce one table with one row per canonical URL and columns for brand, market, language, product, topic, funnel stage, page type, sitemap, canonical, hreflang group, indexability, last significant update, and where the row came from.
Common mistake: treating language as market. One English page may serve six countries. Two pages in the same language may carry different pricing, different products, different legal terms, and different customer proof. Multi-market AI search falls apart the moment those two columns collapse into one.
Step 3: Build a prompt portfolio for each brand and market
A prompt list is not a keyword list, and a random pile of questions will not survive its first review. Build families instead. If that sounds like a lot of writing, it is less than you think, because families repeat across markets.
Cover category and definition questions, problem-aware questions, use cases, product and solution questions, comparisons, alternatives, best-of and shortlist questions, integrations, pricing, and the risk, security, compliance, and procurement questions that come up in real deals. Then add the regional and language-specific ones, plus competitor, brand-recognition, product, and customer-segment prompts.
A useful way to organize them is three diagnostic layers. Does the engine know what the brand is? Does the brand appear for relevant category and problem questions? And does the engine actually recommend or shortlist it? One published starter framework uses fifteen prompts, five per layer, run three to five times each when you want to measure variance. Treat that as a practical starting method, not an industry standard.
Keep a separate portfolio for every materially different product, buyer segment, geography, and language. The framework stays the same everywhere. The inputs change.
Each prompt record should carry the prompt text, an ID and version, language, market, brand or product scope, topic, funnel stage, intent, the competitors you expect to see, the date added, an owner, the reason it exists, and whether it is branded or unbranded.
Lean on the unbranded prompts. Those are your competitive measure. Branded prompts mostly prove the engine has heard of you, and research on AI visibility consistently finds branded recognition running far ahead of unbranded recognition. The gap between those two numbers is the real story.
Sourcing them is less painful than it sounds. Mine sales calls, support tickets, site search, product docs, community threads, and your regional teams. DeepSmith's Discover Prompts generates candidate questions from your stored product, persona, and buyer-stage context, which gives each market team a starting set to edit rather than a blank page.
Done when: every market and business priority has a versioned prompt set, and a reviewer can explain why each prompt exists and what decision it informs.
Common mistake: adding new prompts every week, then reporting that visibility improved. New prompts change the denominator. Re-measure the same stable set before you expand it.
Step 4: Lock the collection protocol and run it on a schedule
This is the step that turns an audit into a system. Run the same portfolio across the engines that matter to you, on a schedule, and store the whole answer rather than a yes or no.
Capture all of it:
- The engine and its version, where you can get it
- The prompt text and version, and the timestamp
- The market and language context
- The full answer, and every brand named in it
- Position or recommendation status, sentiment, and message accuracy
- Every cited URL, tagged as owned, competitor, or third party
- The cited page title and canonical URL
- Rendered evidence where you can capture it
- Collection errors and missing answers
For cadence, daily collection suits active categories, launches, campaigns, and reputation events. Weekly is a reasonable rhythm for stable categories. A new program can use a fourteen to thirty day baseline to learn its normal range. These are operating recommendations, not thresholds handed down from anywhere.
Engines do not behave alike, so do not average them into one number too early. ChatGPT Search may rewrite a question into one or more targeted searches and lean on search partners. Perplexity describes its answers as web-grounded with citations built in. Google says AI Overviews and AI Mode may fan a query out across subtopics, and that the two surfaces can use different models and techniques, so their answers and their links can differ. A win in one Google surface is not automatically a win in the other.
This is where enterprise AEO automation stops being a nice idea. Nobody is running hundreds of prompts across several engines and markets by hand every week. DeepSmith runs tracked prompts on a schedule, queries the engines configured for your plan, captures the full answer, and analyzes mentions and citations as separate measurements, with the answer history kept so you can go back and read what an engine actually said. Engine coverage rises by tier: Pro tracks ChatGPT, Grow adds Perplexity, Scale adds Gemini, and Enterprise and Custom plans cover all ten engines with custom multi-region tracking.
Pro tip: keep collection and interpretation apart. Preserve what the engine returned first, then score it against a written rubric. Score as you read and your expectations quietly become the data.
Done when: any answer can be replayed with its prompt, engine, date, market, and sources attached, and failed runs, empty answers, and API errors are labeled as errors rather than counted as invisibility.
Step 5: Score the answers, then roll them up in stages
Now you turn answers into numbers. Keep the fields simple, and keep them separate. Every scoring problem you will hit starts with two things being measured in one column.
For each answer record: mentioned yes or no, cited yes or no, owned citation yes or no, competitor cited yes or no, recommended yes, no, or unclear, position where an ordered list actually exists, sentiment, message accuracy, source quality, and market fit.
From those you get the reporting metrics: mention rate, citation rate, owned citation rate, citation share of voice, share of voice by brand, first-position share where it applies, recommendation rate, page citation share, competitor displacement rate, message accuracy rate, market-specific visibility, cross-engine variance, and the trend against the last period.
Two definitions are worth writing down and never changing. Mention rate is answers mentioning the brand divided by total relevant answers. Citation rate is answers containing at least one owned citation divided by total relevant answers. Whatever denominator you choose, use the same one everywhere.
Share of voice needs its own guardrail. It only means something when the competitor set, prompts, engines, markets, and collection period are held constant. Do not hold up one market's share of voice against another's unless you have normalized the prompt mix. A market with fewer competitors produces a higher share without a single extra ounce of real visibility.
Then aggregate in stages rather than in one leap: answer, prompt, topic, market, brand, portfolio. Publish the unweighted view alongside any weighted one, and never show a weighted result without showing how it was weighted.
Common mistake: treating a mention as a citation. A brand can be named with no link at all, and a page can be cited in an answer that barely names the brand. Enterprise citation tracking falls over when those two collapse into one metric. A second version of the same error is counting every cited URL as an independent win, when several URLs often support a single answer.
Done when: leadership can ask which brands are invisible, in which markets, for which buyer intents, on which engines, which competitors replace you, which pages earn citations, and whether the problem is visibility, citation ownership, accuracy, market fit, or technical access, and the same model answers all of it.
Step 6: Attribute every result to a page, a source, and a competitor
A brand-level score tells you something is wrong. It never tells you what. So go one level down, and this is the level where the audit starts paying for itself.
For every owned citation, record the exact page and the prompts that produced it. Then rank your pages by how many answers cite them, their share of your owned citations, how many distinct prompts they win, how many markets and engines they show up in, which funnel stages they support, competitor overlap, and whether the citation holds over time.
Now look at the pages nobody cites. A page can be perfectly indexable and still fail because it never states a clear definition, never shows evidence, or never says anything specific to that market.
Do the same on the competitor side, sorting each priority prompt into four buckets. The competitor appears and owns the citation. The competitor appears with no source cited. You appear but a competitor owns the supporting citation. Or nobody appears, which is an open prompt sitting there unclaimed.
Then name a cause for each gap. Missing content, weak category framing, thin topical depth, thin comparison coverage, a third-party authority gap, a crawl or rendering problem, a market or language mismatch, outdated information, ambiguous messaging, or a sentiment issue. Resist the reflex to publish a new page for every gap. Enterprise audits keep turning up years of outdated PDFs, third-party articles carrying wrong facts, key numbers trapped inside images, and content published through JavaScript-heavy systems that crawlers struggle with.
DeepSmith does this attribution work for you. The Pages view shows which of your pages AI actually cites, each page's share of your total citations, and the prompts driving them. Competitor Citations shows who wins the same prompts, on which exact pages, and how each competitor performs by platform. Running AI visibility multiple brands wide means repeating that comparison per brand and per market, never once for the group.

Done when: for each priority gap you can name the prompt, the competitor page, your own page, the source type, the market, and your best current theory of the cause.
Step 7: Validate technical access and market signals
Let the visibility results point the technical audit rather than the other way round. You now know which clusters are invisible, so go look at those.
Work through the usual suspects:
- Robots directives, noindex and nosnippet controls
- Canonical tags and sitemap inclusion
- HTTP status and redirects
- JavaScript rendering, and whether the content is present in the rendered HTML
- Internal links and page language
- Hreflang reciprocity, plus regional pricing and product details
- Structured data where it fits
- PDFs and other non-HTML assets
- Third-party pages carrying stale claims
Google's guidance is reassuring here. There are no special technical requirements and no special schema for AI Overviews or AI Mode beyond normal Search eligibility and ordinary SEO practice. A page needs to be indexed and eligible to appear with a snippet. Eligibility still does not guarantee crawling, serving, or citation, so read it as a floor, not a promise.
The market signals need their own pass. Hreflang annotations must be reciprocal, every language version should list itself and the others, alternate URLs should be fully qualified, a country code never appears without a language code, and x-default belongs only on a genuine fallback. Watch for locale variants being auto-redirected by IP or browser settings, and for near-identical regional pages being treated as duplicates. Google supports the HTML, HTTP header, and sitemap methods equally, so pick the one your team can actually keep accurate at this size.
Common mistake: reading a crawl report as a citation report. A crawler can prove a page was fetched. It cannot prove an answer engine chose that page as evidence for a buyer's question.
Done when: every invisibility cluster has a disposition: technically inaccessible, indexable but not selected, cited through a third party, missing content, wrong market signal, outdated source environment, or still unresolved.
Step 8: Report in three layers and keep the audit running
One dashboard cannot serve a CMO and an analyst. Build three views instead.
The executive portfolio view carries overall mention and citation rate, owned citation share, share of voice against competitors, the trend versus last period, the brands and markets with the widest gaps, priority prompts where competitors appear and you do not, the engines with the most variance, the pages and outside sources driving the most visibility, and any open data-quality risk.
The regional and brand operating view is the same data with filters: brand, country, language, product, engine, prompt family, topic, funnel stage, competitor, owned page, source type, and period.
The investigation view keeps the receipts. Full answers, cited URLs, score rationale, errors, and change history.
Add your Google layer on top. Search Console's generative AI performance reports, announced in mid-2026 and rolled out globally that August, show impressions in AI Overviews, AI Mode, and generative AI features in Discover, broken out by URL, country, device, and date. It is a genuinely useful surface. It is not a replacement for prompt-level collection, because it gives you no answer text, no cross-engine competitor comparison, and no prompt-level source context.
Then set the rhythm. Re-run the same prompt set before adding new ones. Keep a change log recording the date, the page, the market, the message, the product, the competitor, and the result you expected. When a number moves after you publish something, that is evidence of a change, not proof of cause, and saying so out loud protects your credibility.
DeepSmith supports this ongoing pass with engine and time views, seven, thirty, and ninety day trend ranges, answer history, page-level citation attribution, competitor citation views, source tagging, and a Search Console connection. Its Content Map classifies your pages and your competitors' pages onto one shared topic and funnel-stage taxonomy and re-checks sitemaps every twenty-four hours, which is how you tell whether a visibility gap sits on top of a coverage gap. For a large site AEO audit, that combination is the difference between a report you rebuild every quarter and one that keeps itself current.
Done when: the executive view leads to an assigned action, and the investigation view holds enough evidence to explain any number without redoing the research by hand.

What to do next
You do not have to do all eight steps this quarter. Pick your two most important brands and your two most important markets, run the whole loop on those, and let the rest of the portfolio follow the pattern you just proved.
The first baseline is the hardest one. After that, enterprise AEO automation does the collecting and your team spends its time on the gaps.
If you want to see what the evidence looks like for your own brands before you commit to anything, start a free DeepSmith trial and track a small prompt set across a couple of markets. Real answers from your own category beat any framework on a slide.



