DeepSmith

Jul 26 · AEO & AI Visibility

17 min read

First-Party vs Third-Party Citations: Why AI Search Rewards Both and How to Balance Them

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
Monochrome abstract diagram of stacked owned pages on the left and a scattered network of independent sources on the right, both feeding a central answer card with a link icon, under the cover line Owned vs Earned Citations.

Most AI visibility reports surface the same uncomfortable arithmetic. The pages a brand owns, the only URLs its team can edit, restructure, and mark up with schema, account for a small minority of the citations answer engines surface. A synthesis of six independent 2025 studies by AuthorityTech places third-party sources at 82 to 95 percent of all AI citations, depending on engine and vertical. The conclusion many teams draw from that number, that owned content has stopped mattering, does not follow. The question of first-party vs third-party citations is not a budget contest between two channels. It describes two surfaces inside one retrieval system, each performing a job the other cannot.

What follows separates the two, reports what the evidence says about the mix in AI answers, and gives a four-question rule for placing the next unit of effort, so that a team can balance AI visibility channels on evidence rather than instinct. The scope is allocation logic rather than channel-level execution, and every percentage here is directional.

First-Party vs Third-Party Citations at a Glance

DimensionFirst-party (owned, on-domain)Third-party (earned, off-domain)
Who controls the contentThe brand, entirelyA publisher, platform, or community
Typical surfacesProduct, pricing, docs, help center, blog, glossaryReview platforms, news and trade press, Wikipedia, Reddit, YouTube, comparison publishers
Share of the citation mixA consistent minority in every study reviewedThe dominant majority, roughly 82 to 95 percent
Primary job in retrievalCanonical answer and entity groundingIndependent corroboration and category presence
Structured data controlFull: schema, headings, freshness signalsNone
Speed of changeImmediateSlow, measured in months, dependent on third parties
Strongest query typesBranded, factual, documentation, complianceCategory-level, comparison, best-of
Characteristic failure modePresent but unextractable, so never retrievedAbsent from the surfaces engines reach for

The table describes a division of labor, not a ranking. The sections below explain why substitution fails in both directions.

What Counts as a First-Party Citation

A first-party citation is a reference inside an AI answer whose URL points to a domain the brand controls. Product, pricing, and feature pages, documentation, help-center articles, the company blog, glossary entries, and canonical landing pages all qualify. The defining property is control: the brand can revise the language, restructure the page, and attach structured data such as Organization, Product, SoftwareApplication, and FAQPage markup at any point.

That control carries a specific consequence for retrieval. First-party pages are where a brand states its own claims in extractable form, and where it defines the entity an engine must resolve before it can attribute anything at all. Consistent naming, Organization markup, and sameAs links are the mechanism by which a retrieval pipeline concludes that scattered references across the web point to one company rather than to several similarly named ones. Where that grounding is missing, third-party citation share tends to suffer too, because the corroboration has no stable referent. The owned side of the split is a precondition for the earned side registering at all.

What Counts as a Third-Party Citation

A third-party citation points to a domain the brand does not control: review platforms such as G2, Capterra, TrustRadius, and GetApp; news outlets and trade press; analyst reports; Wikipedia and Wikidata; Reddit, Quora, and niche forums; YouTube transcripts; and comparison or best-of publishers. A brand influences these surfaces through coverage, review acquisition, and community participation, but cannot edit them.

Across the studies reviewed, off-domain AI citation sources cluster into four recurring buckets:

  1. User-generated content, forums, and question-and-answer sites. Reddit, Quora, Stack Exchange, and video transcripts. Reddit is consistently a top-three cited domain on ChatGPT, Perplexity, and Google AI Overviews.
  2. Encyclopedic reference. Wikipedia, Wikidata, and the knowledge panels derived from them.
  3. News and trade press. Mainstream outlets alongside industry publications.
  4. Review and comparison sites. G2, Capterra, GetApp, TrustRadius, and their vertical equivalents.

The narrowness of that list is the important part. The corpus engines reach for is far smaller than the open web, so off-domain visibility is less a matter of broad publicity than of presence on a specific and fairly stable set of surfaces.

Mentions and Citations Are Not the Same Signal

The two terms are used interchangeably in most reporting, and the conflation distorts allocation. A mention occurs when a model names the brand in its answer, with or without a link. A citation occurs only when the model links to a specific URL as the source for a claim. A brand can be recommended without a single link pointing at it, and a page can be cited for a fact without the brand being named as a recommendation.

The relationship between the two is directional rather than incidental. In a correlation study of roughly 75,000 brands, Ahrefs found brand mentions correlating with AI citation share at approximately 0.664, against 0.218 for backlinks, making mentions roughly three times more predictive of citation share than links in that dataset. Correlation of that kind does not establish causation, but it supports a defensible working model: third-party mentions behave as the seed layer, first-party pages as the harvest layer. Brands named frequently across independent surfaces tend to accumulate linked citations afterward, provided a well-structured owned page exists for the engine to point at.

The Evidence: Earned Sources Carry the Majority

Three independent measurements converge.

  • An analysis of more than a million AI citations, published in a 2026 state-of-AI-search report, found roughly 85 percent of brand mentions in AI answers originating on third-party pages rather than the brand's own domain.
  • The Observer and Muck Rack Generative Pulse report, published in December 2025, found 89 percent of links cited by AI assistants coming from earned-media outlets.
  • The AuthorityTech synthesis of six independent 2025 studies placed the earned share at 82 to 95 percent across engines and verticals.

The spread between those figures is best read as methodology variance rather than real-world movement, since study designs differ in prompt set, language, vertical, and engine coverage. What survives across all of them is the direction and the order of magnitude: owned vs earned AI citations is a lopsided ratio, lopsided the same way on every engine measured. The practical implication is not that owned pages are unimportant. It is that a team measuring its AEO program purely by citations to its own domain is watching a minority slice of its own visibility and will misjudge both its position and its competitors'.

Why Retrieval Rewards Both Surfaces

The skew follows from how retrieval-augmented answers are assembled. A query is parsed into an intent, candidates are pulled from web indexes, those candidates are reranked on relevance, freshness, authority, and entity clarity, and only the survivors reach the model that drafts the answer and attaches attribution. No major provider publishes the weighting it applies, and source weighting is a subject in its own right. Three consequences matter for allocation.

Corroboration is a trust proxy. When multiple independent sources describe a brand consistently, the pipeline treats that agreement as evidence the brand is the canonical referent for a topic. A single domain asserting its own excellence rarely clears that threshold.

Entity clarity survives reranking. Pages that disambiguate the entity plainly, through structured markup, consistent naming, and explicit about-content, tend to survive the filter. Pages that bury the entity in marketing copy do not.

Extractability beats volume. Models lift short, declarative passages. An answer sitting behind a hero image, a modal, or four paragraphs of narrative frequently goes unretrieved, which produces the familiar pattern of a page that ranks well and is never cited.

Together these explain why the two sides of off-domain vs on-domain AEO are complements. Corroboration is earned off-domain. Extractability and entity grounding are built on-domain. A program strong in only one stalls for a predictable reason.

How AI Citation Sources Differ by Engine

Source preferences vary enough between engines that an aggregate number conceals more than it reveals.

EngineCharacteristic citation behavior
ChatGPTEncyclopedic and community-weighted. Wikipedia is the single most-cited domain at roughly 7.8 percent of citations, with Reddit and YouTube also prominent. Favors evergreen reference content over breaking news.
PerplexityThe most distributed mix of the major engines. Reddit leads at roughly 6.6 percent, followed by Wikipedia and a long tail of news and trade press. Cites more sources per answer, so per-domain share is lower and domain diversity higher.
Google AI Overviews and AI ModeOver-indexes on Reddit, YouTube, Quora, and Wikipedia relative to the classic organic result set, and pulls disproportionately from Google-owned properties. Roughly 88 percent of citations come from URLs outside the traditional top ten organic results.
GeminiTracks closely with AI Overviews on source mix, with a higher share of publisher and reference-site citations and a lighter forum footprint on many queries.
ClaudeThe lowest citation volume of the major models. When it does cite, the profile skews toward editorial, reference, and documentation sources, with a lighter forum presence.

Two qualifications belong with that table. The mix is more stable than it appears week to week: across a thirteen-week window, Semrush found approximately 96.8 percent of cited domains and 97.2 percent of mentioned brands showing zero week-over-week change in citation share, meaning individual URLs rotate inside a largely fixed set of surfaces. The set is not permanent, though. ChatGPT's reliance on Reddit and Wikipedia dropped sharply in mid-September 2025 and has only partially recovered. Optimization should therefore target durable presence on the right surfaces, not any single answer.

When First-Party Citations Win

First-party pages are a minority of the aggregate mix and the dominant source in four situations.

Branded factual queries. When the prompt concerns the brand itself, covering pricing, features, integrations, security posture, or terms, the brand's own pages should be the cited source. Where a third-party page is cited instead for a fact the brand publishes, the usual causes are entity ambiguity or an unextractable page rather than a lack of authority.

Regulated and high-stakes categories. Health, legal, financial, and similarly regulated verticals push engines toward brand-controlled surfaces for specifications, regulatory status, and policy language. The owned page carries the canonical claim by design.

Long-tail comparison and feature queries. Prompts of the form "how does this product handle a specific workflow" resolve most often to the brand's own comparison and feature pages, provided those pages are structured as extractable answer modules. Third-party reviews corroborate; the brand page answers.

Entity grounding. The least visible and most consequential case. Consistent naming, Organization markup, and sameAs links allow an engine to resolve the brand to one referent. Without that resolution, earned coverage accumulates against a blurred entity and converts poorly.

When Third-Party Citations Win

The mirror cases are equally specific.

Category-level and best-of queries. When the buyer asks which tools solve a problem, engines pull from reviews, listicles, forums, and editorial coverage. A brand absent from those surfaces is absent from the answer.

Low brand recognition. Engines do not cite entities they do not recognize. For a brand with thin presence in AI answers, earned coverage addresses the corroboration deficit, which is the faster route to recognition than additional owned pages.

Competitive share of voice. Where competitors hold most of the mentions on the review platforms and comparison publishers an engine reaches for, owned depth does not close the gap. The gap sits off-domain and has to be closed off-domain.

Off-Domain vs On-Domain AEO: A Four-Question Allocation Rule

A fixed ratio is the wrong output here, because the correct split moves with query intent, category, and current standing. Four questions produce a defensible allocation for a specific brand.

1. Is the tracked prompt set branded or category-level? Prompt sets weighted toward branded and documentation questions justify first-party investment; sets weighted toward category and comparison questions justify third-party investment, because that is where the engines are already looking.

2. How mature is the brand's presence in AI answers today? At low presence, earned coverage comes first, since recognition precedes citation. At high presence, first-party depth comes first, as the task shifts to being the source of record.

3. How regulated or high-stakes is the category? In regulated verticals, owned pages carry the canonical claims and third-party coverage corroborates them. Elsewhere, third-party volume is the larger lever and owned pages need the fundamentals rather than exhaustive depth.

4. What is the brand's share of voice on third-party surfaces relative to competitors? Where it is low, earned placement, review acquisition, and inclusion in comparison publishers are the priority. Where it is high, first-party structure defends the lead.

Absent a strong signal from those four questions, the default the evidence supports is roughly 60/40 to 70/30 of incremental effort toward earned surfaces, tilting back toward owned in regulated categories and where prompt sets skew branded. That default is a starting position, not a rule, and should be revised the moment measurement contradicts it.

What Each Side of the Split Requires

The two sides demand different skills, which makes the allocation decision a staffing decision as well. Owned-side work is structural: extractable answer blocks on priority pages, Organization, Product, SoftwareApplication, and FAQPage schema, consistent naming and sameAs links, structured comparison pages, and descriptive internal anchor text. Earned-side work is relational and slower: trade press and analyst placement, sustained review velocity, authentic community participation, encyclopedic inclusion where notability standards are genuinely met, and appearances whose transcripts become citable artifacts. Coordinated inauthentic promotion is an increasing detection target whose reputational exposure outweighs any short-term citation gain, and editing an encyclopedic entry about one's own organization without disclosure violates conflict-of-interest rules.

Measuring Whether the Balance Is Working

An allocation decision without measurement degrades into preference within a quarter. Four measurements make the split visible.

  • Mention rate and citation rate, tracked separately, per engine. Aggregating them hides the seed-to-harvest relationship that makes the earned side worth funding.
  • Share of voice against named competitors. The metric that establishes whether an earned-side deficit exists at all.
  • Page-level citation attribution. Which owned pages earn citations, and for which prompts, separates a page that ranks from a page that gets lifted.
  • Competitor citation sources. Identifying the third-party pages winning citations for a brand's own prompt set converts a vague publicity goal into a finite placement list.

DeepSmith is built around that measurement problem before it is built around production. It tracks mention rate, citation rate, and share of voice across ChatGPT, Gemini, Perplexity, Claude, and Google AI Mode; its Pages view attributes citations to specific owned pages and the prompts driving them; and its competitor view shows which competitor pages win citations for the same prompts, which is the earned-side gap list in concrete form. The same stored brand context feeds production, so owned pages built to close a gap carry the structure, schema, and internal linking that make them extractable.

Engine coverage rises by plan: Pro tracks ChatGPT, Grow adds Perplexity, Scale adds Gemini, and Enterprise covers all five. Single-engine coverage at the entry tier is a real constraint, and it bites less than it appears here, because the surfaces dominating the citation mix recur across engines, so a read of which third-party domains win on one engine identifies most of the surfaces that matter on the others. Two limits deserve plain statement: no platform, this one included, controls what an engine cites, and tracking establishes where a brand stands on owned vs earned AI citations rather than delivering placement, which remains editorial and relational work.

What the Evidence Does Not Settle

Several caveats bound everything above.

Engine opacity. No major engine publishes a specification for how it selects citations. Every percentage in circulation is inferred from observed behavior and partially disclosed architecture, not from policy.

Prompt-shape effects. The mix a brand observes is partly a function of the questions it tracks. A prompt set heavy on pricing and documentation will show a healthier first-party share than one heavy on category comparisons, without either brand doing anything differently, and methodology variance across studies works the same way.

Volatility at the URL level. The domain mix is stable while individual URLs rotate, so a lost citation on one page is weak evidence of a problem while a repeated pattern across a surface is strong evidence. The mid-September 2025 movement in ChatGPT's source mix also showed that a stable-looking distribution can shift in weeks, so percentages should be re-measured rather than inherited.

No guarantees. No tactic described here guarantees inclusion, citation, traffic, or revenue. The guidance is probabilistic and should be funded on that basis.

Which Balance Fits Which Situation

The allocation that fits depends on where a brand currently sits, not on preference.

Early-stage brands with low recognition in AI answers should weight incremental effort toward earned surfaces while holding owned work to the fundamentals: entity markup, a clear about page, and extractable answers on the highest-intent pages. Recognition is the binding constraint, and it is not solved on-domain.

Established brands with strong category presence should weight toward owned depth. The corroboration threshold is cleared, and the upside sits in becoming the cited source of record for the questions buyers ask.

Brands with heavy branded and documentation prompt volume should read the aggregate 82 to 95 percent figure with caution. Their own mix will legitimately skew more first-party than the industry average, and chasing the average would be an error.

Agencies and multi-brand teams meet the split at portfolio level, and should classify each brand rather than apply one house ratio.

Whichever situation applies, the first step is diagnostic: measure the split before adjusting it. Teams that balance AI visibility channels well tend to be the ones that know their own ratio, per engine, before they argue about the target. Start a free 7-day DeepSmith trial to see which prompts name your brand, which pages earn the citations, and which third-party sources are winning the ones you do not.

Frequently asked questions

What is a first-party citation in AI search?

A first-party citation is a reference in an AI answer whose URL points to a domain the brand owns: a product page, documentation, or a blog post. The defining feature is control, since the brand can edit the content and attach structured data at any time, which is why first-party pages carry canonical claims and entity grounding.

What is a third-party citation in AI search?

A third-party citation points to a domain the brand does not control: review platforms, news and trade press, encyclopedic reference sites, community forums, video transcripts, and comparison publishers. A brand influences these surfaces but cannot edit them, which is what makes them function as independent corroboration.

Why do AI engines cite third-party sources more often than brand sites?

Retrieval pipelines treat agreement across independent sources as a proxy for trust, so a claim corroborated by several unaffiliated domains tends to survive reranking better than the same claim asserted on one owned domain. Third-party surfaces also carry broader topical coverage and often fresher signals.

Should a team prioritize owned content or earned mentions?

Both, in a sequence set by the binding constraint. Owned content supplies the canonical answer and the entity grounding without which an engine cannot attribute anything to the brand, while earned mentions supply the corroboration that decides whether the brand surfaces at all. Where recognition is thin, earned work moves the number faster; where it is established, owned structure captures the value.