If you have been watching how ChatGPT, Google AI Mode, or Perplexity answer questions in your category, you have probably noticed the same handful of sites showing up again and again. That observation is correct, and the research backs it up: AI search already gives a small, concentrated group of sources a large share of citation opportunities. What is not confirmed yet is the stronger claim that keeps circulating, that AI search is citing fewer unique websites than it used to, across the board. The evidence grade here is mixed. The concentration is real and well documented. A market-wide decline in the total pool of cited websites is not.
That distinction matters for how you plan your own AI citation concentration strategy. If the pool of cited websites really were shrinking everywhere, the honest answer for a small SaaS brand would be close to "give up." If the real story is concentration, not universal shrinkage, then the honest answer is closer to "the game is more competitive than it looks, and it is not the same game on every engine." This piece works through what the research actually shows, so you can plan around the real trend instead of the version of it that gets repeated on social media.
What "fewer cited sources" actually means
Part of the confusion comes from people using several different measurements as if they were one thing. It helps to separate them before you look at any numbers.
A citation, in this research, is a source link or attributed source shown in an AI answer. Depending on the study, that can be a URL, a page, or a domain. Citation share is how often a domain shows up among an engine's citations. It is not the same as clicks, traffic, or sales. A source can have a high citation share and still send you almost no visitors, especially when the citation points to a card the engine hosts itself rather than to an outside page.
AI citation concentration describes how unevenly citations are spread across sources. A concentrated system gives most of its citations to a small number of domains. Researchers measure this a few ways: top-k citation share (what percentage of citations the top 10 or top 20 domains receive), the Gini coefficient (a 0 to 1 score where 1 means extreme inequality), the raw count of unique cited domains in a sample, and cross-engine overlap (how much the top sources on one engine match the top sources on another).
AI search source diversity has two parts that people also mix up. Breadth is how many different domains show up at all. Distribution is how evenly citations are spread across those domains. A system can technically cite a lot of domains and still be highly concentrated, if most of the citations pile onto a handful of them.
Once you hold these apart, three separate claims come into view: concentration (a small group of domains gets a large share of citations), redistribution (one source loses ground while another gains it), and shrinking diversity (the total number of distinct sites cited actually falls over time). The first two are well supported by the research. The third is the one that is not confirmed.
The evidence that AI citations are concentrated
The strongest single data point here comes from a July 2025 arXiv paper, News Source Citing Patterns in AI Search Systems, which analyzed the AI Search Arena dataset: 24,069 conversations, 366,087 citation URLs, 12 AI search models, and systems from OpenAI, Perplexity, and Google, touching 83,533 unique domains in total. The paper focused specifically on news citations, which made up about 9% of all citations in the dataset, so its numbers describe news sourcing rather than commercial or product pages.
Using the Gini coefficient, the paper found OpenAI's models at 0.83, Perplexity's at 0.77, and Google's at 0.69. In plain terms, OpenAI's top 20 news sources captured 67.3% of its news citations, compared with 31.9% for Google and 28.5% for Perplexity. That is a winner-take-most pattern, not an even spread, and it holds within a provider's models more than it holds across providers.
Two more studies confirm the same pattern outside of news. Digital Applied's April 2026 analysis of 1,000 Google AI Overviews found that roughly 1,200 domains appeared at least once, but the top 1%, about 12 domains, accounted for 47% of all citations, and the next 9% accounted for another 31%. Everyone else split the remaining 22%. The average overview cited 4.2 sources, ranging from 2 to 9, and only 8% of overviews cited more than 7 domains. BrightEdge's April 2026 study, covering ChatGPT, Perplexity, Gemini, Google AI Mode, and Google AI Overviews, found the top 10 domains capturing anywhere from 18.5% of citations (ChatGPT) to 26.7% (Perplexity), depending on the engine.

Put together, this is a real and measurable pattern: a small set of sources gets a disproportionate share of citation opportunities. What none of these studies do is compare an earlier period against a later one and show the unique-domain count falling. They are snapshots of how concentrated the current picture is, not a before-and-after showing that concentration is getting worse over time.
What actually changed over time
The clearest time-based evidence comes from Semrush's November 2025 study, which tracked more than 230,000 prompts and over 100 million AI citations weekly from July 14 through October 12, 2025, across ChatGPT, Google AI Mode, and Perplexity. On ChatGPT specifically, Reddit went from appearing in close to 60% of prompt responses in early August to around 10% by mid-September. Wikipedia fell from roughly 55% to under 20% over the same stretch. Before September, Reddit and Wikipedia each showed up about five times as often as the next most-cited domains on ChatGPT; after mid-September, they sat closer to sources like Medium, Forbes, and LinkedIn.
That is a large, real shift, but it was not uniform. Wikipedia stayed near 3% on Google AI Mode and 0.8% on Perplexity the whole time. Google AI Mode's mix stayed more balanced and stable across the study period than ChatGPT's did. Semrush pointed to Google's removal of the num=100 search parameter around September 11, 2025, as a possible factor, but the study did not establish that as the sole cause, and it said plainly that the changes were not uniform across engines.
A separate shift shows up on the Google side. Profound and Search Engine Journal reported that between April 15 and June 30, 2026, Google's own domain became the second most-cited domain within Google AI Mode, with its citation share rising 8.4 times, mostly through Google-hosted Business Profile cards and Product Knowledge Panels, especially on local and product queries. That is a meaningful redistribution toward Google's own surfaces. It is a citation-share event, not a confirmed drop in the number of outside sites Google AI Mode cites, and it says nothing about whether users still click through to the businesses being described.
Read together, these are documented, platform-specific redistributions where one source's share fell and another's rose. Neither one is evidence that the total pool of cited websites, across the whole AI search market, is shrinking.
Why source diversity differs by engine
The idea of "the AI source pool" as one shared thing does not hold up once you look at more than one engine. BrightEdge's cross-engine work found that the overlap among the top 100 cited-source lists ranged from just 16% to 59% depending on which two engines you compared, and overlap among top named-brand lists ranged from 36% to 55%. A separate 2026 synthesis from 5W estimated that only about 11% of domains were cited by both ChatGPT and Perplexity.
Engines also differ in what kind of source they lean on. In BrightEdge's data, Gemini's citations were about 26% authority sources (like established reference or news sites) with only 0.2% coming from user-generated content, while Google AI Overviews ran closer to 10% authority sources and 18% user-generated content. That is close to opposite profiles between two Google-family products, which is a reminder that "Google" is not one engine with one citation behavior.
A June 2026 Search Engine Journal report on BrightEdge research adds a layer to this: the same source can carry a different job depending on the engine and the question. Reddit showed up alongside authority sources like Mayo Clinic and Healthline in roughly 36% of ChatGPT citations, but only about 6% of the time in Google AI Overviews. LinkedIn showed up in 33% of ChatGPT's "how-to" citations against 22% for Google AI Overviews. Reddit appeared roughly twice as often on ChatGPT as on Google AI Overviews for how-to questions specifically. The pattern BrightEdge describes is that engines seem to assign different sources to different roles: explanation, comparison, verification, professional judgment, or personal experience. That is a strategic read of the data, not confirmed internal logic from the engines themselves, since none of them publish how their retrieval actually weighs a source.
This is what AI search source diversity actually looks like in practice: not one shrinking or growing pool, but several separate pools that overlap only partly. The practical takeaway is that a source strategy built for one engine will not automatically transfer to another. A page that earns citations on ChatGPT is not guaranteed a citation on Google AI Mode, even for a similar question, because the two draw from meaningfully different pools and reward different kinds of sources.
What concentration means for your odds
Here is where it is worth being honest about what the data cannot tell you. None of these studies support a statement like "a brand has a 3% chance of being cited" or "publishing ten articles gets you a 20% chance of a slot." The 4.2 average citations per Google AI Overview is an average citation count, not a probability that any particular brand lands one of those slots. The 47% share going to the top 1% of domains in the Digital Applied sample describes that sample's distribution, not a guarantee about any other query or vertical.
What the research does support is a winner-take-most picture rather than a winner-take-all one. A small set of sources gets a large share of the citation opportunities, but which sources those are changes by engine and by the type of question being asked. That leaves real room for a smaller brand, just not an even playing field.
The more useful way to think about your odds is as a measured competitive position rather than a market-wide probability. For a given prompt set that matters to your business, ask what percentage of tracked answers mention your brand at all, what percentage cite your own pages specifically, which competitors are picking up the rest of the citations, whether you show up on more than one engine or only one, and whether your citation share is moving up or down over time within that prompt set, even while the broader market stays concentrated. That framing turns "AI search feels impossible to crack" into a specific, trackable question: are we gaining ground on the prompts that matter to us, on the engines our buyers actually use.
This is the kind of measurement DeepSmith's AEO tracking is built around: mention rate, citation rate, and share of voice by engine, with a breakdown of which competitors are winning citations on your tracked prompts and which of your own pages are earning them. The point of tracking it this way is not to chase one generic visibility score, but to see your actual position within the specific, concentrated pool that applies to your prompts, so you know where the real openings are instead of guessing.
What would confirm the stronger claim
To actually prove that AI search is citing fewer unique websites over time, market-wide, researchers would need something none of the current studies provide: the same fixed set of prompts, the same definition of a citation, tracked on the same engine version, repeated across multiple periods, reporting the total number of unique domains cited, the average number of unique sources per answer, and a concentration measure like the Gini coefficient calculated the same way each time. Ideally that would be replicated across more than one engine, so a shift on ChatGPT could be checked against what happened on Google AI Mode and Perplexity in the same window.
Right now, the studies covered here differ in engine, country, prompt selection, industry, and whether they measure one snapshot or repeated ones. Some count URLs, others count domains. Some cover a single week; the longest, Semrush's, covers thirteen. None of them can be added together into one continuous timeline, and treating them as if they could is exactly how "concentration is real" turns into the looser, unproven claim that the whole pool is shrinking. Google's own documentation describes AI Overviews and AI Mode as search features that surface relevant links, without publishing a market-wide source-concentration benchmark of its own.
What to do differently
Start by building your own citation-share baseline instead of relying on someone else's market-wide number. Pick a fixed set of prompts that matter commercially to your business, then track, by engine, whether your brand is mentioned, whether your own page is the one cited, which competing domains show up instead, how many citations the answer includes, and how your share of that citation set moves over a few weeks. Keep the engine, the exact prompt wording, and the observation window consistent, because a ChatGPT number and a Google AI Mode number are not interchangeable even when they look similar.
Treat the source pool you are competing in as a portfolio, not a single ladder to climb. Since research points to different sources filling different roles, educational explanation, comparison, verification, professional judgment, personal experience, look at which of those roles your tracked prompts actually reward, and whether your existing pages are built to fill that role or a different one.
Use concentration as a prioritization signal rather than a reason to publish everywhere at once. If a prompt's citations are dominated by the same few domains every time you check, that is a more competitive opening than a prompt where the citation set changes week to week. The sequence that follows from the research is to identify the prompts where citation presence actually matters to revenue, measure who currently holds those slots, figure out what role those incumbents are filling, look for a genuine coverage gap relevant to that role, and then watch whether your citation share moves after you publish. None of the studies here say a brand can force an engine to cite it. They do say that the brands with a real, current picture of their citation position are working from something better than a guess.



