DeepSmith

Jul 26 · AEO & AI Visibility

17 min read

Which Third-Party Sources Google AI Overviews Pulls From (and How to Get Into Them)

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
A monochrome flat-vector cover on a charcoal background showing an answer card at the bottom center linked by thin white lines to eight source tiles arranged around it, each a simple glyph for a source type (video, forum, encyclopedia, news, reviews, official, local), under the white cover line 'Where AI Overviews Look First'.

You searched your own category, read the AI Overview at the top, and counted the sources. None of them was you. That stings, and it is far more ordinary than it feels: across 2025 and 2026 citation studies, roughly 75 to 90 percent of the citation slots in an AI Overview go somewhere other than the brand being discussed. This guide names the Google AI Overviews sources that keep reappearing, then walks you through nine steps to get into AI Overviews through them.

Here is the good news. That source pool is small and repetitive. You do not have to be everywhere. You have to be in the four or five places Google keeps returning to.

The sources Google AI Overviews actually pulls from

AI Overviews and AI Mode both draw on Google's index, and both lean hard on a narrow set of off-domain publishers. The order below reflects how often each category appears in aggregated studies from Ahrefs, BrightEdge, Authoritas, SE Ranking, Semrush, and Similarweb. Exact percentages move study to study. The rank order barely does.

Cited in a majority of eligible queries

  • YouTube. The single most cited domain in AI Overviews, measured at roughly 23 to 30 percent of citation slots depending on the study, and cited about 200 times more than any other video platform. It dominates how-to, demo, recipe, and "show me" queries.
  • Reddit. Cited in roughly 35 to 50 percent of informational, comparison, and "how do people feel about this" queries. Google pays Reddit for a content licensing deal, reported at around 60 million dollars a year, giving its systems richer access than a standard crawl.
  • Wikipedia. Heaviest on definitional and biographical queries, in the range of 18 to 40 percent depending on the query mix. It fades on commercial and YMYL topics.
  • Google's own properties. Google.com answers directly on conversions, definitions, and weather, taking roughly 16 percent of citations.

Cited in roughly 15 to 30 percent of eligible queries

  • Major news publishers. Reuters, AP, NYT, WSJ, BBC, Bloomberg, Forbes, The Guardian. In news-class queries they take 60 to 80 percent of the slots. Forbes leads a 2026 citation-score index of U.S. publications, with Reuters second.
  • Authoritative vertical publishers. Wirecutter, RTINGS, CNET and Consumer Reports in consumer tech. Investopedia, NerdWallet and Bankrate in finance. Mayo Clinic, Cleveland Clinic and Healthline in health. Tripadvisor and Lonely Planet in travel. Edmunds and KBB in auto.
  • .gov and .edu. On YMYL queries where a real authority has jurisdiction, these can take 60 to 90 percent of the slots.

Cited in roughly 5 to 15 percent of eligible queries

  • Q&A and forums beyond Reddit. Quora (one 2025 study ranked it the single most cited AI Overviews domain, against most others that put YouTube first), Stack Exchange, Glassdoor, and active industry forums. Community forums account for roughly 6 to 17 percent of AI citations.
  • Review and comparison platforms. G2, Capterra and GetApp for B2B software. Trustpilot, ConsumerAffairs and BBB for consumer. Yelp for local, where it is the most trusted source for evaluative discovery. About a third of AI Overview responses cite a review platform.
  • Industry trade press. Adweek, Digiday, Modern Healthcare, Law360. These surface on industry-internal queries.

Cited in roughly 1 to 5 percent, but growing

  • Data and statistics sites (Statista, Pew, Gallup), developer documentation, and reference bodies.
  • Press release wires. A fallback, not a channel: releases account for under 2 percent of AI citations, while earned media accounts for 82 to 95 percent.

One more thing before you plan anything. AI Mode is not the same fan as AI Overviews. An Ahrefs analysis in December 2025 found only 13.7 percent URL overlap between the two on identical queries. AI Overviews average roughly 4 to 8 citations per answer and often cluster three or four on one domain. AI Mode averages 12 to 25 per turn and spreads them wider. AI Mode third-party sources skew toward authoritative publishers and comparison sites; AI Overviews skew toward forums and video.

Two jobs? Mostly not. Optimize for AI Overviews first, because its query coverage is far larger. The authoritative coverage you build along the way is what AI Mode rewards.

Step 1: Map the prompts your buyers actually ask

Everything downstream depends on this list, so give it a real morning.

Write out 20 to 50 questions a buyer would type or say when trying to solve the problem your product solves. Use their words, not your category label. Mix the classes deliberately, because query class decides which sources Google reaches for: "is X worth it," "best X for Y," "X vs Y," "how do I X," "what is X," and "X alternatives."

Then sort by how close each question sits to a purchase. Decision-stage prompts are where a citation earns its keep.

Staring at a blank sheet? DeepSmith's Discover Prompts generates a starter set from your product, persona, and buyer-stage context, so you begin with a pool rather than a blinking cursor. You still add the questions only your sales calls would surface.

How you know it is done. You have a written prompt list, tagged by query class and buyer stage, that a colleague could run without asking what you meant.

Where people slip. They write keyword phrases instead of questions. "Project management software" is a keyword. "What is the best project management software for a 12-person agency" is a prompt, and it pulls a different source set entirely.

Step 2: Log which sources get cited for each prompt

Now go find out which sources AI Overviews cites for your actual list. Not in general. For you.

Open an incognito window set to your target locale and run each prompt. For every AI Overview that fires, record the prompt, the date, and every cited URL with its domain. Do the same in the AI Mode tab for your top ten prompts, because the two surfaces will hand you different answers.

A spreadsheet is enough. One row per citation, with columns for prompt, query class, domain, URL, and whether the source is yours, a competitor's, or neutral.

If manual checks stop scaling past 30 or 40 prompts, add tooling. Google Search Console shows your own AI Overview impressions. Ahrefs Brand Radar, Semrush and SE Ranking all track AI visibility from a fixed prompt list, and there are free AI Overview citation checkers for spot checks.

Pro tip. AI Overviews vary by location, history, and personalization, so a source you see once is not a source every user sees. Cross-check two methods before you bet a quarter's budget on a pattern.

How you know it is done. You can answer "which sources AI Overviews cites for our five most valuable prompts" with a list of specific URLs, not a guess.

Where people slip. They log domains and throw away the URLs. The URL is the useful part. A specific Reddit thread or Wirecutter roundup is a target you can act on. "reddit.com" is not.

Step 3: Sort the cited sources into buckets and name your gaps

Group every logged citation into buckets: video, community and forums, encyclopedic, news and PR, vertical review and comparison, official and academic, local, and first-party. Count the slots each bucket won across your prompt set.

That count is your real priority order, and it will not match the industry averages. A regulated health brand sees .gov and medical systems dominate with Reddit almost absent. A B2B software brand sees G2, Reddit, and comparison listicles. A local services business sees Yelp carry the weight. Your Google AI Overviews sources are vertical-specific, every time.

Now mark your presence in each bucket honestly: strong, thin, or nothing. The gaps that matter are the buckets that win the most slots for your highest-intent prompts and where you scored nothing.

Pick two. Maybe three. Not seven.

This is where competitor data pays. DeepSmith's AI Visibility shows which competitor pages earn citations for your tracked prompts and which sources AI cites most across ChatGPT, Gemini, Perplexity, Claude, and Google AI Mode, so you see the pattern instead of inferring it one incognito search at a time. Pair that cross-engine view with your manual AI Overview log, since AI Overviews is a Google Search surface rather than one of the five tracked engines.

How you know it is done. You have a ranked bucket list with slot counts, and two chosen channels with a named owner for each.

Where people slip. They chase the bucket with the biggest industry-wide number instead of the biggest number in their own log. YouTube leads globally and may be nearly irrelevant for your prompt set.

Step 4: Earn Reddit mentions the slow way

For most non-regulated categories, Reddit is the highest-leverage single way to get into AI Overviews, and it is the one people get most wrong.

What gets cited is specific, first-hand experience. "I ran this on a 40-person team for six months, here is what broke" beats any brand post ever written. Numbers, timeframes, and real trade-offs are the texture Google's systems pick up. Comments get cited as often as posts, and a thorough reply on an existing high-visibility thread is usually your fastest yield.

Subreddit fit is the strongest single signal. Find the 10 to 30 subreddits where your buyer actually lives, then search your target prompts and note which threads Google already cites. Those exact threads are your targets.

Post two or three times a week and comment five to ten times, from a real account with history, for months. Mention your product only when it is the honest answer to the question in front of you. Consider an AMA if you have someone with genuine expertise, since the question-and-answer structure gets cited unusually well.

Common mistake. Buying upvotes does not work. Neither does asking employees to post fake praise. Reddit's anti-astroturfing is strong, and clustered positive content from accounts with no history tends to get downweighted rather than cited. Building a credible presence over months is what works. There is no shortcut.

How you know it is done. You have accounts with real history in your target subreddits, and at least a few of your comments sit in threads that already earn AI Overview citations.

Where people slip. They post only under a brand handle and only about themselves. Google reads that as promotional, and promotional content is what these systems are built to discount.

Step 5: Build YouTube answers Google can lift

If your prompt log shows video winning slots, this is your highest-return channel, and the bar is lower than most teams assume.

Title the video the way someone asks the question out loud. "How to clean a dishwasher filter" gets cited. "Dishwasher Filter Cleaning Tutorial!" does not. Answer inside the first 60 seconds, because burying the answer behind an intro is measurably costly. Add chapter markers so the transcript parses into liftable chunks. Fix the auto-captions by hand, since the transcript is what the retrieval layer actually reads.

The most repeatable play: take the blog post that already performs and film the same expert delivering the same content. Chapters, corrected captions, transcript on the page. Clips of 30 to 60 seconds cover quick how-to queries. Eight-minute-plus versions cover the explainers.

How you know it is done. For each of your top video-class prompts, you have a video that opens with the answer, carries chapters, and has a human-corrected transcript.

Where people slip. Clickbait titles. The relevant-content version of a video gets cited noticeably more than the dramatic one, so you are trading citations for thumbnails.

Step 6: Run a PR engine aimed at earned media, not distribution

Earned media accounts for 82 to 95 percent of AI citations. Paid and advertorial content accounts for about 0.3 percent. Press releases sit under 2 percent. That gap is the whole strategy.

Journalists cite people who give a clean, quotable, specific answer. Build two things: a small set of named experts on your team who will actually respond, and proprietary data worth reporting on. Original research is the most cited thing in news, and it compounds, because the article that cites your survey becomes the source a Wikipedia editor reaches for later.

Work the sourcing platforms with a real cadence. Featured is the largest general pool since HARO became Connectively, and Qwoted, Source of Sources, and Help a B2B Writer all still produce. Two to four genuine stories a quarter is realistic for a lean team.

How you know it is done. You have a named expert, a data asset, and a weekly habit of answering journalist requests. Not a press release calendar.

Where people slip. Confusing distribution with earned media. Wiring a release for a non-news event feels like PR and produces almost no citations.

Step 7: Get onto the review and comparison sites your buyers check

For commercial queries, review and comparison content is the most cited format there is. Listicles carry a striking share of citations on product and recommendation queries, and the ones that win have clear ranking criteria, a comparison table, an explicit methodology, and a takeaway near the top.

Your job has two halves. First, be present where the reviews live: keep an accurate, complete profile on G2, Capterra, Trustpilot, or whatever your category's equivalent is, and ask happy customers to review you there as a standing motion, not a campaign. Second, be reviewable: send review units with real technical documentation to the publications that cover your space, and give reviewers the measurable specifics they need to compare you fairly.

How you know it is done. Your profiles are accurate and gaining reviews steadily, and at least one independent publication has covered you in the last two quarters.

Where people slip. Paying for placement and expecting citations. AI Overviews pull from organic user reviews, not from paid quadrant positions, and Google has taken action against manipulated review schemes.

Step 8: Cover the authority layer your vertical demands

Some buckets cannot be bought or posted into, and pretending otherwise wastes a quarter.

Wikipedia does not accept promotional edits, ever. The real path runs upstream: earn coverage in the news articles and papers that Wikipedia editors cite, and the Wikipedia mention tends to follow. If your team has genuine topical experts, they can contribute as editors under the conflict-of-interest rules.

Government and academic sources work the same way. Submit public data to open research repositories, partner with university researchers on published work, file substantive comments in regulatory dockets. Your content becomes the source the .edu page cites, and the .edu page is what gets surfaced.

Selling locally? Yelp deserves its own line. Keep the profile accurate, keep review velocity steady, respond to reviews, and corroborate your location pages elsewhere. AI Overviews now appear on a large majority of local searches, and Yelp is the source they trust most for evaluative ones.

How you know it is done. You have one active, realistic play in the authority bucket your prompt log says matters, and you stopped spending on the ones it says do not.

Where people slip. Hiring a service that promises a Wikipedia page. Wikipedia tracks those accounts and bans them, and the cleanup costs more than the entry was worth.

Step 9: Measure the shift after 60 to 90 days

Give the engines time. Reddit mentions can surface in a fresh query within hours. A news or Wikipedia citation on an evergreen query usually takes days to weeks. Judging this at week three tells you nothing.

At 60 and 90 days, re-run the prompt log from Step 2 and compare citation sets. Look for three things: whether URLs mentioning you now appear, whether the buckets you invested in gained slots, and whether the sources citing you are tied to high-intent prompts. That comparison is the whole of Google AI Overviews AEO measurement, and it is not complicated.

DeepSmith tracks per-prompt mention and citation rates with page-level attribution across the five engines it covers, including Google AI Mode, so the cross-engine half runs on a schedule instead of a spreadsheet. Keep your manual AI Overview checks alongside it.

Then act. Double the channel that moved. Cut the one that did not, without ceremony.

How you know it is done. You can name which channel produced citations and which produced activity, and you have moved budget accordingly.

Where people slip. Measuring clicks. AI Overviews reduce clicks to the results beneath them, and only a small fraction of users click a source link. Track mentions, citations, and share of voice instead, because influence is what this channel produces.

What to do next

Do not build all nine steps this month. Run Steps 1 and 2 this week against ten prompts. That alone teaches you more about your Google AI Overviews AEO position than any industry study can, because it is your query set and your competitors.

Then pick one channel from your gap list and give it a full quarter. One channel done properly beats five started and abandoned. If Reddit is your bucket, that is 12 weeks of real participation. If YouTube, eight to ten videos with corrected transcripts.

If the bottleneck turns out to be production rather than knowing what to do, name it honestly. Mapping prompts is a morning's work. Consistently producing the content that earns and supports those mentions is not. DeepSmith is built for that gap: it tracks where you show up in AI answers, finds the prompts you are losing, and produces publish-ready articles grounded in your own brand context, with heading structure, internal links, and metadata already built in. You can start a 7-day free trial and see real data and real drafts before you pay.

Take it one channel at a time. You are closer than the empty citation list makes you feel.

Frequently asked questions

Which third-party sources does Google AI Overviews cite the most?

YouTube, Reddit, Wikipedia, and Google's own properties lead across most query sets, followed by major news publishers, authoritative vertical review sites, and .gov and .edu on YMYL topics. Yelp leads local discovery. The mix shifts by vertical, so the list of Google AI Overviews sources that matters is the one you build from your own prompts.

Does AI Mode cite different sources than AI Overviews?

Yes, and the gap is larger than most teams expect. One 2025 analysis found only 13.7 percent URL overlap on the same queries. AI Mode returns more citations per answer and spreads them across more domains, and AI Mode third-party sources lean harder on authoritative publishers and comparison sites while AI Overviews cluster on forums and video.

Can I get cited by optimizing my own site without third-party mentions?

Rarely on its own. First-party brand pages take roughly 8 to 25 percent of citation slots, and only when the brand owns genuinely category-defining content. The remaining majority goes to third parties, which makes earned mentions a requirement rather than an extra.

How long until an earned mention shows up as a citation?

It depends on the source and the query. A Reddit thread can be cited within hours when the query is fresh and the semantic match is tight. News and Wikipedia citations on evergreen queries usually land within days to weeks. Re-check at 60 and 90 days rather than at week one.