DeepSmith

Sep 26 · AEO & AI Visibility

14 min read

What Blocking AI Crawlers Actually Costs You: The Evidence So Far

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
A monochrome illustration of small robot crawler icons approaching a website page, with one path blocked by a barred circle icon and another path reaching a chat-bubble answer card, above the text The Cost of Blocking AI Crawlers.

If you've been asking yourself should I block GPTBot, here's the honest answer up front: the cost of blocking AI crawlers is mixed, but it leans toward a real downside. The claim that blocking automatically erases you from AI search is too strong. The claim that blocking is a free, no-cost move is also too strong. What the data actually supports is somewhere in between, and which side you land on depends on which crawler you block, which AI engine you're worried about, and whether you're a large publisher or something else entirely. Call the overall verdict mixed, with a credible downside signal, and let's walk through why.

Not every "AI crawler" is doing the same job

Before any of the numbers make sense, you need to separate what these crawlers are actually for, because "block AI crawlers" gets used to describe several different actions that have different consequences.

Training crawlers, like GPTBot and ClaudeBot, collect pages to help build or improve a model. Search crawlers are a separate thing. OpenAI names OAI-SearchBot specifically as the crawler that surfaces websites in ChatGPT search, and its own documentation says the robots.txt controls for GPTBot and OAI-SearchBot are independent of each other. There's also a third category, the user-triggered fetcher: OpenAI's ChatGPT-User handles requests that happen because a person asked ChatGPT to go look at a specific page, and it isn't used to crawl the web on its own.

This distinction matters because it directly answers the question a lot of people are actually asking, which is whether blocking GPTBot removes them from ChatGPT Search. Based on OpenAI's own documented setup, it doesn't. Blocking GPTBot is a training-access decision. If you want to opt out of appearing in ChatGPT Search results, the crawler that matters is OAI-SearchBot, not GPTBot. Confusing the two is where a lot of the fear around GEO visibility loss actually comes from.

Google draws its own line, and it's simpler: to show up as a supporting link in AI Overviews or AI Mode, a page has to already be indexed and eligible to appear in regular Google Search. Google says there's no separate AI-specific technical requirement, no special markup, and no extra file you need. So a site can block GPTBot entirely and still show up in Google's AI features, as long as Google's own crawler can still reach, index, and serve the page normally.

One more wrinkle worth knowing: a robots.txt rule is a request, not a lock. The Register reported in December 2025 that robots.txt compliance is voluntary, meaning a listed disallow doesn't guarantee the crawler actually stays away, and a site without a listed rule might still be blocked at the server or CDN level instead. Keep that in mind every time you read a statistic that's based on scanning robots.txt files: it's measuring declared intent, not necessarily what's really happening on the wire.

AI crawling grew fast, but it isn't one uniform trend

The case for taking this seriously starts with the AI crawler traffic trends underneath the debate: how much AI crawl volume has actually grown. Cloudflare ran a fixed-cohort comparison of the same customers in May 2024 versus May 2025, specifically to avoid the bias of just counting new customers joining the network, and found combined AI and search crawler traffic rose 18% year over year. Inside that traffic mix, GPTBot's raw request volume jumped 305%, actually outpacing Googlebot's 96% growth over the same period. But don't read that as GPTBot overtaking Google: Googlebot still made up half of the tracked crawler traffic by May 2025, while GPTBot was at roughly 7.7%. Some crawlers moved the other way entirely. ClaudeBot's raw requests fell 46%, and Bytespider's fell 85%, over the same window. AI crawler activity isn't one smooth upward line, it's a set of individual crawlers each doing something different.

Other vendors are seeing the same shift from a different angle. DataDome reported that AI and LLM traffic nearly tripled across its customer base between January and May 2025, going from 2.6% to 7.2% of total bot traffic. Fastly, drawing on visibility into more than 130,000 applications and APIs, found that AI crawlers made up almost 80% of all observed AI bot traffic in its mid-2025 window, with Meta's AI bots and Google each accounting for a large chunk of it. These numbers come from different networks with different customers, so they shouldn't be merged into a single figure, but they all point the same direction in the broader AI crawler traffic trends: AI-related crawling is a meaningfully bigger share of traffic than it was a year or two ago, and it's worth understanding the shape of it on your own site before you decide anything, which is exactly what analyzing your own server logs for AI crawler activity is for.

Cloudflare also surfaced a mismatch that matters for the cost argument. Comparing crawl requests to actual referral visits, it found Google crawled about 14 times for every referral it sent back, while OpenAI's ratio was roughly 1,700 to 1, and Anthropic's was around 73,000 to 1. That gap doesn't prove blocking is free or that allowing crawlers pays off in traffic, but it does show that training-oriented crawlers can pull a lot of content while sending very little measurable traffic back in the short term, which is part of why some site owners have decided the trade isn't worth it.

Blocking has already gone from rare to common

If you're wondering whether you'd be an outlier for blocking, you wouldn't be, especially among publishers. The Reuters Institute looked at the most-used news sites in ten countries and found that by the end of 2023, 48% already blocked OpenAI's crawlers and 24% blocked Google's AI crawler, using archived robots.txt files from the Internet Archive. Blocking rates varied a lot by country, from 79% of US sites down to 20% in Mexico and Poland.

That trend kept climbing. The Register, citing BuiltWith data from December 2025, reported roughly 5.6 million websites listing GPTBot as disallowed, up from about 3.3 million just five months earlier, a jump of nearly 70%. ClaudeBot showed a similar climb, from about 3.2 million to 5.8 million blocked sites over the same stretch. A separate academic study published on arXiv in October 2025 tracked 3,369 reputable news and misinformation sites across six snapshots between September 2023 and May 2025, and found that reputable sites disallowed an average of 15.5 different AI user agents by the later snapshots, compared to fewer than one on misinformation sites. These are all measurements of declared robots.txt rules, not confirmed enforcement, but taken together they show blocking moving from a fringe move to a default one, particularly for publishers who feel the content-scraping side of this most directly. If you haven't checked whether your own edge setup is quietly doing this without a deliberate decision, auditing your CDN or WAF for AI crawler blocks is worth ten minutes.

The strongest evidence that blocking has a real cost

Most of the data so far describes volume and adoption, not consequences. The piece of research that actually gets at consequences is a study by Hangcheng Zhao and Ron Berman, out of Rutgers Business School and Wharton, examining how news publishers responded to generative AI. It combined SimilarWeb visit data, Comscore's US household browsing panel, Semrush traffic estimates, historical robots.txt records, and Internet Archive page histories, and used staggered adoption of blocking rules to compare publishers that had already blocked against publishers that hadn't yet.

The latest version of the study found publishers experienced roughly a 7% traffic decline within six weeks of blocking, consistent across SimilarWeb, Semrush, and Comscore's measures. The Comscore figure carries extra weight here because it's built to capture actual human browsing rather than every automated request, though the authors note that panel is smaller and the estimate less precise as a result. Earlier versions of the same research, looking at a longer post-blocking window for large publishers specifically, reported bigger declines: about 23% lower total visits by SimilarWeb's measure and about 14% lower human-browsing traffic by Comscore's. These aren't contradictory findings so much as different snapshots of an evolving analysis, covering different time horizons and publisher groups, so don't treat either number as the definitive cost of blocking. The honest summary is that blocking was associated with a real, measurable traffic decline in this publisher sample, not that every site loses exactly 7% or exactly 23%.

The study's authors describe reduced exposure in LLM results as a plausible explanation, and they found direct visits declining while organic search referrals held roughly steady before Google rolled out AI Overviews in May 2024, after which both kinds of traffic dropped, which is part of why the main analysis focuses on the earlier period. It's evidence of an association, and a reasonably strong one, but it stops short of nailing down the exact mechanism.

Why "block GPTBot, disappear from AI search" doesn't hold up

Here's the part that keeps this from being a simple cautionary tale. HasData built an AI Crawler Block Index covering nearly 10,900 domains, including almost 1,150 news publishers, and ran it against ten Google AI Mode queries. Among the domains that actually got cited in that general-news query set, 51.9% disallowed at least one AI crawler, more than three times the 15% rate across the full sample. In other words, sites that block AI crawlers showed up in Google AI Mode citations more often than the baseline, not less. HasData's own conclusion was that blocking GPTBot did not keep sites out of Google AI Overviews, which lines up with Google's stated position that ordinary Search eligibility, not any AI-specific rule, is what determines AI Overview and AI Mode inclusion.

That's the strongest evidence available against the idea that blocking AI bots backfire in a simple, direct way. Still, this doesn't mean blocking is harmless, or that it somehow helps. The HasData sample is small, based on just ten queries, and it's observational: it shows blocked sites getting cited, not that blocking caused the citation. What it does responsibly rule out is the strongest, simplest version of the scare story, that flipping a GPTBot disallow switches off your AI-search visibility across the board. The mechanism that actually governs Google's AI features is regular Search indexing, and the mechanism that governs ChatGPT Search specifically is OAI-SearchBot, not GPTBot. Getting cited in the first place depends on a separate set of relevance and structure signals, not on which crawler you allow in. Neither of those maps cleanly onto a single training-crawler block.

What the evidence actually adds up to

Put the pieces together and the honest verdict on the cost of blocking AI crawlers is conditional, not universal. AI crawling has grown into something large enough that ignoring it isn't really an option anymore, but it also isn't one unified channel: some crawlers surged, others shrank, and Googlebot remained the largest single crawler in Cloudflare's tracked mix. None of this replaces a root-cause diagnosis of why your own pages might already be missing from AI answers, block or no block. Blocking has become common, especially among news publishers, to the point where millions of sites now list GPTBot or ClaudeBot as disallowed. The best evidence available on actual cost, the Zhao and Berman publisher study, found a real traffic decline after blocking, though the size of that decline shifts depending on which version and which traffic measure you look at. At the same time, the claim that a GPTBot block wipes you off AI search doesn't survive contact with either the primary documentation from OpenAI and Google or the observational data from HasData.

What that leaves you with is a genuinely conditional trade-off: block, and you likely reduce unwanted automated requests and training access, at some risk of reduced future discovery or referral traffic, a risk that's better documented for large news publishers than for a typical SaaS site, ecommerce store, or local business. Allow, and you keep the door open for whatever citation or referral value an AI system might eventually send back, at the cost of feeding a crawler whose ratio of requests to referrals can be extremely lopsided. Some sites are experimenting with a third option, charging for crawl access instead of blocking it outright, which is a separate trade-off with its own math. Neither side of that trade has a single number attached to it yet, and anyone telling you blocking AI bots backfire always, or never, in a fixed percentage is rounding off a lot of nuance the underlying research doesn't support.

If you already track how AI engines mention or cite your brand, you have a more direct way to check this than guessing from someone else's publisher study: watch your own mention and citation rate before and after any change to your crawler rules, on the specific engines that matter to you. DeepSmith's AI Visibility tracking reports citation rate and mention rate by platform over time, so a change in your own numbers after a robots.txt edit tells you more about your actual situation than an industry-wide average ever will. You can see what that looks like on your own site with a free trial, real data before you pay anything.

What would change this verdict

Any real measure of GEO visibility loss depends on evidence this thin getting a lot better. Right now, the evidence base is still thinner than the size of the decision deserves. It would get a lot more solid with longitudinal studies that separate GPTBot, OAI-SearchBot, Googlebot, Google-Extended, and PerplexityBot from each other rather than lumping them together, with site-level before-and-after comparisons that confirm actual enforcement instead of relying on a robots.txt declaration, and with the Zhao and Berman results replicated now that AI Overviews and AI Mode have had more time to settle in. Until that exists, treat any single percentage you see quoted, including the ones in this piece, as a data point from a specific sample and time window, and as one input into a possible GEO visibility loss, not a universal cost of blocking.

If the question you actually came here with is should I block GPTBot on my own site, that's a separate decision about how you configure your robots.txt file, not what the evidence says happened to news publishers elsewhere, and it deserves its own decision process rather than a borrowed statistic.

Frequently asked questions

Does blocking GPTBot remove a site from ChatGPT Search?

Not based on OpenAI's own documentation. GPTBot and OAI-SearchBot are described as independent crawlers with independent robots.txt controls, and OAI-SearchBot is the one OpenAI connects to ChatGPT Search visibility. A GPTBot-only block shouldn't be treated as an automatic ChatGPT Search opt-out.

Does blocking GPTBot remove a site from Google AI Overviews?

There's no evidence here that it does automatically. Google says AI Overview and AI Mode eligibility runs through ordinary Search indexing and eligibility, and HasData's 2026 baseline found plenty of AI Mode-cited domains that blocked at least one AI crawler.

How much traffic has actually been lost after blocking AI crawlers?

The Zhao and Berman publisher study reported about a 7% decline within six weeks in its most recent analysis, with earlier versions describing larger declines, roughly 23% in total visits and 14% in human traffic, for large publishers over a longer window. These are sample-specific figures from news publishers, not a universal forecast for every site.

Is blocking AI crawlers already the norm?

It's common, particularly among news publishers. Reuters found 48% of prominent news sites across ten countries blocked OpenAI's crawlers by the end of 2023, and BuiltWith data reported by The Register put GPTBot disallows at around 5.6 million sites by December 2025. The exact definitions differ between these sources, so the numbers aren't directly interchangeable, but the direction is consistent.