DeepSmith

Sep 26 · AEO & AI Visibility

13 min read

How the Legal Fight Over 'Publicly Available' Content Is Reshaping AI Crawler Rules

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
A monochrome illustration of a dashed boundary around a document icon, with crawler nodes reaching toward it and a scale in the background, under the text The Legal Line Around Public Content.

A website being reachable without a login does not settle whether an AI company can legally copy what's on it for training. That's the plain answer to a question a lot of site owners are asking right now: is it legal to train AI on my website's content? The honest answer is that nobody can give you a clean yes or no yet, because the answer depends on copyright, how the content was accessed, what it's used for, and where the company doing the crawling is based. What's actually changing, while the legal questions sit unresolved, is how AI companies build their crawlers and what kind of control they hand back to you.

Publicly available does not mean legally unencumbered

It helps to slow down on that phrase "publicly available," because it's doing a lot of quiet work in this debate. The U.S. Copyright Office made the distinction directly in a pre-publication report it released in May 2025: publicly available can simply mean the page is sitting on the open internet. It does not mean authorized. Those are two different ideas that a lot of AI training discussions blur together.

Think about how many different situations get lumped under "public." A page might be reachable with no login. It might be indexed by a normal search engine. It might sit behind a subscription that someone bypassed. It might have been uploaded by a person who never had the right to post it in the first place. All of those pages are technically public in the sense that a crawler can reach them, but the legal story behind each one is different, and that difference matters when you're weighing whether publicly available data AI companies scoop up for training was ever meant to be used that way.

The Copyright Office's report also points out that Common Crawl, one of the largest open repositories used to build AI training sets, holds more than 250 billion pages spanning 17 years. That's a huge number, but it only tells you the size of the pile. It says nothing about whether every page in that pile was fair to copy. Scale and legality are two separate questions, and this whole debate keeps coming back to that separation.

The report also treats copying into a training set as an act that touches copyright's reproduction right, the same right that covers ordinary copying of a book or an article. Whether that copying is allowed usually comes down to fair use, and the Office is careful to say fair use isn't a formula you can run automatically. It's a judgment call that weighs the purpose of the use, the nature of the work, how much was taken, and the effect on the market for the original. None of that gets decided just because a page happened to be sitting online where anyone could load it, which is exactly why the AI training data legal question resists a one-line answer.

Training, search, and retrieval are now separate crawler choices

Here's where things get more useful for a site owner, even without a settled legal answer. The big AI companies have started splitting their crawlers by purpose, so that allowing one thing on your site doesn't automatically mean you've allowed everything else.

OpenAI's documentation lays out three separate bots. One of them, OAI-SearchBot, exists to surface your pages in ChatGPT search results, and OpenAI says it isn't used to train its models. ChatGPT-User visits a page only when someone asks ChatGPT a direct question about it. GPTBot is the one tied to training: OpenAI says that if you block GPTBot, you're telling it your content shouldn't go into that training pipeline. You can let OAI-SearchBot in for search visibility while keeping GPTBot out, and under OpenAI's own stated policy, that's a meaningful distinction, not a technicality.

Google draws a similar line. Googlebot is the crawler behind ordinary Google Search, and it respects robots.txt like it always has. Google-Extended is different: it's a robots.txt control specifically for whether your content trains future Gemini models and grounds Gemini's answers. Google says using Google-Extended has no effect on your regular search rankings or your presence in Google Search. That's the same idea as OpenAI's split, just under a different name.

Perplexity says something similar but frames it around indexing rather than training. PerplexityBot respects robots.txt and won't index the full text of a page that blocks it, though it may still show the domain and a short factual summary. Perplexity also says letting its bot index your page doesn't mean that page feeds a foundation model, since Perplexity says it doesn't build one. It's worth knowing, though, that Perplexity also works with other companies' crawlers to build its index, so how Perplexity crawls and fetches pages matters even if you've only checked one vendor's policy.

Anthropic runs the same kind of setup: ClaudeBot for training collection, Claude-User for answering a specific question a person asked, and Claude-SearchBot for search indexing. Blocking one doesn't automatically block the others, and Anthropic's own page is clear that its policies apply going forward, not retroactively to content already collected. Each of these splits is a company describing its own product, not a court settling the AI training data legal question for anyone else.

If you're trying to decide what to allow, your robots.txt setup for AI crawlers needs to name the specific bot you're trying to stop, because a broad instruction that seems to cover "AI" often misses the exact crawler doing the thing you're worried about. And robots.txt itself has limits worth knowing before you lean on it as your whole answer, which is part of why some sites are also looking at whether llms.txt actually moves the needle as a companion signal.

What the U.S. cases actually establish

A handful of court cases get cited constantly in this debate, and it's worth being precise about what each one actually decided, because none of them hands down a general rule for every website.

The New York Times' amended complaint against OpenAI and Microsoft, filed on April 15, 2025, is still just a complaint. It alleges that OpenAI's training set included content scraped from Times websites, with NYTimes.com showing up 333,160 times in one dataset the Times pointed to. It alleges Microsoft's Bing Chat sometimes reproduced or closely summarized Wirecutter recommendations without a clear link back. Those are allegations the Times is trying to prove, not facts a court has confirmed.

Thomson Reuters v. Ross Intelligence is a case that actually reached a decision. A Delaware court ruled in February 2025 that Ross's use of roughly 25,000 memos built from Westlaw's headnotes was not fair use, finding actual copying of 2,243 of 2,830 headnotes at issue. The court leaned on the fact that Ross used the material to build a directly competing legal research product, which made the use commercial and non-transformative on the record in front of it. That's a specific fact pattern about a competing product, not a blanket statement about crawling websites.

Kadrey v. Meta went the other direction on a different set of facts. In June 2025, a California court granted Meta partial summary judgment over authors' claims that training Llama on their books infringed copyright, even though Meta had downloaded the books from shadow libraries. The court treated training a general-purpose model as somewhat transformative and found the authors hadn't proven the market harm they claimed, partly because Llama couldn't reproduce more than about 50 words from any of the books in testing. The judge was explicit that the ruling was narrow and tied to the record in that case, not a general green light for AI training.

Bartz v. Anthropic is still unresolved. The August 2025 order describes Anthropic downloading millions of books from pirate libraries, later buying and scanning copies for a research library, and recopying subsets for training. The court rejected the idea that a company could copy everything it wanted and rely on the fact that only some of it got used for training, and it kept the case heading toward trial rather than settling the core question. It's one more data point in the AI copyright crawling debate, not the case that resolves it.

Put together, these cases show courts working through specific facts about specific companies, not agreeing on one shared rule about what any AI company may do with any public website. That's part of why the AI copyright crawling debate keeps generating headlines without generating a settled answer. Separately, if you write with AI tools yourself, whether that output can be copyrighted is its own unresolved question worth knowing about.

Regulators are moving toward transparency and machine-readable preferences

While the courts work through specific lawsuits, regulators and infrastructure providers have been building rules and tools that don't require waiting for a final verdict.

AI crawler regulation is furthest along in the EU and UK, where the response so far leans on transparency rather than a ban or a blanket permission. The European Union's AI Act rules for general-purpose AI took effect in August 2025, and they require AI providers to publish a summary of the data used to train their models, including which large datasets and domains were involved. That's a transparency requirement, not a ruling on whether any specific use was lawful, and the EU's AI Office won't take on full enforcement responsibility until August 2026. Separately, the EU's older Digital Single Market copyright directive already lets rightsholders reserve their rights against text-and-data mining through machine-readable signals, which is a more direct lever than the AI Act's disclosure rules.

The UK ran its own consultation on copyright and AI from December 2024 through February 2025, weighing options that ranged from leaving the law as it is to requiring licenses for AI training copies. Its March 2026 report said the government's preferred option, a broad exception with an opt-out, was rejected by most people who responded, and 81% of respondents actually wanted stronger copyright protection with mandatory licensing instead. The government said it would hold off on reform until it was confident any change would work, which leaves the UK's rules about where they started: unsettled, with more evidence gathering ahead.

Infrastructure providers have moved faster than either courts or regulators. Cloudflare announced a permission-based approach for AI crawlers on its network in July 2025, with new domains defaulting to blocking AI crawlers unless the site owner chooses to allow them. Cloudflare also introduced a way for site owners to charge for crawler access, sometimes called pay-per-crawl, which is a business arrangement layered on top of the access question rather than an answer to the underlying legal one. A long list of publishers signed on in support, which tells you how much appetite there is for more direct control, even without new legislation forcing it. Alongside all this, an industry workshop report from the IAB in September 2025 pointed out something practical: AI vendors don't treat robots.txt consistently, and a single instruction doesn't always map cleanly onto training, search, and retrieval the way a site owner might hope.

What this means for a site owner

None of this adds up to a simple rule you can follow, and that's the honest takeaway. What it does give you is a clearer set of separate decisions to make, instead of one big "allow AI or don't" switch.

Start by naming which use you actually care about. Do you want your pages to show up in AI search results? Do you want a chatbot to be able to fetch a page when someone asks about it directly? Do you want your content collected for model training? Those are different questions, and as the OpenAI, Google, Perplexity, and Anthropic examples show, vendors increasingly let you answer them separately instead of forcing one answer for everything.

Next, remember that a page being reachable doesn't tell you how it got that way or what happens to a copy once it's made. Courts and regulators keep circling back to source and access, not just visibility, so treating "it's public" as the end of the analysis skips over the part that actually seems to matter to the people deciding these cases. The publicly available data AI systems gather still carries a separate question about how it was obtained, and that question doesn't go away just because the page loaded without a password.

If you decide you want to restrict some crawlers and not others, a visibility-versus-control framework can help you think through what you give up by blocking a bot versus what you're trying to protect. And once you've set your preferences, it's worth checking that the traffic hitting your site actually matches what it claims to be, since AI crawler identities can be spoofed and a blocked bot doesn't always stay blocked if something else is impersonating it.

It also helps to understand the mechanics behind the policies you're reading. Knowing how ChatGPT fetches and renders pages makes vendor documentation easier to apply correctly. The same goes for knowing how Google's AI Overviews read sites, instead of guessing at what a setting actually does.

Finally, keep licensing in the back of your mind as a separate track from all of this. Regulators in both the US and the UK keep pointing at licensing as the mechanism that might eventually settle some of these disputes without new legislation, and it's a choice that exists independently of whatever you decide about blocking or allowing a crawler today. None of that requires you to have a legal opinion of your own. There's still no single place that answers is it legal to train AI on my website with a plain yes or no, only clearer categories of choice. AI crawler regulation is being written in real time, through court opinions, government reports, and infrastructure defaults, and the most useful thing you can do right now is separate your own preferences clearly rather than wait for one authority to hand you a final answer.

Frequently asked questions

Is content on my public website automatically free for AI training?

No. The U.S. Copyright Office says plainly that publicly available is not the same as authorized. Being reachable without a login is one fact among several. Copyright, how the content was accessed, licensing terms, the purpose of the use, and market effects can all matter to the analysis.

If I allow an AI search crawler, am I also allowing model training?

Not necessarily. OpenAI, Google, Perplexity, and Anthropic each describe separate crawlers or controls for search, user-requested retrieval, and training. The exact effect depends on which vendor you're dealing with and which of its settings you've changed.

Does robots.txt settle the copyright question?

No. Robots.txt communicates a preference, and some commenters have argued that ignoring it could matter to a fair-use analysis, but the Copyright Office and an industry workshop report both note that vendors treat the protocol inconsistently and that it wasn't built with generative AI training specifically in mind.

Have courts actually decided whether training AI on public content is legal?

Courts have ruled on specific disputes, not the whole category. Ross lost its fair-use defense on a record involving a directly competing product. Meta won limited summary judgment on a record involving books and measurable output limits. The Anthropic case is still headed toward trial on its central questions. None of these rulings applies automatically to every website or every AI company.