Is Training AI on Copyrighted Books Legal? Fair Use, Piracy and the $1.5B Question

Explore the distinctions between AI training, unlawful acquisition of books, fair use, market competition, and human authorship.

This content is blocked because it would connect to YouTube.
This content is blocked because it would connect to Spotify.

AI Training on Copyrighted Books: Key Legal Distinctions

Is it legal to train ChatGPT, Claude, Gemini, and other AI models on copyrighted books without authors’ permission? This Short explains why the answer depends on piracy, fair use, transformation, and market competition.

A court treated Anthropic’s AI training as lawful but penalized the company for obtaining books through illegal shadow libraries, leading to a $1.5 billion settlement. In a separate case, Thomson Reuters defeated Ross Intelligence because the competing AI legal platform used Reuters content for substantially the same market purpose.

Copyright law dates to 1976, so courts are applying old rules to generative AI. They consider whether a use is transformative, how much content was taken, its purpose, and whether the new product harms the original market. Fully AI-generated work faces another problem: courts have said copyright requires human authorship.

Source: TechCrunch — August 23, 2026
Author: Amanda Silberling
Expert analysis cited: Cathy Gellis and Jason Henderson

What to watch

Different cases can involve different conduct and market effects. A finding concerning training does not automatically excuse how source material was obtained. The episode explains the issues discussed in the reporting without treating one decision as a universal answer for every model, dataset, or jurisdiction.

Related reading: our coverage of AI-agent safeguards and OpenAI’s training pause.

Watch and listen

Watch the YouTube Short above or listen to the full episode on Spotify.

Scroll to Top