
A US federal judge has given final approval to Anthropic's landmark $1.5 billion settlement with authors who accused the artificial intelligence company of using pirated books to train its AI assistant and large language model Claude. According to reports from Business Standard, the settlement was approved by a judge in San Francisco and comes as AI developers including Anthropic, OpenAI and Meta face growing lawsuits from authors, publishers, news organisations and other copyright owners over the use of copyrighted material to train large language models. The settlement establishes the first major financial and operational blueprint for resolving mass IP litigation in the generative AI era, as reported by RMN Digital, proving that the 'free scraping' era of AI development is officially over. Technology leaders must realize that data acquisition liabilities will increasingly be passed down to commercial enterprise subscribers, with the agreement representing a single milestone in an expanding matrix of global litigation challenging web scraping practices across literature, journalism, music, and software code.
US District Judge William Alsup drew a clear distinction between books Anthropic had legally purchased and books obtained through piracy. As reported by Business Standard, he ruled that training AI models using legally acquired books amounted to fair use under US copyright law because the court found that the books were not being copied to replace the originals, but were being used to teach the AI model how language works. That made the use 'exceedingly transformative' under US copyright law. However, regarding pirated books, he ruled that downloading and maintaining a permanent internal library built from illegally obtained books constituted copyright infringement, even if those books were later used for AI training. According to NPR, Anthropic's deputy general counsel Aparna Sridhar stated that "Training AI on books is fair use under copyright law," emphasizing that the company paid for the books it used. A federal judge also ruled in favor of Meta in a similar case last year involving authors including Richard Kadrey and Sarah Silverman, who sued over pirated copies of their novels being used to train AI models.
Court files exposed Anthropic's internal book scanning program under the codename Project Panama, revealing how the company systematically destroyed millions of books to train Claude. According to the court order, Anthropic purchased millions of used books, sliced off the covers, and scanned every page using hydraulic cutters to remove spines. The program began with cofounder Ben Mann downloading Books3 in early 2021, a library of 196,640 pirated titles, followed by five million more from Library Genesis in June 2021 and two million from the Pirate Library Mirror in 2022. Contractors then cut the bindings, trimmed the pages, and scanned each copy, with the company hiring Tom Turvey in February 2024 to obtain "all the books in the world." The Washington Post reported that Anthropic used brokers Better World Books and World of Books for the book-buying service, while the company raised $65 billion in May at a $965 billion valuation on revenue running at $47 billion annually.
The Anthropic settlement represents a single milestone in an expanding matrix of global litigation challenging web scraping practices across multiple sectors. According to RMN Digital, ongoing cases include OpenAI & Microsoft facing mass scraping of fiction and non-fiction books for ChatGPT base models (consolidated in New York federal court with active discovery phase), Getty Images against Stability AI for scraping millions of stock photos and reproducing watermarks in synthetic images (active parallel litigation in both US and UK courts), and RIAA, Sony, Universal, Warner against Suno & Udio for training music-generation algorithms on commercial sound recordings (labels seeking statutory damages of $150,000 per infringed song). The Chicago Tribune is also pursuing Perplexity AI over systematic news scraping to generate direct answers, while Encyclopedia Britannica has filed against OpenAI for unauthorized scraping of structured reference databases. Some authors and publishers opted out of the Anthropic settlement and continue to pursue separate lawsuits, with the agreement applying only to past claims involving books included in the class list and not preventing future litigation over new copyright issues.