Judge approves 1.5 billion Anthropic settlement over pirated training books

A federal judge has approved a 1.5 billion dollar copyright settlement between Anthropic and thousands of authors whose pirated books were used to train the Claude chatbot. The ruling makes it the largest known copyright recovery in history. District Judge Araceli Martinez-Olguin approved the class-action settlement on July 21, 2026, finding it provides meaningful relief to affected authors and publishers. The settlement covers more than 482,000 books, with approximately 91 percent claimed by a

2 min
Judge approves 1.5 billion Anthropic settlement over pirated training books

A federal judge has approved a 1.5 billion dollar copyright settlement between Anthropic and thousands of authors whose pirated books were used to train the Claude chatbot. The ruling makes it the largest known copyright recovery in history.

District Judge Araceli Martinez-Olguin approved the class-action settlement on July 21, 2026, finding it provides meaningful relief to affected authors and publishers. The settlement covers more than 482,000 books, with approximately 91 percent claimed by authors or publishers who are now due payment.

What happened

Anthropic downloaded millions of pirated books from sources including Library Genesis and Books3 and used them as training data for Claude. The company also purchased millions of print books, scanned them, and built a searchable digital library.

U.S. District Judge William Alsup, who issued the preliminary approval in September 2025 before retiring, delivered a mixed ruling in 2025. He found that training AI chatbots on copyrighted books constituted fair use under copyright law. But he ruled that Anthropic acquisition of those books through pirate websites was unlawful.

Settlement terms

Anthropic will pay approximately 3,000 dollars per book. The company is also required to destroy all pirated datasets. Once lawyers finalize the class list, Anthropic may owe an additional 3,000 dollars for every infringing work beyond the first 500,000.

Plaintiff attorney Justin Nelson called it the largest known copyright recovery in history and said distributions to the class would begin as soon as possible.

Anthropic deputy general counsel Aparna Sridhar emphasized the fair use finding, calling it a landmark showing that training AI on books is fair use under copyright law.

Context

The lawsuit was first brought in 2024 by bestselling thriller novelist Andrea Bartz and two other authors. It is the first major settlement among dozens of AI copyright cases still working through U.S. courts.

The case establishes a split precedent: AI companies can train on copyrighted works under fair use, but acquiring those works through piracy carries massive financial liability. For the broader AI industry, the settlement signals that the era of scraping pirated datasets for training data is over, even as the legal framework for licensed training data remains unsettled.

Sources

- Judge approves a 1.5B Anthropic settlement over pirated books used to train the Claude chatbot - ABC News, July 21, 2026: https://abcnews.com/amp/Technology/wireStory/judge-approves-15b-anthropic-settlement-pirated-books-train-134949964

- Judge approves a 1.5B Anthropic settlement over books used to train Claude - AP News: https://apnews.com/article/ai-anthropic-copyright-settlement-claude-books-bartz-74b140444023898aeba8579b6e9f0d63

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min