Anthropic $1.5 Billion Settlement Over Pirated Book Training Data
Anthropic Settles Copyright Claims for $1.5 Billion
Anthropic has reached a $1.5 billion settlement to resolve legal disputes regarding the use of pirated books in the training of its Claude AI models. The settlement addresses the legality of the data acquisition process rather than the act of AI training itself, providing a financial resolution for affected authors and publishers.
Key Financial Terms of the Settlement
The settlement distributes the $1.5 billion fund across authors, publishers, and legal counsel, though the per-work payout is relatively low compared to the total sum.
- Payout per Title: Eligible titles will receive approximately $3,000 each. For traditional publishing contracts, this amount is typically split between the author and the publisher.
- Legal Fees: The presiding judge reduced the class counsel's fees from 12.5% ($187.5 million) to 6.8% ($101 million).
- Representative Compensation: The three class representatives will receive $15,000 each.
- Payment Structure: The $1.5 billion is paid in installments, which effectively makes the class of plaintiffs creditors of Anthropic, linking the payout's completion to the company's ongoing solvency.
Legal Distinction: Piracy vs. Fair Use
A critical aspect of this case is the legal distinction made by Judge Alsup between the acquisition of data and the use of that data for machine learning.
The ruling establishes that while training Large Language Models (LLMs) on books may be considered fair use, the act of using pirated copies of those books to conduct that training is not.
This distinction means that Anthropic was found liable for the piracy involved in building its training library, but the resulting model—and the act of the AI "learning" from that data—does not necessarily violate copyright law. This outcome allows Anthropic to keep its models operational while paying a fine for the method of data collection.
Industry Implications and Critical Perspectives
The settlement has sparked significant debate among legal experts, authors, and the tech community regarding the adequacy of the penalty and the precedent it sets.
The "Cost of Doing Business"
Many critics argue that a $1.5 billion settlement is insufficient for a company of Anthropic's valuation. Some observers characterize the payment as a "speeding ticket" or a "cost of doing business," suggesting that the financial penalty is negligible compared to the commercial advantage gained from the training data.
Impact on Future AI Competition
There is concern that such settlements create a "financial moat" for established AI labs. Large companies can afford to pay billions in settlements to "clean up" their data provenance after the fact, whereas new startups may lack the capital to resolve similar copyright disputes, potentially stifling competition in the frontier model space.
Author and Creator Concerns
Authors have expressed frustration over the payout structure and the lack of ongoing royalties.
"A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas."
Some authors suggest that a more sustainable model would be similar to the UK's Public Lending Right, which provides ongoing payments to authors when their works are accessed via libraries, rather than a single lump-sum settlement.
Comparison to Individual Piracy
Discussion has highlighted the disparity between how corporate and individual piracy are treated by the legal system. Commenters frequently cited the case of Aaron Swartz, who faced severe federal prosecution for downloading academic articles, contrasting his experience with Anthropic's civil settlement for the mass use of pirated works.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch