AI Training and the Physical Destruction of Books
AI Companies are Purchasing and Destroying Physical Books for Training Data
AI companies are reportedly acquiring millions of secondhand physical books through intermediaries, scanning them for training data, and subsequently destroying the physical copies. This practice is driven by the need for high-quality training data "untouched by machines"—specifically content published before 2022—to avoid the noise of AI-generated content now saturating the internet.
According to a report from Anna’s Archive, Anthropic’s "Project Panama" is a primary example of this trend. The project allegedly involved spending tens of millions of dollars to purchase, scan, and destroy millions of paper books to train the Claude LLM. The motivations for this strategy are three-fold:
- Competitive Advantage: Preventing competitors from scanning and using the same physical sources.
- Legal Mitigation: Reducing legal risks associated with digital copyright infringement.
- Cost Efficiency: Destroying books is often cheaper than maintaining lossless scanning archives.
The Risk of Knowledge Privatization
The systematic scanning and destruction of physical books creates a risk where human knowledge is permanently monopolized on private corporate servers. While AI assistants may become more capable, the underlying primary sources may disappear from the public domain. This creates a paradox where companies promising to make knowledge accessible are simultaneously dismantling the physical carriers of that knowledge.
Anna’s Archive has called for a global volunteer effort to scan and upload rare books, journals, newspapers, and magazines to a "digital library of Alexandria" to ensure that human civilization's memory remains free and accessible. This urgency is amplified by the following factors:
- AI Content Saturation: Since early 2025, AI-generated content has accounted for more than half of new internet publications, making it harder to distinguish human-authored work.
- Data Monopoly: If AI companies are the sole possessors of the digital scans of destroyed books, they control the access to that information.
Critical Perspectives and Counter-Arguments
Community discussion on Hacker News reveals significant skepticism and nuance regarding these claims. Several key counter-arguments have emerged:
The Role of Copyright Law
Some argue that the destruction of books is not a choice made by AI companies for malice, but a requirement of "Kafkaesque" copyright laws. Under certain legal frameworks, digitizing a book requires the destruction of the original to ensure the license is transferred rather than copied.
Scale and Rarity
Many critics question whether "rare" books are actually being targeted. They argue that AI companies are buying bulk secondhand books that were otherwise rotting on shelves or destined for pulping by professional book dealers. As one commenter noted:
"Actual professional specialized book dealers pulp books by the millions... model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys."
Institutional Preservation
Others point out that national libraries (such as the Library of Congress in the US or the British Library in the UK) have legal deposit requirements, meaning a copy of almost every published book is already preserved in a state-funded archive, making the total loss of knowledge unlikely.
The "Preservation" Paradox
Some argue that the act of scanning and training a model on a book is a form of preservation. If a book was obscure and unread, transforming it into weights within an LLM may be the only way its information reaches a broader audience, even if the physical copy is destroyed.
Summary of the Conflict
The tension centers on a fundamental disagreement over how knowledge should be stored and accessed. On one side is the push for open, decentralized shadow libraries to prevent corporate monopolies. On the other is the reality of a commercial book market and a legal system that incentivizes the destruction of physical media during the digitization process.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch