AI Training and the Destructive Scanning of Rare Books

AI Training and the Destructive Scanning of Rare Books

AI Companies Employ Destructive Scanning for Training Data

AI companies are bulk-buying rare and out-of-print books, scanning them using high-speed machines that cut the spines off, and shredding the original physical copies. This process, referred to in some court documents as "Project Panama," aims to obtain vast amounts of high-quality text for AI training while bypassing expensive licensing fees from publishers.

Key details of the operation include:

  • Facilitation via ISBNdb: A service called ISBNdb reportedly facilitates orders of up to one million books and maintains buyer anonymity.
  • Targeting Pre-2022 Content: Books published before 2022 are considered premium training data because they are free of AI-generated text.
  • Strategic Hiring: Anthropic reportedly hired the former head of Google Books partnerships to assist in the goal of obtaining "all the books in the world."
  • Operational Secrecy: The process is often shielded by NDAs and framed as "digital preservation" to avoid public backlash.

Legal Justification and the "Fair Use" Ruling

A federal judge has ruled that the practice of destructive scanning is legal under the doctrine of fair use. The ruling suggests that because the original physical copy is eliminated, only one copy of the work exists at any given time (the digital version), which mitigates certain copyright infringement claims.

Critics argue that this is a loophole created by restrictive U.S. copyright laws. Some observers suggest that if copyright laws were more flexible—allowing for AI training without the destruction of materials—companies would not feel compelled to use destructive methods to avoid "extortion fees" from copyright holders.

Risks to Cultural Heritage and Historical Accuracy

The irreversible nature of shredding rare books creates several critical risks for the preservation of human knowledge:

Loss of Primary Sources

Unlike scraping a website, which can be reversed or re-uploaded, the destruction of a physical book is permanent. Once the original is shredded, there is no physical primary source left to verify the accuracy of the digital scan.

Erasure of Physical Artifacts

Digital scans capture text but lose the physical context of the book, including:

  • Marginalia: Handwritten notes in the margins that provide historical context.
  • Materiality: The specific printing, binding, and paper quality of the era.
  • Provenance: The physical evidence of the book's history and ownership.

Potential for Historical Revisionism

There is significant concern that controlling the only existing copy of a text allows for the electronic rewriting of history. Without a physical original to check against, digital copies could be altered to fit specific political or corporate narratives without detection.

Community Perspectives and Counterpoints

Discussion among technical and archival communities reveals a divide in how this practice is perceived:

Arguments Against the Practice

Many view this as a modern form of "book burning" or a "Library of Alexandria 2.0" event.

"You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data."

Arguments in Defense or Mitigation

Some argue that the value of a book lies in its information, not the physical medium, and that digitizing rare works makes them more discoverable via electronic search than they would be in a locked archive.

Others point out that libraries have long practiced "weeding" (the systematic removal of low-demand materials), and that the AI companies are simply applying a commercial version of this process to books that are often out of print and ignored by the public.

Proposed Alternatives

Suggestions to mitigate the loss include:

  • Mandatory Public Release: Requiring AI companies to make the digital copies of destructively scanned books public after a certain period.
  • Sovereign Datasets: Moving toward public-private partnerships where the public owns the training datasets and underlying models, while private companies handle the production of the LLMs.

Sources