Thursday, 23 July 2026 · Europe
EUR/USD 1.141 EUR/GBP 0.8534 EUR/CHF 0.9268 EUR/PLN 4.33 All rates →
Sign in · Join
EUROPES The European Report
European Edition Thursday, 23 July 2026
LATEST
Tech & Startups

AI labs secretly buy and destroy old books for clean training data

AI labs secretly buy and destroy old books for clean training data

Artificial intelligence companies are quietly purchasing and destroying millions of pre-2022 printed books to secure uncontaminated training data, exposing a stark legal divide with European copyright standards.

Data broker ISBNdb has begun sourcing physical books in bulk for AI laboratories to scan. The process requires workers to slice the spine off each book so the loose pages can feed through high-speed machines. Because the destruction is permanent, ISBNdb offers clients strict secrecy. “Strict NDA on every engagement,” the company states, ensuring buyers’ names are “never disclosed.”

The rush to acquire physical paper stems from a fundamental flaw in the AI supply chain. So much of the open web is now machine-generated that models risk training on their own output, a documented degradation known as model collapse. Printed books published before 2022 predate the large language model era. “The world’s best AI training data is sitting on a shelf,” ISBNdb advertises, describing the material as “dense, edited, authoritative.”

Physical archives also protect AI developers from a growing counteroffensive by authors. Writers are increasingly using tools like Nightshade to inject hidden characters into their text that corrupt AI models during training. Pre-2022 print runs cannot be retrospectively poisoned, giving laboratories a verifiable, clean chain of custody.

For European publishers and investors, the strategy underscores a stark transatlantic legal divergence. A US judge recently approved Anthropic’s $1.5bn settlement over pirated books and ruled that training on purchased, scanned books constitutes fair use, partly because the original physical copy is destroyed.

This US judicial blessing stands in sharp contrast to the European Union’s text and data mining rules, which generally require rights holders to opt out rather than opt in. Ingram, the largest book distributor in the US, has already warned publishers about the bulk scanning and offered an opt-out mechanism to protect their catalogs.

Despite the US legal cover, the industry remains deeply aware of the reputational damage. “The optics problem is real,” ISBNdb warns on its website. “‘AI company destroys two million books’ is not a headline that generates sympathy.” The firm suggests rebranding the destruction as digital preservation. Yet as the synthetic text flood rises, the market has made its calculation: the scarcest resource in artificial intelligence is now a sentence guaranteed to have been written by a human.

More from Tech & Startups