Exchanges

The Shredded Library: How AI Companies Turn Physical Books Into Training Data Ash

CryptoLion

Anthropic spent millions buying millions of physical books. Then they shredded them. The pages were scanned, OCR'd, and fed into a model. The paper became pulp. The covers were recycled. This is not a metaphor. It is the latest evolution of the AI training data pipeline.

Context

The hunger for high-quality, human-generated text has driven AI companies to scrape the entire open web, purchase closed datasets, and even license copyrighted works. The problem: most of that data is contaminated by AI-generated content, modern data poisoning techniques, and legal uncertainty. Enter the physical book. In 2025, a U.S. court ruled that digitizing a lawfully purchased physical book for non-distribution purposes — and destroying the original — constitutes fair use. The logic: if the digital copy exactly replaces the physical one (one-to-one swap), no additional copies exist, and the rights holder's market is not harmed.

This ruling unlocked a new data supply chain. ISBNdb, a database service, began offering a "destructive scanning" service: buy books, cut off the bindings, scan each page, then shred the original. AI developers are the target customers. Anthropic is the confirmed early adopter, having spent "millions of dollars" on millions of books. The service promises legally binding NDAs and verifiable destruction of the physical originals. The marketing material touts that "physical books printed before 2022 are less exposed to AI-generated text and modern data poisoning techniques, making physical publication catalogs attractive as a source for human-generated text."

Core: Systematic Teardown

Let me dissect this pipeline with the same cold logic I applied to Compound Finance's interest rate model back in 2020. This is a system with distinct failure modes.

The One-to-One Myth. The court's reasoning assumes that a digital copy can perfectly "replace" a physical one. In physical space, destroying the original ensures scarcity. In digital space, a copy is trivially reproducible. The moment the digital file exists, the constraint is broken. The only enforcement is legal — contract and burn certificates. From my years auditing DeFi protocols, I know that any system relying solely on off-chain enforcement for an on-chain (or in this case, digital) asset is a single point of failure. One rogue employee, one backup hard drive, one cloud sync — and the "one-to-one" constraint evaporates. The audit trail is only as strong as the trust in the destructor.

The Hidden Cost of OCR. The article mentions scanning but omits the post-processing pipeline. A million scanned pages produce hundreds of terabytes of raw images. Optical character recognition (OCR) is required to convert those images into machine-readable text. OCR introduces errors: misread characters, broken lines, lost formatting. Then cleaning: removing headers, footers, page numbers, indexing metadata. Then deduplication. Then quality filtering. Then labeling. Each step requires engineering hours and capital. The cost of this post-scanning pipeline likely exceeds the price of the books themselves. "s heart." The published numbers — millions spent on books — are only the tip of the cost iceberg.

Data Distribution Bias. The book supply is not random. ISBNdb allows filtering by ISBN, subject, publication year. What books are available for purchase in bulk? Remainders, overstocks, textbooks from a decade ago, self-published duds, library discards. These have a strong skew toward Western, English-language, and often outdated content. The model trained on this corpus will inherit those biases. It will know the 2018 edition of a statistics textbook but not the 2025 version. It will understand the historical context of the 1990s dot-com bubble but not the modern cryptocurrency landscape. "Empty metadata, full wallets." The data source is "clean" but not comprehensive.

The Scalability Ceiling. The number of physical books available for destructive purchase is finite. The total stock of retail remainders, library withdrawals, and used book inventories is a finite pool. Estimates suggest a few hundred million books globally. A single large language model training run might consume tens of billions of tokens. A typical book has 50,000–100,000 tokens. To reach even 10 billion tokens, you need 100,000 books. To compete at the frontier, you need millions. The supply cannot support multiple players doing this at scale. This is a race for a depletable resource. "Optimization is often obfuscation." The industry's clever legal workaround hides the fact that it is consuming a non-renewable physical asset.

Legal Exposure. The 2025 ruling is not final. It applies only to non-distributed digital library copies. Once that digital copy is used to train a language model and that model is deployed — serving thousands of users — does that constitute distribution? The court did not rule on that. Additionally, Anthropic still faces a pending claim for "pirated copies of central library copies" that was not dismissed. If that case decides that Anthropic's data includes illegally obtained content, the entire pipeline could be retroactively poisoned. The legal risk is not zero; it is unknown but material.

The Cultural Loss Blind Spot. The article from CryptoSlate notes that "no specific titles of rare, unique, near-extinct books have been named" in public records. But that is the problem: the ones that are rare are precisely the ones likely to be targeted by a premium destructive scanning service. A first edition of a seminal work, a signed copy, a book with unique marginalia — these are unique artifacts. The one-to-one logic ignores the difference between a generic trade paperback and a cultural artifact. "s heart." The law treats all printed paper as fungible. Cultural preservationists do not. The conflict is structural.

Contrarian: What the Bulls Got Right

To be fair, the reasoning behind destructive scanning is not purely cynical. The data purity argument stands: books are written by humans, edited by humans, and contain long-form arguments that are increasingly rare on the internet. In an era where synthetic text pollutes web crawls, physical books offer a statistically distinct distribution. For models that need to generate coherent, factual prose — legal documents, medical papers, historical analyses — this data could be genuinely superior. The court ruling provides a clear legal path, which reduces uncertainty for investors. Compared to the shadowy world of copyright lawsuit settlements (see: OpenAI and The New York Times), this model offers a clean audit trail: here is the purchase receipt, here is the destruction certificate, here are the scans. It is a legible supply chain. "Code is law until it isn't." Here, the physical destruction is the enforcement.

However, this advantage is temporary. The reputational cost is already surfacing. ISBNdb's own procurement articles acknowledge "headlines about AI companies destroying books" as a reputational issue. As public awareness grows, the cultural backlash will intensify. An AI company that markets itself as ethical while literally shredding books faces a contradiction that cannot be finessed.

Takeaway

The destructive scanning of physical books is a short-term data play that sacrifices long-term trust for immediate token purity. The parallel to the NFT metadata hollowing I documented in 2021 is exact: a clever legal or technical trick creates an illusion of scarcity and quality, while the underlying structure is fragile. "s heart." The AI industry will either abandon this practice under public pressure, or regulators will step in to protect cultural heritage. Either way, the one-to-one legal loophole will close. The question is not whether this model will survive, but how much irreplaceable knowledge will be turned into pulp before it does.