ISBNdb, a book records database that now also brokers physical bulk copies to AI developers under strict NDAs, runs a pipeline built to acquire large volumes cheaply.
High-speed scanners at AI labs cut the bindings off physical books, separate the pages, and photograph both sides at speed. The digital text becomes training data for a chatbot. The original paper volume goes to a bin. The books arrive in anonymous bulk shipments, under non-disclosure agreements, from ISBNdb, a service that started as a database of book records and now also brokers physical bulk copies to AI developers. The 404 Media investigation that surfaced the pipeline in July describes the practice in operational detail; the Aio Books case reported by Fortune on 2026-07-31, in which a Dutch bookseller was offered payment to scan and destroy approximately 3,000 of its stock, shows the approach in the wild.
The books the pipeline targets are not new. ISBNdb's customers want books published before roughly 2022. The reason is a data-quality problem the industry calls "model collapse." AI-generated text became common in training corpora after 2021, so books written before then are valued for being untouched by chatbot prose. 404 Media's reporting describes the books ISBNdb sources as "dense, edited, authoritative" and free of LLM-generated content. The pattern is broader than ISBNdb: 404 Media has separately reported that AI companies are buying up old books precisely because they are free of AI-generated text. Model collapse is the documented quality loss that happens when an AI trains on text written by other AIs. Buying older books is one way to step around it. Licensed scans from libraries such as HathiTrust and the Internet Archive already produce training-grade text; the spine-cutting pipeline exists because it is faster and cheaper, not because it is the only path.
The supply chain is sealed by contract. The 404 Media investigation found that ISBNdb's agreements with AI developers include non-disclosure agreements, NDAs, that prevent either side from naming the other. ISBNdb has acknowledged to 404 Media that the secrecy is reputational, not commercial. The result is a market for physical books that no one outside the immediate transaction can observe, and that libraries, authors, and the publishers who still hold rights cannot audit. TechTimes reported on 2026-08-01 that ISBNdb issued a policy reversal under public pressure that did not change the underlying practice.
The rare-book loss is a structural side effect, not the goal. 404 Media's reporting says the bulk pipeline is built to acquire large volumes cheaply, not to target rare holdings. Some books entering the pipeline are reportedly extremely rare, including volumes that have survived institutional and private collections for decades. The 404 Media investigation does not name specific rare titles, and the reporting that has surfaced does not document a specific destroyed first edition. The risk is the structure itself: a supply chain that takes in physical books, destroys them in the scanning process, and operates under secrecy has no mechanism to keep rare volumes out. A volume that survived a war or a private collection for two centuries can end up in a paper-shredder bin because no one in the pipeline is required to look.
The legal question splits in two. Bartz v. Anthropic PBC, 4:24-cv-05417, is the live US case testing whether training on copyrighted books without permission is fair use. The ISBNdb pipeline raises a different question: the books are being bought and physically destroyed, so the issue is not whether training on them is fair, but whether the market for bulk physical copies is being used to bypass licensing. Anthropic's motion to dismiss is pending in the same court.
The destruction is a chosen path, not a technological one. The Internet Archive's controlled digital lending and HathiTrust's library partnership scans are working models of how to digitize books without destroying them. Provenance-tracked corpora with attribution to authors and rights-holders are buildable today. The reason the industry is using spine-cutting scanners on purchased physical copies is that the supply chain is fast, the books are cheap, and the NDAs are tight. None of those conditions is a property of the technology. The industry could choose licensed scans with attribution and provenance, and pay more per book to do it. 404 Media's reporting suggests the labs have not, because they did not have to.
The scanner does not check whether the binding is rare before it cuts.