ISBNdb, once a back office tool for libraries and bookshops, now brokers 1,000 to 1,000,000 title orders to AI companies, marketing pre 2022 print copies as free of LLM contamination and offering discretion.
ISBNdb, the self-described "world's largest book database," is now pitching itself to AI labs as a bulk broker of used physical books, per 404 Media: 1,000 to 1,000,000 titles per transaction, marketed as "structurally clean" of LLM-era contamination, with the buyer's name kept off the paperwork.
The pitch works because of a documented precedent that landed the same week. A federal judge in San Francisco entered the [Final Approval Order in Bartz v. Anthropic](https://www.courtlistener.com/docket/69058235/bartz-v-anthropic-pbc/) on July 20, 2026, closing a class action over how Anthropic built its training set. The court accepted Anthropic's account: the company bought used print books, fed them through a hydraulic cutting machine to separate the pages, then ran those pages through industrial imaging rigs to produce training data. Anthropic paid $1.5 billion to settle; it did not admit infringement. The physical books did not survive the process.
The legal stack that made the settlement possible is the same stack ISBNdb is now selling against. Under the first-sale doctrine, a buyer of a lawful copy can do what they want with that copy: resell it, lend it, cut it up. The digitization of a lawfully acquired physical book, the court accepted, can be transformative fair use: a new purpose (training a language model) that does not substitute for the original market. Combine the two, and a lab can lawfully acquire a book, destroy it, and ship the resulting training signal into a model without paying the author or the publisher a licensing fee.
That is the entire ISBNdb pitch in two sentences. "Structurally clean" means something specific in broker copy: printed before 2022, the books predate the public release of capable LLMs, so their text was written without AI assistance and is uncontaminated by it. To a lab worried that a corpus is silently recycling its own outputs back into training, that property is worth a premium. The destruction is the feature. Once the scan is done, the physical object is gone, and with it the only thing the rightsholder could point to as a unique, identifiable copy.
The supply chain is industrializing. A single ISBNdb transaction of one million titles is not a researcher's haul; it is the input for a pre-training run. The broker sits between the used-book market and the GPU cluster, takes a margin, and absorbs the embarrassment of physically handling the materials. Confidentiality, offered in the same package, lets the buyer keep the order out of the public record, a meaningful offer in a market where the only comparable case to date has already produced a $1.5 billion settlement.
Meta, by the parallel route, has been accused of using pirated digital books to train its models. The legal theory is different: no first-sale defense applies to a copy the lab never lawfully acquired. The end state is the same: a large language model trained on copyrighted text, with no royalty paid on the input side. The ISBNdb market is the cleaner version of the same problem, sold openly to the customers who want to avoid the messier version.
The $1.5 billion Anthropic settlement closed a lawsuit. It did not close the market. As long as a lab can buy a book, scan it, destroy it, and call the result "transformative," the supply of pre-2022 print copies is a finite, depleting resource being converted into model weights one hydraulic cycle at a time. The structural questions are still open: whether first-sale should extend to a copy whose only use is ingestion, whether "transformative" should survive the destruction of the original, and whether a clean licensing market for pre-LLM text is reachable before the used-book supply runs out. The brokers are not waiting for the answers.