Every AI vendor trains their chatbots on the Common Crawl, a copy of the whole public internet. But that’s just the starting point. AI vendors are desperate for more data — in the hope of making their chatbot suck a bit less.
Last year, we covered how Anthropic was buying physical books, scanning them, and destroying them to feed their Claude chatbot. Judge William Alsup ruled this fair use. In particular, he said mulching the books was perfectly fine because making them digital was “exceedingly transformative”: [Order on Fair Use, 2025, PDF]
The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company.
The great plan here was literally spelled out in the words of Anthropic executives: “Project Panama is our effort to destructively scan all the books in the world.” But also, “We don’t want it to be known that we are working on this.” [Washington Post, archive]
Don’t want it to be known? Wonder why?
Anthropic also downloaded pirate book libraries, which the court found not to be fair use. This is where Anthropic lost part of the case and why they had to pay out the authors $1.5 billion.
But the lawsuit didn’t stop Anthropic. To the contrary, it gave them license to keep going. They’re still buying rare books. As rare as possible. And feeding them into the book mulcher.
We tried tracking this down a few months ago. There’s one book buyer called Zoom Books in Canada and Nevada who are buying up such large volumes that booksellers around the world were left scratching their heads. The booksellers report Zoom buying up dusty piles of old books nobody wanted and they weren’t able to move — and paying good money for them, too. [Amazon Seller Central; Amazon Seller Central; blog post]
We tried and failed to quite find a smoking gun linking Zoom Books directly to Anthropic. But Zoom did have a blog — with AI-related posts. These posts have all disappeared, but there’s an Internet Archive link to one post called “AI Models Can Reproduce Books Verbatim — What New Research Means for Training Data Compliance,” which nobody has a surviving copy of. Must have been mulched. [Zoom, archive]
Zoom denies they’re doing the actual scanning and destruction of books: “To be unequivocally clear: Zoom Books does not digitize or destroy used or new books for the purpose of training AI models, nor for any other purpose.” [Publishers Lunch]
Fine, but what Zoom’s denying is not the claim. The claim is that it’s Anthropic doing the scanning and shredding — and that Zoom is buying the books on their behalf.
Finally, Emanuel Maiberg from 404 Media has nailed down the story. He got a great interview from a bookseller who spilled the beans. Maiberg’s story, “AI Companies Are Buying Tons of Old Books Because They’re Free of AI Slop,” doesn’t specifically name Zoom Books, but he’s identified other bulk book buyers. [404, archive]
The star of the story is ISBNdb, a book catalogue platform, indexing the ISBN (international standard book number) printed on the back cover of almost every book since 1970. But ISBNdb is now selling itself hard to AI companies. It’s also got scare stories that authors — who are apparently the AI vendors’ greatest enemies — are poisoning their books to mess with the models! Oh no! [ISBNdb, archive]
So ISBNdb is telling AI companies they should train on print books from before 2022: “This pre-2022 inventory represents a vast, provably clean corpus that exists nowhere else in quite the same form.”
ISBNdb also reassures its AI customers that it can totally keep their names secret.
To be clear, ISBNdb is not buying the books itself. Various bulk book buyers order the books online. But ISBNdb is a vital part of the chain.
According to 404’s Mailberg:
One bookseller told me that similarly large orders of books were coming through another marketplace called Biblio. Customers can provide Biblio with a spreadsheet of ISBNs they want to purchase and the company takes it from there.
ISBNdb loudly advocates strip-mining the world’s old books, scanning them for AI training, and then pulping them — and leaving the only remainder of the original text as some weights in a large language model. And all this for an unnamed large AI company that’s going broke in a year.
ISBNdb understands the key issue here — it’s bad public relations: [ISBNdb, archive]
“AI company destroys two million books” is not a headline that generates sympathy.
We wonder why it doesn’t generate sympathy. You, foolish pleb, might think preserving old knowledge and not strip-mining it might be some sort of public good.
But it’s fine, because, according to ISBNdb:
A physical book is a delivery mechanism for information. Once that information has been extracted and encoded into an AI model, the delivery mechanism has served its purpose.
Served its purpose. You heard them: the purpose of a book is to train AI.
Also, ISBNdb gets paid. And that’s what counts.









