- cross-posted to:
- technology@beehaw.org
- cross-posted to:
- technology@beehaw.org
Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.
ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.
“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”


What’s the point of destroying books after scanning them instead of reselling them apart from *hurr durr we are evil*?
you snip off the book and scan the binding
Easier to scan if you cut the binding off, and it costs money to re-bind a book you purchased for 10 cents as part of a lot.
Have you ever tried scanning a book? It’s a huge PITA.
They just cut the binding off so they can be run through a document feeder.
Do make sure that knowledge is only available through their model and otherwise lost to humanity.
Same with bombarding small websites until they give up and pull the plug.
Everything is fucked
Cause to easily scan them they remove the binding and rebinding is more work %han destroyibg so the point is capitalism
Copyright. They mustn’t make a copy. There is legal precedent that confirms it’s okay to transform the copy you bought into digital format. The Internet Archive relies a lot on that. I think they actually litigated it in the first place. So that’s why the copyright heads are going so absolutely apeshit. If no one’s charging you rent for using some data, then it’s “unethical”.
The Internet Archive absolutely does not destroy books. They scan them the hard way with the binding still intact.
I was misremembering. Google won the precedent, when they were suing over Google Books. IA had only 1 big lawsuit and were forced to settle. The non-destructive scanning is dicey.
I mean have you tried donating books? They just dump them into a storage room with thousands of other books donated that nobody wants.
To keep their competitors from buying them and scanning them.
That seems to be the most likely reason. Bastards.
Sorry, but I did an oopsie. I watched the video linked, they destroy the books to scan them. It’s really sad. Though I’m sure they would destroy them anyway.
I assumed the process destroyed the book but keeping the contents from the competitors would definitely be seen as a business goal.
Fair use loophole, they claim they convert one physical book into one digital with no illicit copies, while also supporting the delusion of AI being an equivalent to a human reader. It is weird on so many levels but no one of noticeable weight asked them wtf.
That is not required by law. It is just easier to cut the bindings off. The other way is slow: https://archive.org/details/eliza-digitizing-book_202107
They’re pissing in the well so their competitors can’t access the same information.
To avoid competitors also scanning the book to get some kind of edge.