Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.

ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.

“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”

  • mirshafie@europe.pub
    link
    fedilink
    English
    arrow-up
    25
    ·
    13 hours ago

    If they did have a copy in Anna’s Archive, why the fuck would AI companies find physical copies, buy them, rip them apart, scan them, and then digitize them before destroying them?

    Remember, you don’t get the digital copy. It’s only used to train their models.

    • Jason2357@lemmy.ca
      link
      fedilink
      English
      arrow-up
      4
      ·
      5 hours ago

      Absolutely. Every AI company has Anna’s Archive in their training data already. They want to differentiate their models by having more training data than their competitors. There are way more books out there that have never been scanned than those that have.

      • Kairos@lemmy.today
        link
        fedilink
        English
        arrow-up
        2
        ·
        3 hours ago

        OpenAI just settled out of court for $1.5 Billion perhaps some see this as a way to avoid that legal exposure. But IDK.