HN Debrief

AI companies are shredding rare books

  • AI
  • Copyright
  • Libraries
  • Culture
  • Regulation

The post points to reporting that AI companies and their vendors are buying large quantities of older and low-circulation books, cutting off the spines, sheet-feeding the pages through scanners, and keeping the scans for model training. The alarmed framing was that this could destroy rare books for private AI corpora. The comments landed on a narrower and more credible version of that claim. Destructive scanning is real and common because it is fast, cheap, and produces clean scans. What is much less supported is the idea that labs are routinely pulping irreplaceable 18th-century works or singular historical artifacts. Several people dug up the cited examples and found recent ISBN-bearing books from the late 1960s onward, often niche or foreign-language titles, not medieval manuscripts or obvious museum pieces.

If you care about preservation, push past the outrage headline and focus on policy. The practical gap is simple: current law and market incentives favor private ingestion over public archiving, so expect more cultural material to be scanned into closed systems unless libraries, archives, or regulators force a preservation copy and public access path.

Discussion mood

Angry and distrustful. Most commenters disliked the optics and secrecy of shredding books for private AI training, but the dominant frustration was aimed at copyright law, DRM rules, and the lack of a public preservation requirement more than at the mechanics of scanning themselves.

Key insights

  1. 01

    The cited books were mostly modern ISBN titles

    The strongest factual correction was that the concrete examples being circulated were recent, niche, ISBN-tagged books, not ancient artifacts. That matters because it punctures the viral image of frontier labs feeding unique 18th-century volumes into shredders. The actual risk is narrower and more mundane. Low-circulation technical, academic, and foreign-language books can still be hard to replace, but the evidence in hand points to obscure modern print runs rather than museum-grade heritage objects.

    Treat future claims in this area as two separate questions. Ask first whether destructive scanning is happening, then ask what kinds of books are actually being sacrificed before you generalize from “obscure” to “irreplaceable.”

      Attribution:
    • dpark #1 #2
    • Cider9986 #1
    • tokai #1
  2. 02

    Copyright and DRM make the dumb workflow legal

    The perverse part is not that scanners require destruction. It is that the cleanest legal path for in-copyright material is often to buy print copies and convert them by destroying the original, instead of cracking ebooks protected by the Digital Millennium Copyright Act or trying Internet Archive style lending models that courts have already rejected. Non-destructive scanners exist. They just do not fit the lowest-risk, lowest-friction legal pipeline for bulk ingestion.

    If your company handles digitization, preservation, or training data, do not assume the ugliest workflow is driven by technology. Audit where law and licensing are pushing you into worse operational choices, because that is where reform or partnership work can actually change behavior.

      Attribution:
    • TeMPOraL #1 #2
    • tourmalinetaco #1
    • voakbasda #1
  3. 03

    Public release is the missing preservation step

    Several comments converged on a practical compromise. If companies want to scan books at scale, especially public-domain or low-circulation works, they should have to deposit the resulting scans in a public archive or release them after a delay. That would turn destructive scanning from private extraction into actual preservation. One useful detail added here is that the US already has mandatory deposit for copyrighted works, which shows the legal system already accepts some preservation obligations in principle.

    When this issue comes up in policy or partnerships, ask for deposit and access requirements, not just bans. A delayed-release archive, library deposit, or escrowed preservation copy is more actionable than a broad moral objection to scanning.

      Attribution:
    • JumpCrisscross #1
    • s1artibartfast #1
    • Legend2440 #1
    • qingcharles #1
  4. 04

    Out-of-print scholarship really does get trapped

    The best case against the casual “nobody needs these books” line came from people working in specialized fields. They pointed out that a great deal of twentieth-century scholarship is still under copyright, never republished, and still actively useful in domains like linguistics, history, and archaeology. In those cases the missing issue is not sentimental attachment to old books. It is real access to working knowledge that lives in old monographs, field reports, and regional publications.

    If you rely on specialized literature, build your own preservation and access plan instead of assuming the market or publishers will do it. The vulnerable category is not blockbuster fiction. It is narrow, older scholarship that remains useful but commercially dead.

      Attribution:
    • TFNA #1
    • jolmg #1
  5. 05

    Book durability nostalgia skips the economics

    Comments from people familiar with binding and libraries cut through a recurring myth in the thread. Many older books that survive were luxury objects or later custom bindings, and the reason they are visible now is partly that the flimsy ones already disintegrated. Modern cheap bindings are not proof that publishers forgot how to make durable books. They reflect mass-market demand and cost tradeoffs. Preservation has usually depended on libraries, rebinding shops, and archival processes rather than every retail copy being built to last centuries.

    Do not turn preservation policy into a fight about making all books premium physical objects. The scalable levers are archival copies, library-grade rebinding where it matters, and better digital preservation standards.

      Attribution:
    • kasey_junk #1
    • tookmund #1
    • ctolsen #1

Against the grain

  1. 01

    Digitization can still beat quiet decay

    A minority view held that even destructive scanning may be a net gain when the alternative is obscurity, weeding, or landfill. For books that are commercially dead and barely accessible, a decent scan at least keeps the text alive somewhere, and many neglected works were never going to receive curatorial treatment from libraries or publishers. This does not justify closed corpora, but it does challenge the idea that the untouched physical copy is automatically the best preservation outcome.

    When evaluating preservation tradeoffs, compare against the real baseline, not an imagined museum future. Some low-value print artifacts will disappear either way, so the key question is whether scanning creates durable public access or merely private leverage.

      Attribution:
    • dpedu #1
    • DarkIye #1
  2. 02

    Internet Archive was not a clean precedent

    Some commenters pushed back on claims that publishers crushed a reasonable digital-lending model and caused this mess. They noted that the Internet Archive case was legally weak even before the pandemic-era unlimited lending move, and courts had already signaled hostility to digital first-sale theories through cases like ReDigi. That does not make the outcome good, but it means this was not a stable public-rights regime that AI companies later distorted.

    If you are looking for legal models to support digital access, do not assume controlled digital lending has durable footing under current US law. Preservation strategies that depend on courts blessing digital first sale remain fragile.

      Attribution:
    • ndiddy #1
    • kmeisthax #1
    • kingstnap #1
  3. 03

    Some physical copies are less valuable than people admit

    A few commenters argued that for many post-industrial print books, the content is the thing and the artifact itself carries little standalone historical value. On that view, destroying one copy of an obscure 1970s children's test run or a niche technical book is not comparable to destroying a manuscript, a signed limited edition, or a culturally singular object. The point is not that all shredding is fine. It is that preservation priorities should be ranked, not treated as absolute.

    Build a tiered preservation policy instead of a blanket rule. The case for intervention is strongest when scarcity, provenance, or physical features create value beyond the text itself.

      Attribution:
    • stuartjohnson12 #1 #2 #3

In plain english

DRM
Digital rights management, technology used to restrict copying or access to digital files like ebooks, music, or video.
Format-shifting
Converting content from one medium to another, such as scanning a printed book into a digital file.
ISBN
International Standard Book Number, a standardized identifier assigned to published books and editions.
Mandatory deposit
A legal requirement that publishers provide copies of certain works to a national library or copyright office for preservation and recordkeeping.
OCR
Optical character recognition, software that turns scanned images of text into machine-readable text.
ReDigi
A US court case about resale of digital music files that is often cited to show how weak digital first-sale rights are under current copyright law.

Reference links

Reporting on book scanning and AI

Copyright and legal background

Preservation and scanning resources

Books and fiction referenced in discussion