HN Debrief

AI companies destroy physical books – let's scan rare books before it's too late

  • AI
  • Copyright
  • Libraries
  • Regulation
  • Public Access

Anna’s Archive argued that AI companies are buying secondhand books, disbinding and scanning them, then keeping the scans private for training while the physical copies disappear. The post framed this as a race to digitize obscure print before it gets locked inside corporate training pipelines. That matters because many older books, especially from the pre-ebook era, were never digitized at all, and some are out of print with murky or effectively abandoned rights. The immediate trigger was the recent Anthropic litigation, where a court treated scanning purchased print books and destroying the originals as a more defensible kind of format shift than building a permanent library from pirated ebooks.

If you care about preserving niche or out-of-print knowledge, the practical bottleneck is copyright and access, not scanning tech. Watch for policy fights around format-shifting, orphan works, legal deposit, and whether AI-era digitization ever gets routed into public archives instead of private datasets.

Discussion mood

Concerned and frustrated. The dominant mood was that private scans locked inside AI companies are a bad outcome, but the thread was also skeptical of the most apocalyptic framing and annoyed by claims that implied AI labs were literally buying and destroying the last copy of major works without evidence.

Key insights

  1. 01

    Google Books already showed the policy failure

    Google Books is the clearest proof that mass book digitization was technically feasible years before the AI boom, and that the real blocker is rights, not scanning. The lost opportunity was not just Google’s limited previews. It was the collapse of a proposed settlement that could have made out-of-print and orphan works searchable and purchasable at scale. The result was the worst version of the future. Google scanned a vast corpus, nobody built a competing legal system for access, and the public still cannot read most of it.

    If you want these books available, focus less on scanner drama and more on orphan-work and out-of-print licensing rules. Any preservation strategy that ignores rights clearance will recreate Google Books: huge corpus, tiny public benefit.

      Attribution:
    • ndiddy #1
    • allturtles #1 #2
    • cube00 #1
  2. 02

    Internet Archive lost on the exact theory people wanted

    A lot of people still talk as if libraries should be able to digitize a print book they own and lend one digital copy at a time. The problem is that this specific controlled digital lending theory already got crushed in court. The publishers did not just object to the emergency library stunt. The ruling rejected the idea that owning print entitles a library to make and lend a replacement ebook without a publisher license. That makes the legal path far narrower than preservation advocates want to admit.

    Do not assume a library can legally solve this by scanning its shelves and lending digital substitutes. If your plan depends on controlled digital lending in the US, it currently rests on losing law.

      Attribution:
    • ndiddy #1 #2
    • allturtles #1
    • scarmig #1
  3. 03

    Destruction strengthened the format-shifting argument

    The books were not destroyed only because feed scanners are faster. Destruction also made the legal story cleaner. The Anthropic ruling leaned on the idea that a purchased print copy had been converted into a digital library copy for storage and searchability, with the source copy gone and no extra copy distributed outside the company. That is a very different fact pattern from borrowing library books and returning them, which is why Google’s workflow does not map neatly onto AI training pipelines.

    Expect companies to keep choosing the workflow that gives them the strongest fair-use posture, not the one archivists prefer. If policymakers want better preservation behavior, they need to change the incentives, not ask firms to volunteer for weaker legal defenses.

      Attribution:
    • shagie #1
    • jujube3 #1
    • thread_id #1
  4. 04

    A scan for training is not preservation

    Archivist-minded comments drew the sharpest line in the whole discussion. What gets lost is not just paper. It is typography, artwork, layout, metadata, provenance, citation quality, and sometimes even foldouts or physical structure. More importantly, ingesting a book into model weights is a lossy transformation. Even keeping a private PDF is not enough if no one outside the company can inspect or verify it. Preservation means a recoverable, citable source, not a statistical echo inside an LLM.

    When vendors talk about "capturing knowledge," ask whether they are preserving inspectable source materials or only extracting training value. For historical, legal, academic, or technical use, that distinction is everything.

      Attribution:
    • chefandy #1 #2
    • nuancebydefault #1
  5. 05

    LLMs are bad substitutes for exact text

    Several comments cut through the fuzzy "knowledge is preserved" claim with a simple test: can you retrieve the exact passage you need. In practice, no. Models refuse verbatim copyrighted output, hallucinate quotes, paraphrase badly, and blur provenance. That makes them poor substitutes for books in any use case that depends on exact language, like scholarship, legal citation, translation, literary study, or just finding one sentence correctly.

    Treat model access as derived access, not source access. If your work depends on exact wording, quotations, or trustworthy citations, you still need the actual scan or text corpus.

      Attribution:
    • radu_floricica #1
    • jjulius #1
    • SkyBelow #1
    • zahlman #1
    • iLemming #1
  6. 06

    The long tail is where real loss happens

    The strongest pro-preservation argument was not about famous first editions. It was about obscure academic monographs, conference proceedings, local histories, niche manuals, and low-print-run works that are expensive, hard to find, and often absent from the web. Those books are precisely the kind of thing that can be unimportant commercially and still irreplaceable to historians, researchers, or specialist practitioners later. Popularity is a terrible filter for future value.

    If you care about preservation, prioritize the long tail over canonical titles. The books most likely to vanish are not the famous ones. They are the ones with tiny audiences and no current business case.

      Attribution:
    • nloomans #1
    • TFNA #1
    • afpx #1
    • arttaboi #1
  7. 07

    Legal deposit exists but does not solve access

    A lot of people reached for the Library of Congress as the obvious answer, and the thread clarified why that is only a partial fix. Mandatory deposit rules already exist in the US and many other countries, which means some published books should already have archival copies somewhere. But deposit does not mean good digital access, complete coverage, easy rights status, or public availability. It also does little for foreign works, orphan works, poor scans, and items that never got digitized in the first place.

    Do not confuse national-archive custody with usable public preservation. A collection can be legally deposited and still be practically inaccessible for decades.

      Attribution:
    • brookst #1
    • shagie #1 #2 #3

Against the grain

  1. 01

    AI is a tiny part of book destruction

    The sharpest pushback was that AI companies are being singled out for something the book ecosystem already does at much larger scale. Libraries weed collections. Donations get dumped. Used-book sellers pulp dead inventory. Against that background, scanning one copy before disposal can look better, not worse, than the normal lifecycle of unwanted books. This does not make private archives good, but it does puncture the idea that AI created book destruction as a new phenomenon.

    If this issue matters to you, broaden the lens beyond AI. Preservation efforts that ignore library deaccessioning, donation waste, and routine pulping will miss the majority of actual losses.

      Attribution:
    • tptacek #1 #2 #3
  2. 02

    Evidence for last-copy loss is still weak

    A credible skeptical line ran through the thread: there is plenty of anxiety but very little hard evidence of unique or culturally major works actually being wiped out by AI scanning. Reported examples were often hypothetical, conditional, or vague about titles. Without concrete ISBNs, editions, or known last-copy cases, some readers saw the story as opportunistic rhetoric attached to a real but unquantified practice.

    Demand inventories and examples before treating this as a proven extinction event. Better transparency from booksellers, scanners, and reporters would separate real preservation failures from fear-driven speculation.

      Attribution:
    • quietsegfault #1 #2
    • joshstrange #1
    • demibabs #1
  3. 03

    Selective nondestructive scanning is feasible

    One practical counterpoint to the broader outrage was that the problem does not require stopping bulk scanning. It requires triage. Commodity books can be destructively scanned. Truly scarce or high-value editions can be flagged for slower handling. Price, catalog scarcity, and library holdings are imperfect but usable signals. The failure is not digitization itself. It is running an indiscriminate industrial process with no preservation gate.

    A workable policy is to force or incentivize rarity checks, not ban scanning. If you run a digitization pipeline, build a hold-and-review path for books with low holdings or unusual edition data.

      Attribution:
    • NishanStepak #1
    • qarl2 #1
    • hakanensari #1

In plain english

controlled digital lending
A library theory that a library can digitize a print book it owns and lend one digital copy at a time in place of the physical copy.
Google Books
Google’s large-scale book scanning project that made many books searchable online, often with limited previews.
Internet Archive
A nonprofit digital library that preserves websites, books, media, and other digital materials.
Library of Congress
The national library of the United States, which also receives many published works through legal deposit rules.
LLM
Large Language Model, a machine learning model trained to generate and analyze human-like text.
orphan works
Works that are still under copyright but whose current rights holder cannot be identified or located.

Reference links

Legal cases and copyright background

Preservation and digitization projects

Book access and donation tools

Commentary and reporting on the controversy

Copyright policy and legal deposit references