HN Debrief

Wellington second-hand bookstore's mysterious orders

  • AI
  • Copyright
  • Books
  • Public Infrastructure

The RNZ piece reports that Wellington used-book shop Book Haven noticed a stream of unusual overseas orders for niche New Zealand books, the same pattern other sellers have linked to AI training supply chains. The concern is that these books are being bought because they are hard to find elsewhere, then scanned in ways that often destroy the binding, with the resulting digital copies kept inside private model-training pipelines rather than released as readable archives.

If your business depends on archives, research, or rights, the practical issue is not book destruction by itself. It is valuable source material being digitized into closed systems while the law blocks public access to the scans.

Discussion mood

Mostly negative and uneasy. People were less upset about AI training in the abstract than about rare or obscure books being funneled into private corpora, with copyright law seen as the mechanism turning preservation into closed access.

Key insights

  1. 01

    Destruction may be tied to fair use

    Cutting off the spine is not just a cheap scanning trick. Several comments pointed to the emerging legal logic that destroying the purchased copy helps argue no extra market substitute was created, which could strengthen a fair use defense for scanning. That makes book destruction part of the incentive structure, not just an operational shortcut.

    Watch copyright doctrine, not only AI company behavior. If courts reward destructive scanning as cleaner fair use, preservation groups and rights holders will need a policy response fast.

      Attribution:
    • johntb86 #1
    • buildbot #1
    • jandrese #1
  2. 02

    Researchers lose a citable primary source

    For local memoirs, railway histories, war records, and similar niche books, the loss is not abstract cultural harm. It is the loss of a source a scholar can inspect and cite directly. A model output cannot stand in for a page number, an edition, an illustration, or exact wording, so closed scans effectively erase these works from serious research even if some of their content survives statistically inside an LLM.

    If you run research, legal, or knowledge products, treat source traceability as a hard requirement. A model trained on a document does not preserve the document for any workflow that depends on evidence.

      Attribution:
    • tmnvix #1 #2
    • stereolambda #1
  3. 03

    Legal deposit is not enough anymore

    National libraries and legal deposit systems do preserve some printed works, but comments from New Zealand and elsewhere stressed the gap between preservation and access. Copies may be held on site only, not loaned, and small print runs or self-published works can still slip through. That means the existence of a deposit copy does not solve the problem created when the market copies vanish and the scans stay private.

    Do not assume institutional archives have this covered. If a collection matters to your field, verify what is actually preserved, who can access it, and whether digital lending exists.

      Attribution:
    • kzrdude #1 #2
    • jiehong #1
    • tmnvix #1
  4. 04

    Private digitization is brittle preservation

    A physical book in a decent archive can outlast many digital storage stacks. Comments noted that plenty of 1990s digital material is already gone, which undercuts the idea that scanning into a private system automatically preserves knowledge for the long term. If the files are proprietary, inaccessible, or abandoned after a business failure, the public may end up with neither the book nor a usable scan.

    For preservation work, durability and governance matter as much as digitization. Prefer formats, custodians, and access rules that survive product shutdowns and media churn.

      Attribution:
    • croes #1
    • larodi #1

Against the grain

  1. 01

    A private scan may beat shelf obscurity

    A few comments argued that many of these books were already effectively inaccessible, sitting unread in scattered shops or private collections. From that view, getting their contents into models creates at least some chance of retrieval by future users, which may be better than near-total obscurity. The stronger version of this argument paired it with a call for much shorter copyright terms and a public archive so the same material would not stay locked up forever.

    Do not dismiss the access argument outright. If you oppose closed AI ingestion, pair that position with a realistic plan for public digitization, shorter rights windows, or digital lending.

      Attribution:
    • johnsmith1840 #1
    • rob74 #1
    • anal_reactor #1
  2. 02

    Bookshops should capture the digitization value

    One line of argument treated the bookstore as the wrong loser in this market. If obscure inventory now has training value, sellers could in theory scan and license it themselves rather than letting AI companies extract the upside. That immediately runs into copyright, which is exactly why some commenters said the commercial opportunity and the preservation problem are both symptoms of outdated rights rules.

    If you operate archives or niche content businesses, reassess whether your rights strategy matches new demand. The value chain around scanning and licensing is shifting faster than the law.

      Attribution:
    • nullbio #1 #2
    • rob74 #1

In plain english

AI
Artificial intelligence, software systems that perform tasks such as analyzing code or generating text.
copyright
A legal right that gives creators control over how their works are copied, distributed, and used for a period of time.
ephemera
Short-lived or minor printed material, such as pamphlets, guides, newsletters, or local publications, that often becomes historically valuable later.
fair use
A legal doctrine in United States copyright law that allows some uses of copyrighted works without permission under limited circumstances.
legal deposit
A law requiring publishers to provide copies of published works to a national library or archive.
LLM
Large Language Model, a machine learning model trained to generate and analyze human-like text.
provenance
Information about where a piece of text or data came from and who or what produced it.

Reference links

Copyright and fair use references

Book discovery and sales channels

  • Bokbörsen
    Given as an example of an online marketplace where many second-hand books are supplied by used bookstores.

Related New Zealand coverage

Wealth data references