HN Debrief

AI companies destroy physical books – let's scan rare books before it's too late

  • AI
  • Books
  • Preservation

Anna’s Archive published a short appeal centered on a simple claim: rare physical books are being cut apart and fed through high-speed scanners to create digital copies and AI training data, and scarce works should be digitized carefully before the remaining copies are damaged or lost. The post frames this as a preservation emergency rather than a normal library digitization effort. It is about the physical fate of hard-to-replace books, not just about who gets to read the resulting files.

If you care about archives, training data, or niche technical literature, the practical signal is to watch the earlier discussion and treat physical preservation as a separate problem from copyright or scraping policy. The immediate issue is not just access to text, but whether one remaining copy survives the scanning process at all.

Discussion mood

No real sentiment is visible here because the comments were moved to another submission before any substantive exchange happened.

In plain english

AI training data
The text, images, audio, or other material used to train an artificial intelligence model.
Digitization
Converting physical material such as books or documents into digital files.

Reference links

Related discussion