When Anna's Archive published a blog post this year, it landed on Hacker News and split the comment section in two. "This is a cultural crime" and "this is completely rational business operations" showed up almost simultaneously. The post drew 475 points and 822 comments — and at first it seemed strange that two camps looking at the same facts could arrive at such different conclusions.

The post's core claim was this: AI companies acquired rare books to build training data, and once the digital scans were done, they physically disposed of the originals. The "crime" reaction came from the standpoint of rare-book preservation; the "rational" reaction came from the standpoint of corporate operating logic. There's a reason both can be true at the same time. Following the story further shows something worth seeing: what AI companies actually took from those books, and what they threw away.

Rare-Book Collectors and AI Companies Want Different Things From the Same Book

Look closely at why rare books fetch high prices at auction, and a common pattern emerges: either the original print run was small, or the surviving copies have dwindled over time. In both cases, the text itself makes up only a small share of the book's value.

Take the Gutenberg Bible. The biblical text it contains is worth almost nothing on its own — that text is available anywhere. The value lies in the fact that it was printed with that specific metal type, on that specific paper, in the 1450s. The rare-book market trades, in a sense, in "status as an object." It buys existence, not content.

The AI training-data market is different. What's valuable there is the text. What era's ink was used, what paper, what binding method — none of that is a meaningful input for current AI training pipelines. The target is the vocabulary, the sentence structures, the period-specific language, the distinctive narrative style contained in the book.

In the end, the two markets are trading different objects while looking at the same book. One buys "the existence of this copy"; the other buys "the content of this copy." Until that distinction is clear, it's hard to understand why scan-then-destroy happens at all.

The Premise That Makes Destruction Logical — and Where It Cracks

Disposing of the original once the scan is complete is, within an AI company's economic logic, entirely consistent. The target — the digital file — has already been secured. What's left is a stack of paper that takes up storage space, costs money to transport, and can create legal complications when disposed of. In that situation, destruction is the simplest way to close the loop.

But for that logic to hold, one premise has to be true: that nothing of value is lost in the transfer from physical original to digitization.

Examine that premise, and the story gets more complicated. Library science and bibliography treat a book's physicality as carrying information distinct from its text. How worn a given page is, what a previous owner scrawled in the margins, which era's binding conventions it follows — none of this is text, but all of it is evidence of how that book was read and circulated. It's not the text, but it records the text's history.

Current scanning technology struggles to capture this information fully, and even when it is captured, it isn't systematically processed by AI pipelines focused on text training. If Anna's Archive's claims are accurate — and they haven't yet been independently verified by a third party — then AI companies extracted only the value they themselves defined, and discarded the rest. Whether that remainder held some other kind of value appears to have been outside the scope of their judgment.

Text That's Already Shared Doesn't Buy Competitive Advantage

This raises another question: why would AI companies bother acquiring physical rare books when digitized modern material is already so abundant?

There's a reason. Web pages, news articles, public papers, e-books — text that's already on the internet is material that most AI models have trained on in common. Models trained on the same material tend to hit similar ceilings in capability. Competitive advantage comes from data other models don't have.

Text that hasn't yet been digitized — out-of-print specialized works, unpublished material held by a specific academic institution, rare documents in a particular language — becomes, in this context, a training resource that can be monopolized. It's acquired through public auctions, direct purchases from collectors, or data partnerships struck with libraries and academic institutions.

In the partnership model, an institution signs an agreement allowing its holdings to be scanned for AI training purposes. What compensation the institution receives, whether it retains access to the originals after the deal, how long the scanned data remains exclusive to the AI company — these terms sit inside the contract and are rarely disclosed publicly. To the institution, it can look like a deal that simply cuts digitization costs, but it's not easy to know in advance exactly what rights are changing hands in the process.

The physical destruction of rare books that occurs along the way is, in a sense, the logical endpoint of this deal structure. If the goal is to monopolize the data, it makes sense to eliminate the possibility that someone else could scan the original again. Destruction is the surest way to prevent a re-scan.

Publishing Contracts Still Have No Clause for AI Training Use

This is where the problem lands squarely on publishers, authors, and content directors.

Traditional copyright frameworks are built around reproduction. Reprinting a book, translating it, adapting it into film or video — all of these require the author's consent and compensation. The framework rests on the premise that the author holds rights over "reusing the content."

AI training doesn't work the same way. In the process of converting text into model weights, the text itself doesn't survive intact in the final output. Whether this counts as "reproduction" is being decided differently by courts in different countries, and as of 2026, several ongoing lawsuits in the United States have yet to reach a conclusion.

Separate from the legal question, the economic structure is already shifting. A published work can now hold value in two markets simultaneously: one where readers consume it as content, and one where AI companies process it as a training resource. The two markets price the same work completely differently, and consume it in completely different ways.

In Korean publishing contracts, treating "whether AI training use is permitted" as its own separate clause is not yet standard practice. A contract without this clause can be read either way. If a dispute arises later, and one side claims "we never permitted it" while the other claims "it was never prohibited," the contract itself offers no basis for resolution.

When a publisher and author sit down to negotiate a contract, how to handle this clause is a matter of choice. A full ban, conditional permission, a separate compensation requirement — any of these is possible. But if no choice is made, the gap simply remains.

What this trail makes clear, and what it leaves open, are both worth stating plainly. What's clear: in the rare-book trade, the value AI companies want and the value preservation institutions are trying to protect never overlapped to begin with. Since the two sides wanted different things from the same book, whatever one side walked away with after the deal is exactly what the other side lost. Scan-then-destroy sits at the end of that structure.

What remains open: Anna's Archive's specific claims — which companies, at what scale, through what channels — have not been independently verified, and how existing copyright law will apply to AI training is still being decided in courts around the world.

And one question remains at the end of this trail: when you write the next contract, how will you word the clause on AI training use? The difference between including that clause now and leaving it out is the difference in how that contract can be interpreted later.