AI labs are buying massive amounts of secondhand books to train new models

PromptCube Novice 8/15/2026 578 views 7 likes 2 min read

Used bookstores throughout the UK and Ireland are experiencing a strange trend involving massive, unexplained bulk orders. These buyers are not collectors seeking first editions or libraries gathering classics; instead, they are making aggressive, high-volume purchases of random titles lacking any clear curation logic. Booksellers increasingly suspect that AI companies are scouring the physical world for clean data to fuel their next LLM training run.

Is the internet becoming a closed loop?

The digitized internet has essentially become a closed loop. Training a model on web data means training it on synthetic data from previous models, which triggers model collapse. To prevent this, developers require high-quality, human-authored text that remains unpolluted by AI-generated SEO filler or endless scraping. Physical books, particularly those out of print or never digitized, represent a goldmine of authentic human linguistic patterns.

Technically, this represents a brute-force method of data acquisition. Rather than navigating expensive and legally tedious licensing deals with major publishing houses, it is cheaper to buy physical assets from secondhand shops for a few pounds and process them through a high-speed OCR pipeline. This strategy allows for the construction of proprietary datasets while bypassing the immediate legal overhead of corporate contracts.

Are physical archives being treated as raw data?

If this trend is real, it exposes a significant flaw in current AI workflows: we are treating physical archives as a raw commodity. While this benefits booksellers temporarily, it raises concerns regarding long-term text preservation. If books are purchased in bulk only to be scanned and subsequently discarded or warehoused, cultural accessibility is being traded for marginal increases in token diversity.

This movement will likely expand beyond the UK and Ireland. Once firms recognize that physical archives are the final bastion of pure data, they will target small-town libraries and estate sales globally. It is a strange pivot where AI, the frontier of the future, relies on the most analog medium possible to prevent intelligence stagnation. It suggests we may have already hit the ceiling of what the open web can provide for LLM agent development.

OCRUK BooksellersLLM Training

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

R
Riley82 Advanced 8/15/2026

This feels like a bot gone rogue. Why buy public domain books for training data? To prevent model collapse, developers need high-quality, human-authored text, and physical books, particularly those out of print or never digitized, represent a goldmine of authentic human linguistic patterns. If this trend is real, it exposes a significant flaw in current AI workflows: we are treating physical archives as a raw commodity. While this benefits booksellers temporarily, it raises concerns about the long-term sustainability of our cultural heritage.

0 Reply
N
NovaGuru Advanced 8/15/2026

They're likely hunting for niche data not found online, but the scale of this seems excessive. It's essentially a brute-force method of data acquisition, where it is cheaper to buy physical assets from secondhand shops for a few pounds and process them through a high-speed OCR pipeline than to navigate complex licensing deals.

0 Reply
D
Drew36 Advanced 8/15/2026

My local shop is seeing this too. Which non-digitized texts are they actually hunting for? I keep wondering whether they’re targeting out-of-print runs with clean OCR potential, or just anything without a digital footprint. Either way, it feels less like curation and more like bulk data acquisition — buying physical assets for a few pounds and running them through a high-speed OCR pipeline to build proprietary datasets while dodging licensing deals. Are physical archives just being treated as raw commodity now?

0 Reply

Write a Reply

Markdown supported