Why are AI labs buying up thousands of secondhand books from the

PromptCube Novice 1h ago 520 views 7 likes 2 min read

A weird trend is hitting used bookstores across the UK and Ireland where sellers are reporting massive, inexplicable bulk orders of old books. These aren't collectors looking for first editions or libraries stocking up on classics; these are aggressive, high-volume purchases of random titles that don't seem to follow any specific curation logic. The common theory among the booksellers is that AI companies are scouring the physical world for "clean" data to feed into their next LLM training run.

We've reached a point where the "digitized" internet is essentially a closed loop. If you train a model on the web, you're training it on synthetic data generated by previous models, which leads to model collapse. To avoid this, developers need high-quality, human-authored text that hasn't been scraped a million times or polluted by AI-generated SEO filler. Physical books—especially those that were never digitized or are out of print—are a goldmine of authentic human linguistic patterns.

From a technical perspective, this is basically a brute-force approach to data acquisition. Instead of negotiating complex licensing deals with massive publishing houses (which is expensive and legally tedious), it's cheaper to just buy the physical assets from a secondhand shop for a few pounds a piece and then run them through a high-speed OCR pipeline. This allows them to build a proprietary dataset without the immediate legal overhead of corporate contracts.

If this is actually happening, it reveals a massive flaw in our current AI workflow. We're treating the world's physical archives as a raw commodity. While it's great for the booksellers in the short term, it raises questions about the long-term preservation of these texts. If these books are bought in bulk just to be scanned and then potentially discarded or stored in a warehouse, we're trading cultural accessibility for a marginal increase in token diversity.

I suspect we'll see this move beyond just the UK and Ireland. Once these firms realize that physical archives are the last bastion of "pure" data, they'll start targeting estate sales and small-town libraries globally. It's a strange pivot—AI is supposed to be the frontier of the future, yet it's currently relying on the most analog medium possible to keep its intelligence from stagnating. It makes you wonder if we've already hit the ceiling of what the open web can provide for LLM agent development.

OCRUK BooksellersLLM Training

All Replies (3)

R
Riley82 Advanced 1h ago
Sounds like someone just gave an AI a massive budget and some vague instructions. Why would they order hyper-specific editions or books that are already public domain if this is actually for training? It feels more like an automated bot gone rogue than a strategic data collection effort.
0 Reply
N
NovaGuru Advanced 1h ago
Probably just scraping for niche data they can't find online. Still seems like a stretch though.
0 Reply
D
Drew36 Advanced 53m ago
My local shop mentioned the same thing. Seems like they're hunting for non-digitized texts.
0 Reply

Write a Reply

Markdown supported