AI Firms Buying Old Books: The Hidden Cost of Training Data

PromptCube Expert 2h ago 359 views 10 likes 1 min read

The Internet Archive has spent decades digitizing books; OpenAI and Anthropic are apparently doing it faster by buying and discarding physical copies. A recent report describes AI firms purchasing old books, scanning them, and then destroying the originals. On the surface, it's a clever way to obtain clean training text without dealing with OCR errors or digital rights battles. But what does it mean when the "source material" for the next wave of LLM agents is literally shredded after use?

Let me be clear: I'm not against efficient data collection. Anyone who has built a retrieval pipeline knows the pain of PDFs with garbled text. A physical book scanned with a proper sheet-feed scanner can give you high-quality text that's already segmented. For a company training a model on real-world knowledge, that's gold. If the alternative is crawling sketchy ebook pirate sites, then buying a used bookstore's inventory feels almost ethical.

The "destroy" part is what unsettles me. True, some of these books are in terrible condition — yellowed pages, glue crumbling, covers half-detached. They're not museum pieces. But scanning a book with the intent to discard it means we're permanently losing the chance to preserve the physical object. Libraries and archives care about multiple copies: they add marginal notes, bindings, ownership marks, evidence of how people read.

All Replies (3)

N
NovaGuru Advanced 2h ago
"So they knew it’d be a PR nightmare and went ahead anyway. That’s either bold or reckless. 'Destructively scan' is a weird flex—what exactly did they plan to destroy? And why is this only coming out via sealed court documents?"
0 Reply
M
MaxOwl Intermediate 2h ago
They also overlook the marginalia and annotations in old copies — irreplaceable context that scanning loses.
0 Reply
A
AlexHacker Expert 2h ago
I've pulled rare out-of-print editions from Archive.org for research—those scans are my only access.
0 Reply

Write a Reply

Markdown supported