- AI labs are buying pre-2022 physical books in lots of 1,000 to 1 million, scanning them, and in many cases destroying the originals, because print published before the generative-AI era is guaranteed free of machine-written text.
- The driver is model collapse, a documented degradation that sets in when models train on the synthetic output of earlier models, with each generation losing quality and diversity.
- A June 2025 US ruling in Bartz v. Anthropic held that buying physical books, digitizing them, and training on the copies can qualify as fair use, clearing the legal path for the practice to scale.
Labs are paying for paper the internet cannot supply
The most valuable input in artificial intelligence right now is not compute or talent. It is text that no machine has touched. As the open web fills with AI-generated writing, the clean human corpus that trained the first wave of large language models is becoming scarce, contested by its owners, and expensive to replace. The industry's response has been to go offline. Data brokers now source physical books in bulk for AI labs, which scan them into training sets and, in a growing number of cases, pulp the originals afterward. Firms are buying pre-2022 print in batches of 1,000 to as many as 1 million at a time.
The logic is simple and slightly grim. Books printed before roughly 2022 predate the widespread deployment of generative models, so they cannot contain AI output. They were written, edited, and published through human processes, and much of what they contain has never been posted online for free. In a market where the web is increasingly suspect, a shelf of old paperbacks is a certified-clean data source.
Model collapse is the fear driving the spending
The reason a search giant would rather scan a used bookstore than crawl another billion web pages is a failure mode with a name. Model collapse describes what happens when models learn from the output of earlier models rather than from original human data. Errors and blind spots compound generation over generation, rare patterns at the tails of the distribution vanish first, and the system drifts toward a blander, narrower version of itself. The effect was demonstrated under controlled conditions in peer-reviewed work published in Nature in 2024, which showed that recursively training on generated text degrades performance until the model loses the ability to represent the original data at all.
That research turned an abstract worry into an operational constraint. If the web is now saturated with machine text, then every fresh crawl carries a rising share of AI exhaust, and every model trained on it inherits a little more of the previous generation's decay. Verified pre-2022 human writing is the antidote, and there is a fixed amount of it in existence. The scramble for old books is what a supply ceiling looks like when it starts to bite.
Clean data is becoming the moat compute used to be
For three years the assumed source of durable advantage in AI was compute: whoever could marshal the most chips would train the best models. That story is now incomplete. Capability increasingly depends on access to high-quality, uncontaminated, legally clean data, and unlike compute, it cannot simply be manufactured on demand. There is no fab for a 1960s first edition. The labs quietly building the largest scanned archives of pre-AI human text are accumulating an asset their rivals cannot replicate by spending more on hardware.
The legal ground under that asset firmed up considerably last year.
"The training use was a fair use. The technology at issue was among the most transformative many of us will see in our lifetimes."Judge William Alsup, ruling in Bartz v. Anthropic, June 2025
| Pre-2022 books labs buy in a single batch, at the high end | 1,000,000 |
| Cutoff year before which print is treated as slop-free | 2022 |
| Bartz v. Anthropic held book-scanning training can be fair use | June 2025 |
The Bartz decision found that lawfully purchasing a book, converting it to a digital copy, and training on that copy is transformative under US copyright law, provided the original is not retained as a competing library. It gave labs a defensible route to turn physical books into training data at scale, which is precisely why the buying accelerated in the months after.
What this shifts for the companies building models
The second-order effect is a change in what it takes to reach the frontier. A well-funded newcomer can rent enough GPUs to train a large model, but it cannot retroactively acquire a decade of clean, human-authored text that competitors locked up first. Data provenance, once a compliance footnote, becomes a strategic function with its own budget and its own supply chain. Expect the leading labs to keep their archives private, to sign exclusive deals with publishers and estates, and to treat the composition of their training corpus as a trade secret rather than a disclosure.
The uncomfortable irony sits underneath all of it. The technology that promised to generate unlimited content has made genuinely original content more valuable, not less, and the companies that flooded the web with synthetic text are now paying to buy back the human writing that came before them. The frontier is no longer only a race for the biggest model. It is a race for the last clean words.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.