AI firms are facing a shortage of quality data for model training. Their unexpected solution? Acquiring millions of old, printed books. A 2025 court …
1/8
As the online space fills with AI-generated content, developers are revisiting physical libraries. These traditional texts represent a “premium source” of genuine human-written material essential for developing smarter AI systems.
2/8
AI companies aim to prevent “model collapse,” a situation where AI performance diminishes after being trained exclusively on AI-generated text. To counteract this, they seek “clean” data from publications released before the generative AI surge of 2022.
3/8
Books published prior to 2022 are highly valued as they contain professionally edited, humanwritten content ensured to be free from AI influence. This has transformed libraries and storage spaces into the tech sector’s new “gold mines.”
4/8
In June 2025, a landmark ruling in a U.S. federal court allowed AI businesses to utilize legally acquired physical books for training without needing consent from authors. This ruling differentiates these bought physical copies from pirated digital formats, prompting companies to purchase books in bulk.
5/8
Interest in physical books has skyrocketed, with certain retailers noting a five-fold increase as of April 2026. Companies like ISBNdb are marketing themselves as large-scale sources, facilitating transactions of up to a million books in one go.
6/8
To efficiently digitize these books, companies utilize a destructive scanning method. By removing the spine, they can scan hundreds of pages rapidly, a process much more expedient and cost-effective than non-destructive alternatives.
7/8
After removing the spines and scanning, optical character recognition (OCR) translates the images into digital text for AI model training. The original physical pages, now separated and scanned, are usually discarded or recycled.
8/8
As these books are transformed into AI models, the loss of their physical forms raises concerns about the dwindling availability of rare and out-of-print editions. Many AI companies, aware of the “optics problem” related to being perceived as book destroyers, operate under stringent non-disclosure agreements to hide their identities and the extent of their acquisitions.