Why AI Companies Are Acquiring Millions of Vintage Books

Why AI Companies Are Acquiring Millions of Vintage Books
By

CNBCTV18.com
 August 1, 2026, 5:18:10 PM IST (Published)

AI firms are facing a shortage of quality data for model training. Their unexpected solution? Acquiring millions of old, printed books. A 2025 court …

CNBCTV18 on Google

Image count
1/8

As the online space fills with AI-generated content, developers are revisiting physical libraries. These traditional texts represent a “premium source” of genuine human-written material essential for developing smarter AI systems.

Image count
2/8

AI companies aim to prevent “model collapse,” a situation where AI performance diminishes after being trained exclusively on AI-generated text. To counteract this, they seek “clean” data from publications released before the generative AI surge of 2022.

Image count
3/8

Books published prior to 2022 are highly valued as they contain professionally edited, humanwritten content ensured to be free from AI influence. This has transformed libraries and storage spaces into the tech sector’s new “gold mines.”

Image count
4/8

In June 2025, a landmark ruling in a U.S. federal court allowed AI businesses to utilize legally acquired physical books for training without needing consent from authors. This ruling differentiates these bought physical copies from pirated digital formats, prompting companies to purchase books in bulk.

Image count
5/8

Interest in physical books has skyrocketed, with certain retailers noting a five-fold increase as of April 2026. Companies like ISBNdb are marketing themselves as large-scale sources, facilitating transactions of up to a million books in one go.

Image count
6/8

To efficiently digitize these books, companies utilize a destructive scanning method. By removing the spine, they can scan hundreds of pages rapidly, a process much more expedient and cost-effective than non-destructive alternatives.

Image count
7/8

After removing the spines and scanning, optical character recognition (OCR) translates the images into digital text for AI model training. The original physical pages, now separated and scanned, are usually discarded or recycled.

Image count
8/8

As these books are transformed into AI models, the loss of their physical forms raises concerns about the dwindling availability of rare and out-of-print editions. Many AI companies, aware of the “optics problem” related to being perceived as book destroyers, operate under stringent non-disclosure agreements to hide their identities and the extent of their acquisitions.

Previous Article

Muthoot Finance proposes Alexander George as the Managing Director starting October 1.

Next Article

Tech Roundup: Spotlight on Dell XPS 13, OnePlus N6x, POCO M8 Power, and Vivo X300e