Amazon reportedly destroys rare books to scan them for AI training data

AI Models18.Aug.2026 05:362 min read

A report says Amazon has been buying rare and out-of-print books, cutting off their spines, and scanning them at a warehouse facility to create training data for AI systems. The claim highlights the growing scramble for high-quality, pre-LLM text as major model developers exhaust easily available internet data.

Amazon reportedly destroys rare books to scan them for AI training data

Amazon is reportedly acquiring rare and hard-to-find books, physically dismantling them, and scanning the contents for use in AI-related products and services, according to reporting cited by TechCrunch from 404 Media.

The report says the company has been purchasing books through commercial channels, then removing their spines so the pages can be digitized efficiently at an Amazon facility in Las Vegas known as VGT3. Amazon told 404 Media that it buys books "to improve the products and services customers use," but did not publicly detail the specific models or datasets involved.

Why rare books matter in the AI race

The alleged practice points to a broader issue in generative AI: access to fresh, high-quality text. Most large language models have already been trained on enormous portions of the public web, making offline and out-of-print material increasingly attractive. Rare books can provide text that is not widely available online and is less likely to have been contaminated by AI-generated writing.

That distinction matters because AI companies are trying to avoid training future systems on synthetic output produced by earlier models. Researchers have warned that excessive reliance on AI-generated text can degrade model quality over time, a risk often described as model collapse.

A new pressure point for AI data sourcing

If the reporting is accurate, Amazon's book-scanning effort underscores how valuable legacy print archives have become as training inputs. It also shows how the economics of AI development are reshaping older media supply chains, including secondhand and rare-book markets.

The story may also intensify debate over the provenance of AI training data. Model developers are already under scrutiny over their use of copyrighted and pirated material, and new attention is now falling on whether legally purchased physical books can be converted into large-scale training corpora without broader ethical or cultural concerns.

For the tech industry, the significance is larger than Amazon alone. As easily accessible web data becomes less useful, AI companies appear to be moving deeper into private, licensed, physical, and otherwise scarce sources of text. That shift could become one of the defining competitive battlegrounds in the next phase of model development.