Amazon is buying rare books in bulk, slicing off their spines, and scanning the pages to train AI models, according to a 404 Media investigation that planted a tracking device inside a rare book and followed it to an Amazon facility in Las Vegas. The site, known internally as VGT3, is identified by a logo of a dinosaur clutching a book. Amazon confirmed the underlying activity in a statement, saying it "purchases books through commercial channels to improve the products and services customers use."
The company that began in 1994 as an online bookseller is now destroying the physical objects that once defined it. Cutting the spine off a book is the standard preparation for high-speed sheet-fed scanning, which converts bound volumes into machine-readable text faster and more accurately than non-destructive methods. The books do not survive the process.
The economics are straightforward. Frontier labs have already scraped most of what the open internet offers, and the remaining supply of high-quality human-written text is a hard constraint on further scaling. Rare and out-of-print books represent a source of clean training data that competitors cannot easily replicate by crawling the web.
“purchases books through commercial channels to improve the products and services customers use.”— Amazon, Company statement to 404 Media
Key facts
- 01404 Media planted a tracker in a rare book and traced it to Amazon's VGT3 facility in Las Vegas, marked with a dinosaur clutching a book.
- 02Amazon confirmed it buys books through commercial channels to improve its products and services.
- 03Books printed before 2022 are prized training data because they predate the era of LLM-generated text.
- 04Training on AI-generated text can trigger model collapse, where output quality degrades over successive generations.
- 05Anthropic has previously trained on pirated books, part of an industry-wide scramble for clean text.
Pre-2022 texts carry a specific premium. Anything printed before the release of ChatGPT is guaranteed to be free of LLM-generated content, which matters because models trained heavily on synthetic text risk what researchers call model collapse — a degradation in output quality as the training corpus becomes recursively contaminated with prior model outputs. A book printed in 1978 is, by definition, uncontaminated.
Amazon is not alone in facing this data problem, and the industry's track record on sourcing is uneven. Anthropic previously trained Claude on pirated books, a fact that has surfaced in ongoing copyright litigation. Amazon's approach — buying physical copies through commercial channels before destroying them — is more legally defensible than torrenting a shadow library, though the copyright status of scanning purchased books at scale for model training remains contested and largely untested in court.
The VGT3 operation also raises a preservation question that is separate from the copyright one. Rare books, particularly those out of print or with limited surviving copies, have cultural value beyond their text. A destructive scan produces a digital record but eliminates the artifact. Libraries and archives have spent decades building non-destructive digitization workflows for exactly this reason. Amazon's pipeline is optimized for throughput, not preservation.
The broader picture is a training-data arms race that is running out of easy inputs. OpenAI, Anthropic, Google, and Meta have all faced lawsuits over how they acquired the text their models learned from. Synthetic data pipelines, licensing deals with publishers, and now physical book acquisition are the responses. Amazon's willingness to run a dedicated facility for destructive scanning suggests the marginal value of clean, pre-2022 text is high enough to justify the operational cost and the reputational risk.
Amazon has not disclosed the scale of the VGT3 operation, how many books it processes, or which of its AI products — the Nova model family, Alexa's underlying systems, or Bedrock-hosted offerings — are trained on the scanned corpus. The company's statement stopped at the general acknowledgment that it buys books for product improvement.
The reputational math for Amazon is worth watching. The company built its brand on making books more accessible; a story about it destroying rare copies to feed a model runs directly against that origin story. But the training-data shortage is real, and every frontier lab is going to have to answer the same question about how it acquired its text. Amazon is one of the first to answer it with a facility address and a dinosaur logo. Expect competitors to have quieter versions of the same operation, and expect the copyright bar to keep pressing on whether purchased-and-scanned counts as fair use at industrial scale.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




