Investigation: Amazon buys and destroys rare books to scan them for AI training

Amazon is buying large quantities of books, scanning them for AI training data, and destroying the printed copies in the process, according to a 404 Media investigation published Aug 17, 2026. The investigation revealed Amazon's book-buying operation, which had not been previously reported, by placing a tracking device inside a rare book suspected of being acquired by an AI company. The shipment traveled from California through Milwaukee, Kenosha, Colorado Springs, and a truck stop in Grand Jun

2 min
Investigation: Amazon buys and destroys rare books to scan them for AI training

Amazon is buying large quantities of books, scanning them for AI training data, and destroying the printed copies in the process, according to a 404 Media investigation published Aug 17, 2026.

The investigation revealed Amazon's book-buying operation, which had not been previously reported, by placing a tracking device inside a rare book suspected of being acquired by an AI company. The shipment traveled from California through Milwaukee, Kenosha, Colorado Springs, and a truck stop in Grand Junction, Colorado before reaching its final destination: an Amazon warehouse in Las Vegas, Nevada.

Abstract diagram of rare books arriving at a warehouse and being scanned into a digital archive

Employees at the Las Vegas facility, identified by the Amazon team label VGT3, told 404 Media that all they do is receive massive shipments of printed books, cut the bindings off to scan them more quickly, and discard the books. The printed book is destroyed in the process. The team's logo is a dinosaur holding a book.

Amazon confirmed the purchases in a statement: Amazon purchases books through commercial channels to help develop and improve the products and services our customers use. The company did not dispute that books are destroyed during scanning.

The reporting follows a July 404 Media story in which booksellers described a sharp, seemingly random spike in bulk orders over the past year, which they suspected were driven by AI companies seeking training data. Printed books are valued as training data because much of their text is not available on the internet, is conveniently organized, and, if printed before 2022, is free of AI-generated text that can degrade models through a recursive problem known as model collapse.

Why it matters

The episode sharpens a debate over how AI companies source training data. Unlike libraries and universities, the anonymous bulk buyers were not price sensitive, and booksellers could not identify them because marketplaces keep buyer identities hidden. The investigation confirms that at least one major AI buyer is destroying physical books as it digitizes them.

Sources

We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility - 404 Media, Aug 17, 2026 (Emanuel Maiberg): https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

Written by

More to read

  • Automated Prompt Optimization in Production: Signatures, Teleprompters, and Metric-Driven Compilation with DSPy

    Manual prompt engineering remains one of the largest sources of technical debt in modern LLM applications. Teams routinely spend weeks hand-crafting multi-paragraph system prompts, hardcoding few-shot examples, and tweaking phrasing to extract reliable outputs from specific model checkpoints. When the underlying model is upgraded, migrated to an open-weight alternative, or integrated into a multi-step pipeline, these hand-crafted strings break, requiring another cycle of trial-and-error adjustme

    1 min
  • Model Merging in Large Language Models: How Task Arithmetic, TIES, and DARE Combine Checkpoints Without Training

    Fine-tuning foundation models for specialized tasks typically produces isolated checkpoints. A model adapted for mathematical reasoning retains high numerical precision but often degrades in general dialogue or code generation. Traditionally, unifying these capabilities required multi-task training: gathering mixed datasets, re-running optimization across multiple GPUs, and managing gradient conflicts during backpropagation. Model merging provides an alternative paradigm. By operating directly

    1 min
  • Harvey Introduces Tenet, Its First In-House Legal LLM Trained on Moonshot's Kimi K3

    Legal AI startup Harvey has announced Harvey Tenet, its first proprietary, in-house foundation model tailored for legal workflows. The release marks a strategic shift for the $11 billion legal tech company, which has historically relied on API access to third-party frontier models from OpenAI and Anthropic. Tenet is post-trained on top of Kimi K3, an open-weights model released in July 2026 by Chinese AI lab Moonshot AI. The initiative is part of a broader platform update titled Harvey II, whic

    1 min