Amazon’s training pipeline now runs on disassembled rare books, according to an investigation 404 Media published Monday. The outlet slipped a tracking device inside a rare volume and watched it travel to VGT3, an Amazon facility in Las Vegas that brands itself with a dinosaur clutching a book.
Amazon said it buys books through commercial channels to improve the products and services customers use, and declined to say how many titles have been processed or which models consume the scans.
The logic of the program is simple: the internet has already been mined dry. Out-of-print and hard-to-find texts offer untouched human writing, and anything published before 2022 predates the AI-generated flood, so training on it carries less risk of model collapse, the degradation that hits models trained on their own output.
The practice reopens the copyright battles that have shadowed the AI boom. Anthropic previously settled claims that it used pirated book collections, and publishers have sued other labs over scraped text. Amazon’s approach differs in one notable way: it buys physical copies, which raises the question of whether ownership of a book confers the right to dismantle it for machine learning.