A shipment of rare books, a hidden AirTag, and a warehouse in Las Vegas have put Amazon’s book scanning practices under scrutiny. According to an investigation by 404 Media, Amazon buys printed books in bulk, scans them for AI training data, and destroys the physical copies during the process.
The reporting connects the books to Amazon’s Nova models and to a warehouse team called VGT3. It also places Amazon inside a broader debate over how AI companies turn printed knowledge into private training material.
What the AirTag investigation found
404 Media placed an AirTag inside a shipment of rare books and followed the shipment across the country. The tracker eventually led to an Amazon warehouse in Las Vegas.
That warehouse is home to a team called VGT3. The team’s logo features a Tyrannosaurus rex holding a book, a detail that stands out because workers there say the process used on books is destructive.
According to workers cited in the source article, the team cuts off book spines to make scanning faster. Once the spine is removed, the copy is destroyed as a usable printed book. Amazon then uses the scanned data to train its Nova models.
A spokesperson said Amazon buys books through commercial channels to improve its products. The source does not describe the exact purchasing process beyond that, but the central point is clear: physical books are being acquired, digitized, and consumed by a training pipeline.
Why printed books matter for AI training
The source article explains why printed texts are especially attractive to AI companies. Many books do not exist online, which means their contents may be missing from the large bodies of digital text commonly associated with AI training.
Another factor is timing. Printed texts often predate 2022, meaning they are free of AI-generated content. For companies building models, that makes older printed material valuable because it can represent human-written language that was not shaped by the recent wave of generative AI output.
Booksellers suspect that AI companies are trying to systematically scan every book by ISBN number. The source presents that as a suspicion from booksellers, not as a confirmed plan from Amazon. Still, it shows why the book trade is watching bulk purchases closely.
If the goal is broad coverage of printed culture, rare books become especially important. They may contain material that is hard to find elsewhere, and once scanned they can become part of a private dataset rather than remaining only on a public shelf or in the hands of collectors and booksellers.
Amazon is not the only company named
The source article also points to Anthropic as another company that ran a similar operation. A lawsuit by book authors revealed a program called "Project Panama," in which Anthropic bought books on marketplaces, cut off their spines, and digitized them.
In that case, the judge ruled that the scanning qualified as fair use and did not violate copyright. The source says that part of the reasoning was that the printed originals were destroyed and therefore were not copied and resold.
That detail matters because it shows how destruction can become part of the legal and practical logic of scanning. The copy is not preserved as a resold book; it is converted into data and removed from circulation as a physical object.
The source does not say Amazon’s operation is legally identical to Anthropic’s, and it does not provide a court ruling about Amazon. What it does show is that destructive scanning is not limited to one company.
The core concern is access to knowledge
The controversy is not only about copyright or corporate purchasing. It is also about what happens when sometimes rare books are turned into private training material.
When a company scans a book and destroys the original, the contents may help improve a closed AI model. But the physical copy is no longer available to a reader, a bookseller, a library, or another buyer. That tradeoff is especially sensitive when the books may be irreplaceable.
The source frames the concern as knowledge being pulled off public shelves and locked inside the closed AI models of a single corporation. That is a different outcome from ordinary digitization meant to preserve or share access. Here, the result described is private model training, not public availability.
The facts reported by 404 Media leave several questions open, including how many books are involved and how Amazon chooses them. But the practice described is concrete enough to raise a larger issue for the AI industry: printed books are not just inputs. They are physical artifacts, and destroying them changes who can access the knowledge they contain.
Amazon’s book scanning for Nova models shows how the demand for AI training data is moving beyond the open web. The next frontier may be the shelves of printed material that have not already been absorbed into digital systems.