An AirTag, 1,000 Rare Books, and What It Reveals About AI Training-Data Provenance
A 404 Media investigation tracked a bulk book order to an Amazon facility that destructively scans books for AI training — a reminder that most organisations can't answer a basic governance question: where did our model's training data actually come from?
Key Takeaways
- 404 Media hid an Apple AirTag in a roughly 1,000-book order placed through the marketplace Biblio and traced it to VGT3, a unit inside Amazon's LAS8 facility in Las Vegas.
- Workers described a destructive process: bindings are cut off so pages can be scanned faster, destroying the physical book — reportedly feeding Amazon's Nova model training data.
- Pre-2022 printed books are valuable precisely because they predate the flood of AI-generated text online, making them 'clean' training data — but the acquisition itself was anonymous and undisclosed to sellers.
- For security and governance teams, the real story isn't the book-shredding — it's that most organisations still can't produce a defensible answer to 'where did this model's training data come from.'
What 404 Media found
According to reporting highlighted by Simon Willison, 404 Media placed a hidden Apple AirTag inside a book that was part of an order of roughly 1,000 volumes, bought from a bookseller on the marketplace Biblio by an anonymous, price-insensitive buyer in July. The tag's route ended at LAS8, an Amazon facility in Las Vegas, inside a unit workers refer to as VGT3 — marked, per the reporting, by a logo of a dinosaur digging its claws into a book.
Amazon employees told 404 Media that VGT3's job is to receive large shipments of printed books, cut the bindings off so the pages can be scanned faster, and discard the now-destroyed physical copies. The scanned text reportedly feeds training data for Amazon's Nova model family, according to The Decoder's coverage of the same investigation.
Why rare, pre-2022 books specifically
The bulk-buying pattern behind this — anonymous buyers paying above market rate for large lots of physical books — has been documented before; Willison previously covered Anthropic's own book-scanning activity in June 2025. The appeal is straightforward: printed texts that predate 2022 largely don't exist online, and they're free of the synthetic text that now saturates the open web. As foundation-model teams run low on clean, high-quality human-written data, out-of-print books are one of the few remaining large, unpolluted corpora — and physical scanning sidesteps the copyright and licensing friction of scraping text that's already digitised and controlled.
The governance angle security teams should care about
None of this is a breach in the conventional sense — no CVE, no exploited vulnerability. But it's a clean illustration of a problem that AI-security and AI-governance practitioners deal with constantly: training-data provenance is opaque by default, and most organisations deploying or fine-tuning models cannot produce a verifiable record of what data went in, how it was acquired, or under what rights.
- Undisclosed acquisition: the booksellers reportedly didn't know the buyer's identity or purpose — a pattern that, inside an enterprise, would fail almost any data-sourcing audit.
- No chain of custody: once a book is destructively scanned, there is no way to later verify what was actually digitised, in what condition, or with what fidelity.
- Copyright and rights exposure: scanning copyrighted physical works at scale for model training sits in the same contested legal territory as web-scraped training data, with active litigation shaping the outcome.
- Downstream trust: any model trained partly on undocumented sourcing inherits a provenance gap that's very hard to close retroactively — a live issue for organisations trying to attest to responsible AI sourcing under frameworks like ISO/IEC 42001.
ISO 42001's AI management system requirements explicitly call for documented data provenance and lifecycle controls. Stories like this are a useful prompt for any organisation building on top of third-party foundation models to ask a narrower, more answerable question: not 'was this specific dataset acquired ethically,' but 'do we have a documented basis for trusting the data our vendor's model was trained on, and does our own AI governance process demand that evidence before we deploy?'
FAQ
Frequently Asked Questions
Did Amazon confirm it destroys books to train its Nova AI models?
404 Media's reporting is based on tracking a book shipment to Amazon's LAS8 facility in Las Vegas and interviews with workers describing the scanning process; the outlets covering the story, including The Decoder, report the scanned data feeds Nova model training, but this is sourced to worker accounts and investigative reporting rather than an Amazon public statement.
Is buying and scanning physical books for AI training illegal?
It sits in legally contested territory. Courts are still working through whether training on copyrighted works — scraped or scanned — qualifies as fair use, and outcomes so far have varied by jurisdiction and use case, so no blanket answer currently applies.
Why would AI companies prefer physical books over digital text scraping?
Printed books that predate roughly 2022 are largely absent from the open web and untouched by AI-generated content, making them a comparatively 'clean' data source as models increasingly need to avoid training on synthetic text generated by earlier models.