METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Amazon's Rare-Book Shredding Scan Facility Exposed by AirTag Tracker

404 Media planted a tracker in a book shipment and traced it to a Las Vegas warehouse, where spines were cut off books later used to train Nova

Amazon's Rare-Book Shredding Scan Facility Exposed by AirTag Tracker

Image: METAL

Summary

  • 404 Media hid an Apple AirTag in a box of rare books and traced it to Amazon's "VGT3" team at a Las Vegas warehouse
  • Warehouse workers cut off book spines to speed up scanning, destroying the originals in the process, with Amazon using the resulting data to train its Nova model
  • Anthropic was found to have carried out similar book purchasing, cutting, and scanning under an internal effort called "Project Panama," revealed during a copyright lawsuit, and a court ruled it fair use partly because the physical originals were destroyed

A hidden tracker in a box of books

404 Media hid an Apple AirTag inside a box of rare books and followed its shipping route to the end. The box arrived at an Amazon warehouse in Las Vegas. The team responsible for scanning books at this warehouse is called VGT3, and its team logo features a Tyrannosaurus holding a book in its mouth, according to 404 Media's investigative report.

Cutting spines to speed up scanning

Workers at the warehouse said they cut off the spines of books to speed up scanning. Rather than scanning books page by page, cutting the spine to separate the pages makes machine scanning much faster. The problem is that this process effectively destroys the original book beyond repair. Amazon uses the text data obtained this way to train its own large language model, Nova. A company spokesperson said only that the books were purchased through commercial distribution channels and used for product improvement, without directly addressing the cutting and shredding practice itself.

Why print copies — suspicions of an ISBN-by-ISBN sweep

Booksellers suspect AI companies are scanning nearly every book that exists, tracked by ISBN number. Print copies are especially valuable for two reasons: many books have never been posted online as text, and books published before 2022 were written before generative AI became widespread. There's a shared industry concern that text flooding the web since ChatGPT's debut is mixed with AI-generated writing, undermining its reliability as training data — a contamination that print originals are free from.

Anthropic walked the same path — Project Panama

This approach isn't unique to Amazon. Anthropic was found to have carried out similar work, revealed during a copyright lawsuit filed by book authors. The internal project was code-named "Project Panama," and Anthropic bought books from secondhand markets, cut off the spines, and digitized them. The judge overseeing the lawsuit ruled that this scanning work constituted fair use and did not infringe copyright. One factor cited in the ruling was that the physical originals were destroyed rather than being reproduced or resold.

ItemAmazonAnthropic
Internal project name(undisclosed, team name VGT3)Project Panama
MethodPurchase books, cut spines, scanPurchase books, cut spines, scan
Training purposeIn-house model, NovaData for language model training
How it was exposed404 Media's AirTag trackingAuthors' copyright lawsuit
Legal outcomeNot confirmedRuled fair use, based partly on destruction of originals

Editor's take

This story sits at a clear inflection point. There's a shared understanding across the industry that clean, unpolluted text on the web is running out. Books published before 2022 are verified, "pure" data — guaranteed not to contain a single line written by generative AI. Once companies decide that value far outweighs the cost of buying and shredding a book, they build warehouses instead of libraries. The fact that the VGT3 team's logo deliberately features a dinosaur holding a book suggests this work has already become formalized practice within the organization.

Data collection used to conjure images of web crawling or licensing deals. Now the picture has shifted to manual labor in a logistics warehouse, cutting book spines. What makes this shift interesting is that it's AI companies reaching backward from "digital" to "analog." Also worth noting is how Anthropic's court ruling lent legitimacy to this trend. The logic that destroying an original and not reselling it in the market means no copyright infringement gives other AI companies more incentive to adopt the same approach.

There are two practical lessons for domestic companies here. One is that whether training data is "contaminated" is becoming an increasingly significant variable in model quality. The other is that when negotiating contracts to hand over copyrighted content for AI training, an "original destruction" clause could serve as a legal line of defense. Publishers, libraries, and organizations holding archives should recognize that, amid this trend, their print assets carry more negotiating leverage than they might think.

It's likely that similar cases involving other Big Tech companies will surface in the coming weeks. With two companies already implicated, it would be more surprising if follow-up investigations into Google's or Meta's logistics networks didn't emerge.

Comments