One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Amazon's Rare-Book Shredding Scan Facility Exposed by AirTag Tracker

404 Media planted a tracker in a book shipment and traced it to a Las Vegas warehouse, where spines were cut off books later used to train Nova

책에서 글자와 디지털 코드가 쏟아지는 도서관 콜라주 이미지

이미지: The Decoder

Summary

  • 404 Media hid an Apple AirTag in a box of rare books and traced it to Amazon's "VGT3" team at a Las Vegas warehouse
  • Warehouse workers cut off book spines to speed up scanning, destroying the originals in the process, with Amazon using the resulting data to train its Nova model
  • Anthropic was found to have carried out similar book purchasing, cutting, and scanning under an internal effort called "Project Panama," revealed during a copyright lawsuit, and a court ruled it fair use partly because the physical originals were destroyed
추적 방식
404 Media가 희귀 도서 상자에 애플 에어태그를 숨겨 배송 경로 추적
도착지
라스베이거스 아마존 창고, 팀명 VGT3(로고: 책을 문 티라노사우루스)
파쇠 방식
노동자들이 스캔 속도를 높이려 책등을 절단, 원본 훼손
훈련 대상
아마존 자체 모델 노바(Nova)
아마존 입장
대변인, 상업적 유통 경로로 도서를 구매해 제품 개선에 활용한다고 답변
유사 사례
앤스로픽 '프로젝트 파나마' — 도서 구매 후 절단·디지털화, 저작권 소송에서 드러남
법원 판단
앤스로픽 사건에서 원본 파괴로 재복제·재판매가 없었다는 점을 근거로 공정 이용 인정

A hidden tracker in a box of books

404 Media hid an Apple AirTag inside a box of rare books and followed its shipping route to the end. The box arrived at an Amazon warehouse in Las Vegas. The team responsible for scanning books at this warehouse is called VGT3, and its team logo features a Tyrannosaurus holding a book in its mouth, according to 404 Media's investigative report.

Cutting spines to speed up scanning

Workers at the warehouse said they cut off the spines of books to speed up scanning. Rather than scanning books page by page, cutting the spine to separate the pages makes machine scanning much faster. The problem is that this process effectively destroys the original book beyond repair. Amazon uses the text data obtained this way to train its own large language model, Nova. A company spokesperson said only that the books were purchased through commercial distribution channels and used for product improvement, without directly addressing the cutting and shredding practice itself.

Why print copies — suspicions of an ISBN-by-ISBN sweep

Booksellers suspect AI companies are scanning nearly every book that exists, tracked by ISBN number. Print copies are especially valuable for two reasons: many books have never been posted online as text, and books published before 2022 were written before generative AI became widespread. There's a shared industry concern that text flooding the web since ChatGPT's debut is mixed with AI-generated writing, undermining its reliability as training data — a contamination that print originals are free from.

Anthropic walked the same path — Project Panama

This approach isn't unique to Amazon. Anthropic was found to have carried out similar work, revealed during a copyright lawsuit filed by book authors. The internal project was code-named "Project Panama," and Anthropic bought books from secondhand markets, cut off the spines, and digitized them. The judge overseeing the lawsuit ruled that this scanning work constituted fair use and did not infringe copyright. One factor cited in the ruling was that the physical originals were destroyed rather than being reproduced or resold.

ItemAmazonAnthropic
Internal project name(undisclosed, team name VGT3)Project Panama
MethodPurchase books, cut spines, scanPurchase books, cut spines, scan
Training purposeIn-house model, NovaData for language model training
How it was exposed404 Media's AirTag trackingAuthors' copyright lawsuit
Legal outcomeNot confirmedRuled fair use, based partly on destruction of originals

Editor's take

This story sits at a clear inflection point. There's a shared understanding across the industry that clean, unpolluted text on the web is running out. Books published before 2022 are verified, "pure" data — guaranteed not to contain a single line written by generative AI. Once companies decide that value far outweighs the cost of buying and shredding a book, they build warehouses instead of libraries. The fact that the VGT3 team's logo deliberately features a dinosaur holding a book suggests this work has already become formalized practice within the organization.

Data collection used to conjure images of web crawling or licensing deals. Now the picture has shifted to manual labor in a logistics warehouse, cutting book spines. What makes this shift interesting is that it's AI companies reaching backward from "digital" to "analog." Also worth noting is how Anthropic's court ruling lent legitimacy to this trend. The logic that destroying an original and not reselling it in the market means no copyright infringement gives other AI companies more incentive to adopt the same approach.

There are two practical lessons for domestic companies here. One is that whether training data is "contaminated" is becoming an increasingly significant variable in model quality. The other is that when negotiating contracts to hand over copyrighted content for AI training, an "original destruction" clause could serve as a legal line of defense. Publishers, libraries, and organizations holding archives should recognize that, amid this trend, their print assets carry more negotiating leverage than they might think.

It's likely that similar cases involving other Big Tech companies will surface in the coming weeks. With two companies already implicated, it would be more surprising if follow-up investigations into Google's or Meta's logistics networks didn't emerge.