
이미지: The Decoder
Summary
- 404 Media hid an Apple AirTag in a box of rare books and traced it to Amazon's "VGT3" team at a Las Vegas warehouse
- Warehouse workers cut off book spines to speed up scanning, destroying the originals in the process, with Amazon using the resulting data to train its Nova model
- Anthropic was found to have carried out similar book purchasing, cutting, and scanning under an internal effort called "Project Panama," revealed during a copyright lawsuit, and a court ruled it fair use partly because the physical originals were destroyed
- 추적 방식
- 404 Media가 희귀 도서 상자에 애플 에어태그를 숨겨 배송 경로 추적
- 도착지
- 라스베이거스 아마존 창고, 팀명 VGT3(로고: 책을 문 티라노사우루스)
- 파쇠 방식
- 노동자들이 스캔 속도를 높이려 책등을 절단, 원본 훼손
- 훈련 대상
- 아마존 자체 모델 노바(Nova)
- 아마존 입장
- 대변인, 상업적 유통 경로로 도서를 구매해 제품 개선에 활용한다고 답변
- 유사 사례
- 앤스로픽 '프로젝트 파나마' — 도서 구매 후 절단·디지털화, 저작권 소송에서 드러남
- 법원 판단
- 앤스로픽 사건에서 원본 파괴로 재복제·재판매가 없었다는 점을 근거로 공정 이용 인정
A hidden tracker in a box of books
404 Media hid an Apple AirTag inside a box of rare books and followed its shipping route to the end. The box arrived at an Amazon warehouse in Las Vegas. The team responsible for scanning books at this warehouse is called VGT3, and its team logo features a Tyrannosaurus holding a book in its mouth, according to 404 Media's investigative report.
Cutting spines to speed up scanning
Workers at the warehouse said they cut off the spines of books to speed up scanning. Rather than scanning books page by page, cutting the spine to separate the pages makes machine scanning much faster. The problem is that this process effectively destroys the original book beyond repair. Amazon uses the text data obtained this way to train its own large language model, Nova. A company spokesperson said only that the books were purchased through commercial distribution channels and used for product improvement, without directly addressing the cutting and shredding practice itself.
Why print copies — suspicions of an ISBN-by-ISBN sweep
Booksellers suspect AI companies are scanning nearly every book that exists, tracked by ISBN number. Print copies are especially valuable for two reasons: many books have never been posted online as text, and books published before 2022 were written before generative AI became widespread. There's a shared industry concern that text flooding the web since ChatGPT's debut is mixed with AI-generated writing, undermining its reliability as training data — a contamination that print originals are free from.
Anthropic walked the same path — Project Panama
This approach isn't unique to Amazon. Anthropic was found to have carried out similar work, revealed during a copyright lawsuit filed by book authors. The internal project was code-named "Project Panama," and Anthropic bought books from secondhand markets, cut off the spines, and digitized them. The judge overseeing the lawsuit ruled that this scanning work constituted fair use and did not infringe copyright. One factor cited in the ruling was that the physical originals were destroyed rather than being reproduced or resold.
| Item | Amazon | Anthropic |
|---|---|---|
| Internal project name | (undisclosed, team name VGT3) | Project Panama |
| Method | Purchase books, cut spines, scan | Purchase books, cut spines, scan |
| Training purpose | In-house model, Nova | Data for language model training |
| How it was exposed | 404 Media's AirTag tracking | Authors' copyright lawsuit |
| Legal outcome | Not confirmed | Ruled fair use, based partly on destruction of originals |
Editor's take
This story sits at a clear inflection point. There's a shared understanding across the industry that clean, unpolluted text on the web is running out. Books published before 2022 are verified, "pure" data — guaranteed not to contain a single line written by generative AI. Once companies decide that value far outweighs the cost of buying and shredding a book, they build warehouses instead of libraries. The fact that the VGT3 team's logo deliberately features a dinosaur holding a book suggests this work has already become formalized practice within the organization.
Data collection used to conjure images of web crawling or licensing deals. Now the picture has shifted to manual labor in a logistics warehouse, cutting book spines. What makes this shift interesting is that it's AI companies reaching backward from "digital" to "analog." Also worth noting is how Anthropic's court ruling lent legitimacy to this trend. The logic that destroying an original and not reselling it in the market means no copyright infringement gives other AI companies more incentive to adopt the same approach.
There are two practical lessons for domestic companies here. One is that whether training data is "contaminated" is becoming an increasingly significant variable in model quality. The other is that when negotiating contracts to hand over copyrighted content for AI training, an "original destruction" clause could serve as a legal line of defense. Publishers, libraries, and organizations holding archives should recognize that, amid this trend, their print assets carry more negotiating leverage than they might think.
It's likely that similar cases involving other Big Tech companies will surface in the coming weeks. With two companies already implicated, it would be more surprising if follow-up investigations into Google's or Meta's logistics networks didn't emerge.



