
이미지: The Decoder
Summary
- Analysis of 14,419 self-published ebooks found that heavily AI-generated books make up 20% of the catalog but account for only 12.1% of revenue
- Even looking only at books with zero detected AI text, per-title revenue fell in 7 of 8 genres, confirming a "dilution effect"
- Researchers say the findings could serve as long-missing evidence of market harm in copyright infringement lawsuits
- 분석 대상
- 2023년 1월~2026년 3월 출간 자가출판 전자책 14,419권
- AI 판별 도구
- Pangram v3.3(오탐률 0.04%), 이후 Pangram 4 출시
- AI 다량 함유 도서
- 카탈로그의 20%, 판매량 12.1%, 매출액 11.3%
- AI 텍스트 없는 도서
- 카탈로그의 62.9%, 매출액 72.5%
- 카탈로그·매출 성장률
- 2023년 1분기~2026년 1분기 카탈로그 38.3배, 분기 매출 8.9배
- 권당 매출 하락
- 8개 장르 중 6개 하락, AI 텍스트 없는 책만 봐도 7개 장르 하락
- 예외 장르
- 판타지/초자연/호러 - 무AI 도서 권당 매출 35% 상승
Catalog grew 38-fold, but revenue grew only 9-fold
The numbers tell a simple story about what happened in Amazon's self-publishing market from 2023 to early 2026. The number of titles being sold grew 38.3-fold, while the revenue those titles shared grew only 8.9-fold. That's the finding from researchers at Stony Brook University, who conducted a full analysis of 14,419 self-published ebooks. The sample was drawn from an internal database maintained by one of the "Big Five" US publishers, which reportedly tracks roughly 500,000 Amazon titles and covers about 95% of daily ebook sales.
Unlike previous studies that judged AI involvement based only on short book previews, the researchers fed entire manuscripts into a detector. The tool used was Pangram v3.3, which its developer states has a false-positive rate of 0.04%. A newer version, Pangram 4, has since been released, but this analysis was conducted using v3.3. Books were sorted into three tiers based on the proportion of AI-generated text: "none," "minor" (25% or less), and "substantial" (more than 25%).
AI books make up 20% of the catalog but only 12% of revenue
Books with substantial AI text accounted for 20% of the catalog under study, but only 12.1% of units sold and 11.3% of revenue. In contrast, books with zero detected AI text made up 62.9% of the catalog while capturing 72.5% of revenue.
| Category | Share of catalog | Share of units sold | Share of revenue |
|---|---|---|---|
| No AI text | 62.9% | - | 72.5% |
| Substantial AI text (25%+) | 20% | 12.1% | 11.3% |
Taken at face value, these numbers point to a familiar conclusion: AI-written books remain low-quality "slop" stuck at the bottom of the market. But the researchers argue this interpretation misses what's actually happening in the market.
Even AI-free books lost revenue in 7 genres
Comparing books published in 2023 with those published in 2025 over the same post-release time window, per-title revenue fell in 6 of 8 genres. What's notable is that even when looking only at books with zero detected AI text, revenue fell in 7 of 8 genres. This means the phenomenon can't be fully explained simply by poor-selling AI books dragging down the average.
The researchers call this phenomenon "dilution." They caution that this is an observed correlation, not a relationship proven through experimental causation. The lone exception was the fantasy/paranormal/horror genre. In this genre—where AI text arrived latest and spread most slowly—per-title revenue for AI-free books actually rose by 35%. The researchers point to this reversal as evidence that overall market trends aren't the driving factor. The dilution pattern was more pronounced in genres with heavy Kindle Unlimited usage, where readers draw from a single subscription pool of books.
The "market harm" evidence that copyright lawsuits have needed
The researchers believe this analysis could provide substantive material for copyright infringement lawsuits against AI companies. Plaintiffs have long argued that AI companies harmed the market by training on copyrighted works, but have lacked data to quantify that harm. This study presents a concrete pattern showing that per-title revenue is declining even for books by human authors, separate from the growth in AI-generated books.
In a separate analysis released by the same research team on August 11, they used the Allen Institute for AI (Ai2)'s infini-gram engine to examine overlap in rare phrasing between AI-generated text and existing published works. At the time, 200 Amazon bestsellers with high proportions of AI text showed 4.4 percentage points more rare-phrase overlap than a general comparison group, and the gap widened to 22.5 percentage points when compared against award-winning and nominated literary works. Taken together with this revenue analysis, evidence is accumulating on both fronts—"textual similarity" and "market harm"—relevant to copyright infringement disputes.
Editor's Take
What makes this study interesting is that its conclusion flips on closer inspection. At first glance, the numbers offer a reassuring, familiar story: "AI books are still low-tier slop." Making up 20% of the catalog but only 12% of revenue seems to confirm the narrative that readers ultimately filter out low-quality AI text. But when the researchers dug one layer deeper and isolated books with zero AI text, even those books saw declining revenue in 7 of 8 genres. That's not a quality problem with AI books—it's purely a volume problem. While the size of the pie grew 9-fold, the number of mouths sharing it grew 38-fold, creating a structure where human authors get a smaller slice even while writing books just as good as before.
Something similar already happened in stock photography, YouTube Shorts, and spot-content markets before this hit publishing. When a new channel emerges, quality initially acts as a filter—but once volume crosses a critical threshold, overall unit prices fall regardless of quality. The self-published ebook market simply hit that threshold unusually fast because it has virtually no barriers to entry. It also makes sense that this dilution effect appears stronger in flat-rate subscription pools like Kindle Unlimited, where the pie itself is fixed—so as more people share it, everyone's slice shrinks.
For Korea's content industry—especially those running or publishing on web-novel and ebook platforms—these numbers shouldn't be dismissed as someone else's problem. Volume hasn't piled up to Amazon's scale yet, but a market with the same structure is likely headed in the same direction. There are two practical responses available now. One is to focus not on publishing speed but on building direct channels with readers—subscriptions, newsletters, fandoms—since dilution shows up strongly in anonymous search/recommendation traffic but affects creators who already have a fanbase relatively less. The other is to pay attention to this trend of proving harm with data, as this study does. Once creator groups in Korea start compiling similar revenue data, the copyright debate may soon shift from an emotional argument to a battle of numbers.
In the coming months, this study is likely to be cited as evidence in ongoing AI copyright lawsuits, and platforms including Amazon will likely face growing pressure to address how AI-generated content is labeled and surfaced in search.



