One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

35% of Web Pages Published Since ChatGPT Show Signs of AI Authorship

Pew Research analyzes 500,000 Common Crawl pages — .com domains show AI authorship rate 10x higher than .edu or .gov

노트북 화면에 오픈AI 로고와 이름이 클로즈업되어 있다

이미지: TechCrunch AI

Summary

  • Pew Research Center analyzed roughly 500,000 English-language web pages from the Common Crawl archive and found that 35% of pages published since ChatGPT's launch showed signs of AI authorship.
  • In a random sample of 10,000 pages collected in July 2026, only 10% of all pages were flagged as suspected AI-authored, a lower figure because the sample included pages predating ChatGPT.
  • By domain, .com sites showed an AI authorship rate roughly 10 times higher than .edu and .gov (about 1% each), while .org came in at 4.6%.
발행 주체
퓨리서치 연구소(Pew Research), 2026년 8월 20일 목요일 공개
분석 대상
커먼크롤(Common Crawl) 아카이브의 영어 웹페이지 약 50만 건, 최근 5년치
탐지 도구
오픈 팽그램(Open Pangram)
무작위 표본 결과
2026년 7월 수집 1만 건 중 10%가 AI 저작 의심
챗GPT 출시 이후 필터링 결과
35%가 AI 저작 흔적 (2022년 11월 이후 발행분 한정)
도메인별 격차
.com은 .edu·.gov(각 약 1%) 대비 약 10배, .org는 4.6%
함께 언급된 지표
클라우드플레어, 봇 웹트래픽이 인간 트래픽을 추월했다고 보고

One in Three Web Pages Is Now Written by AI

In a study released Thursday, Pew Research Center found that 35% of English-language web pages published since ChatGPT's debut in November 2022 show signs of AI authorship. That means text written or heavily edited by AI has become more common than text written entirely by humans. The Pew Research report noted that these findings align with conclusions from other recent studies as well.

이미지: TechCrunch AI

Methodology

Pew Research pulled roughly 500,000 English-language web pages spanning the past five years from Common Crawl, an internet archiving project. The collection window began several years before ChatGPT's launch, so it also included pages from an era when AI writing tools didn't yet exist. To determine AI authorship, researchers used detection technology from Open Pangram.

When researchers randomly sampled 10,000 pages collected in July 2026, 10% were flagged as suspected AI-authored. But that sample was diluted by older pages written before AI writing tools existed. When Pew filtered the data to include only pages published after ChatGPT's launch, the figure jumped to 35%.

이미지: TechCrunch AI

A Sharp Divide Across Domains

The rate of AI authorship varied sharply depending on the type of domain. Commercial .com domains showed an AI authorship rate roughly 10 times higher than academic and government domains (.edu and .gov), which each hovered around just 1%. Nonprofit .org domains fell somewhere in between at 4.6%.

DomainSuspected AI Authorship Rate
.com~10% ██████████░ 72
.org4.6% ████░░░░░░░ 33
.edu / .gov~1% █░░░░░░░░░░ 7

The gap is likely explained by the fact that university and government sites tend to have relatively tighter author verification and editorial review processes. Commercial sites chasing ad revenue or search visibility, by contrast, appear to be embracing AI writing tools more aggressively for their ability to scale content production.

An Internet Where Chatbots Read Chatbots

This study pairs neatly with another recent report by shifting the focus from "who is reading the web" to "what is actually written on the web." Internet infrastructure company Cloudflare recently reported that bot-generated web traffic has already overtaken human-generated traffic — sooner than even Cloudflare itself had expected. Taken together, the two studies paint a single picture: a loop in which AI writes web pages, and AI crawlers and bots in turn read those pages, has already taken firm hold.

Pew also pointed to other signals suggesting rising AI authorship, including increased use of em dashes between clauses, higher rates of Oxford comma usage, and phrasing patterns like "It's not X, it's Y" appearing more frequently over time.

A Caveat on the Numbers

AI detection tools, including Pangram, can sometimes misclassify human-written text as AI-generated. Pew acknowledged this limitation but said the overall trend holds up given the scale of the dataset. In other words, the key takeaway isn't any single precise percentage but the broader trajectory: the share of AI-authored content online is rising rapidly.

Editor's Take

What makes this figure alarming isn't the absolute number but the direction it points. Before November 2022, the internet was essentially a repository of human-written text — one that both search engines and AI models drew on as training material. But if more than a third of new content being added to that repository is now AI-written, the next generation of AI models will be trained not on human writing, but on the output of previous-generation AI. A structure in which models feed on other models' output can only erode the diversity of information over time.

This shift is already visible in search results. A few years ago, searching for a specific question would surface answers from personal blogs or specialized forums near the top. Today, those spots are increasingly filled by summary-style pages with strikingly similar sentence structures and tone. The finding that .com domains show an AI authorship rate 10 times higher than .edu or .gov captures this shift precisely — commercial sites chasing search traffic and ad revenue face the greatest pressure to mass-produce content.

The practical lesson for content and media teams is clear. First, any strategy built around producing content at scale for search visibility should now assume that most competitors are already doing so with AI. Second, maintaining .edu/.gov-level credibility will require strengthening editorial review, not cutting it — that becomes the differentiator. Third, running AI detection tools on one's own content matters less than how rigorously a publication backs its work with verifiable sources, evidence, and citations that readers can actually trust.

More studies like this one are likely to emerge in the coming months. Detection tools such as Pangram are also continuing to improve, which should sharpen future figures as false-positive rates decline. And each time new numbers surface, they're likely to intensify the debate over how search engines and AI companies filter the sources of their training data.

Comments