Roughly 35% of web pages published after the November 2022 launch of ChatGPT show significant signs of AI authorship, according to a Pew Research study released Thursday. The firm analyzed nearly half a million English-language pages pulled from the Common Crawl archive, spanning about five years of the open web starting a couple of years before ChatGPT's debut. In a random 10,000-page sample drawn in July 2026 — one that included older pages predating any usable AI writing tools — around 10% carried markers of AI authorship. Filter out the pre-ChatGPT pages and the share of AI-touched content on the newer web jumps to more than a third.
Pew ran its detection using Open Pangram's classifier, which flags pages that were either written by or, in the study's phrasing, substantially edited by AI. The 35% figure covers both categories, so it captures human writers running drafts through a model as well as fully machine-generated text. That framing matters: the study is not claiming that a third of the post-2022 web is pure bot output, but that a third of it bears the fingerprints of a model at some point in the pipeline.
The domain breakdown is where the study gets sharper. Pew found .com URLs showed signs of AI authorship at roughly 10 times the rate seen on .edu and .gov domains, both of which came in near 1%. Non-profit .org pages sat in the middle at 4.6%. The gap suggests commercial publishers — the sites paid by clicks and ad impressions — have moved to AI-assisted content production far faster than universities or government agencies, whose publishing incentives and review processes lean against machine drafts.
Key facts
- 01Pew found 35% of web pages published after ChatGPT's November 2022 release show significant signs of AI authorship.
- 02In a random 10,000-page July 2026 sample including older pages, roughly 10% showed AI authorship markers.
- 03.com domains showed AI authorship at 10x the rate of .edu and .gov domains, both near 1%; .org sat at 4.6%.
- 04Pew analyzed nearly 500,000 English-language pages from the Common Crawl archive spanning about five years.
- 05Detection ran on Open Pangram's classifier; Pew acknowledged the tool can misclassify human-written pages as AI.
The report lands days after Cloudflare said bot web traffic has overtaken human traffic, a threshold the infrastructure provider hit sooner than it had projected. Pew's contribution is on the other side of the transaction: not who is reading the web, but what is being read. Stacked together, the two data points describe a feedback loop in which bots increasingly browse pages that other bots wrote, with human authors and human readers occupying a shrinking middle.
Pew flagged other stylistic tells that have grown more common on post-ChatGPT pages, including heavier use of em dashes, Oxford commas, and rhetorical constructions such as "it's not X, it's Y." These are the kinds of surface features model outputs tend to overuse, and they show up in classifier training data. Whether they are reliable evidence of AI authorship in a specific page is another question — human writers use all three, and some have used them for decades.
Which points to the study's main caveat. Open Pangram, like every AI-detection tool on the market, misclassifies. Human-written pages get flagged as AI. Lightly edited AI pages sometimes pass as human. Pew acknowledged the limitation and framed the results as directionally correct rather than precise. At a sample size of nearly half a million pages, the direction is what matters — but any single page in the dataset could be misclassified.
The findings reinforce what search engines, publishers, and model developers have all been navigating for two years: training data pulled from the open web is increasingly contaminated with model output. That creates a well-documented model-collapse risk for future training runs, and it complicates efforts by companies including Anthropic and OpenAI to build clean corpora. Anthropic's disclosure this month that its Claude watermarks were broken within four hours of launch, which we covered last week, underscores how thin the technical defenses against synthetic-text contamination currently are.
For search platforms and content marketplaces, the .com concentration is the practical headline. If commercial pages are 10 times more likely to be AI-authored than institutional ones, ranking algorithms that penalize thin AI content have a domain-level heuristic they can lean on, and advertisers buying against commercial inventory have a new quality signal to price. Publishers that produce human-written commercial content now have a scarcity argument they did not have two years ago.
The Pew numbers also complicate the ongoing copyright fights over model training. Plaintiffs suing model providers have argued their human-written work was ingested without license. If a growing share of new web content is itself model-generated, the provenance question at the heart of those cases gets harder to litigate — future training crawls will pull in text that no human author can claim.
One open question is whether the 35% figure keeps climbing or plateaus. Detection tools improve, but so does the sophistication of models and the humans prompting them. If AI-assisted writing becomes the default workflow for commercial publishers — the way spellcheck and grammar tools did before it — the useful measurement may stop being "is this AI?" and start being "is this any good?"
Pew's data is the clearest snapshot yet of how quickly generative models moved from novelty to default authoring tool on the commercial web, and the domain split is the tell worth watching. The gap between .com and .gov is not going to close in the government's direction. It is going to widen as the commercial incentive to publish more, faster, at lower cost compounds against the institutional incentive to publish slowly and defensibly. For anyone building search, advertising, or training infrastructure on top of the open web, the underlying corpus is no longer the corpus they were designing for.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




