TECH
Pew Research Center: AI-generated content on the web exceeds 35% of recent pages
The presence of AI-generated content on the web surpassed 35% on pages published after the launch of ChatGPT, according to a Pew Research Center study released on August 20, 2026. The analysis evaluated the evolution of algorithm-synthesized text on the English-language internet between January 2021 and July 2026. The results demonstrate artificial intelligence's transition from a niche tool to one of the primary drivers of text production within the digital ecosystem.
Across the entire active sample collected in July 2026, the average proportion of pages showing signs of automation stands at 10%. This sample includes the historical archive from before 2022, which lowers the overall average. However, isolating data from the period following the market arrival of large language models reveals a continuous acceleration in the rate of automated production.
The Pew Research Center study on AI confirms that more than one-third of recent text production on the internet involves direct algorithmic intervention. The investigation gathered 490,000 web pages from the public Common Crawl repository, divided into 49 monthly samples of 10,000 pages each.
The team of data scientists submitted the texts to the open-weights model *editlens_Llama-3.2-3B*, developed by the organization Open Pangram. The classifier assigned each document a score between zero and one based on statistical probabilities regarding authorship. Pages with a score of 0.2 or higher were categorized as showing signs of algorithmic drafting or substantial editing.
Samuel Bestvater, a senior data scientist at the Pew Research Center and the study's lead author, stated in the report that "artificial intelligence models learn to mimic human writing patterns and may end up using certain words, phrases, or linguistic quirks more frequently than human authors do." The research center acknowledges margins of error in individual texts but maintains the statistical robustness of the aggregate sample collected over five years.
Distribution by domain and prevalence of AI-written web pages...The rate of AI-generated content on the web is concentrated primarily in commercial .com domains, reaching 9.4% of all active pages in 2026. In contrast, institutional domains maintain negligible levels of automation.
The distribution of AI-written web pages varies according to the purpose of each internet address:
.com domains: Record 9.4% of pages showing signs of AI in 2026, following a steady rise from the 1.09% observed in 2021.
.org domains: Show a rate of 4.6%, a figure corresponding to half the incidence observed on commercial platforms.
.edu domains: Maintain a presence of 1.0%, with minimal variation from the 0.57% recorded in early 2021.
.gov domains: Stand at 0.8%, the lowest percentage among all analyzed categories.
Pressure to generate mass traffic and optimize for search engines explains the acceleration observed in the commercial sector. Conversely, government and higher education platforms maintain formal human review processes that limit the direct publication of algorithmic text.
AI text detection relies on identifying statistical deviations in punctuation, vocabulary, and syntax. The proliferation of generative models altered the very formal structure of text available on the internet between 2023 and 2026.
Data from the Pew Research Center identifies four key changes in online writing:
Explanatory dashes: The frequency of this punctuation mark rose from 5.79 to 11.19 occurrences per 10,000 words between 2023 and 2026.
Oxford comma: The use of the comma before the final conjunction in lists saw a 63% increase during the same period.
Characteristic vocabulary: English terms such as *delve*, *interplay*, *tapestry*, *pivotal*, *testament*, *underscore*, and *intricate* doubled their rate of occurrence across the open web.
Negative rhetorical parallelism: Comparative structures such as “it is not just X, it is Y” nearly tripled in statistical incidence.
These elements do not constitute definitive proof of AI-generated writing on their own, as they are also part of human language usage. However, their cumulative density allows classifiers to calculate the mathematical probability of algorithmic intervention in a document.
Structural risks and the impact of artificial intelligence on the internet...The impact of artificial intelligence on the internet poses challenges to preserving the integrity of public data and to the development of future computing systems. The primary technical risk identified by experts is "model collapse."
The expansion of AI-generated content on the web means that future generations of language models will ingest synthetic text as primary training material. This closed feedback loop can amplify factual errors, reduce lexical diversity, and lead to a gradual degradation in the quality of responses from digital assistants.
The proliferation of automated text also alters search engine indexing strategies. The saturation of the web with low-cost articles necessitates stronger quality filters to distinguish original sources from mechanical syntheses.
A cycle of self-pollution...This is not merely a story about "declining content quality"; it strikes at the very foundations of the AI industry.
In 2024, *Nature* published a paper by a research team from Oxford and Cambridge proving that AI models suffer from "model collapse" during recursive training (using AI-generated data to train the next generation of AI): outputs gradually drift away from the real-world data distribution, lose rare patterns found in the distribution's tails, and—after several iterations—generate increasingly homogeneous and even nonsensical content. Researchers at Epoch AI predict that the supply of high-quality, human-authored text suitable for training AI could be exhausted between 2026 and 2032.
The cycle can already be described: AI models read the human internet and generate new web pages in bulk; these new pages are indexed by search engines and archived by crawling tools like Common Crawl; and next-generation AI models continue to be trained on this data. With each cycle, the concentration of "AI-writing-for-AI" content in the training data increases, while signals derived directly from human experience are diluted.
If the internet's signal-to-noise ratio continues to deteriorate, content creators who rely on search traffic will be the first to feel the impact.
Search results are becoming saturated with homogeneous, AI-rewritten content, making it increasingly difficult for readers to distinguish between the "original" and the "copy." Google is already replacing some click-through links in search results with "AI Overviews," and a Pew study from July of this year found that only 20% of users consider AI-generated search summaries "very useful," while only 6% deemed them "very trustworthy." As AI-generated rewrites become ubiquitous, the media outlets that provide content with "scarcity value" will be those offering: firsthand information gathered on-site; exclusive data and documents; opinions and judgments traceable to specific human sources; and a distinct editorial voice that readers are willing to pay for.
From the flip side, the metrics used to measure a media outlet's competitiveness are quietly shifting.
The volume of published articles is no longer a barrier to entry; AI can generate a thousand articles a day. The real barrier is "verifiable human authorship": was the information in the article gathered by journalists or created by the model? Was the content shaped by industry-experienced editors or merely by a combination of prompts?
In an internet landscape where 35% of new web pages show traces of AI, the ability to answer that question effectively represents a competitive advantage.
mundophone
No comments:
Post a Comment