Ethics

More than one-third of new English webpages contain AI-written text, Pew analysis finds

A Pew Research Center analysis of 500,000 webpages published since ChatGPT's launch found machine-written text in over one-third of new English web content, with commercial sites leading adoption. The findings raise questions about content authenticity, detection accuracy, and the need for clearer disclosure standards.

By Michael C ·

More than one-third of new English webpages contain AI-written text, Pew analysis finds
SUPERBASH_ editorial image.

More than one-third of English webpages published since the launch of ChatGPT contain machine-written text, according to analysis by the Pew Research Center. The finding, based on examination of approximately 500,000 webpages, signals a significant shift in how online content is produced and distributed at scale. The report does not claim these pages were created entirely without human involvement. Researchers acknowledged the uncertainty inherent in detecting machine-written text through automated tools, and the one-third figure represents a snapshot based on current detection capabilities rather than a definitive count.

The analysis found substantial variation across different types of websites. Commercial .com domains were far more likely to incorporate AI-generated text than government or academic sites. The pattern suggests that cost reduction and content production speed may be driving adoption in the commercial sector, while institutional sites maintain stricter editorial standards or face different regulatory pressures. The Pew researchers did not provide a breakdown of specific industries or business models most heavily reliant on the technology, leaving open questions about which sectors are most affected. The Decoder report documents the reporting behind this account.

Detection of machine-written text remains technically uncertain. The Pew analysis used automated detection methods, which have known limitations in distinguishing between human and machine authorship, particularly as language models become more sophisticated. False positives and false negatives are both possible. According to The Decoder report, the scale of adoption has accelerated dramatically since late 2022, outpacing the development of reliable detection standards. Pew Research Center offers useful technical background for evaluating the claim.

Transparency and accountability

The proliferation of machine-written text on the web raises fundamental questions about transparency and public trust in online information. Most websites containing detected AI text do not disclose this fact to readers. No enforced standard exists requiring publishers to label AI-generated content, leaving audiences without clear signals about authorship or editorial process. This absence of disclosure creates potential problems for readers trying to assess source credibility, journalists evaluating reference material, and regulators seeking to understand the composition of online information ecosystems. The operational tradeoff is also reflected in Common Crawl.

Distribution of AI-written text detection across website domains and publishing categories in Pew Research Center sample. Image: SUPERBASH_.
Distribution of AI-written text detection across website domains and publishing categories in Pew Research Center sample. Image: SUPERBASH_.

Institutional accountability mechanisms have not kept pace with adoption. Publishers, platforms, and technology vendors have few clear obligations to track or report AI use in content production. Industry groups like C2PA have proposed technical standards for content provenance, but adoption remains limited and voluntary. Without mandatory disclosure or third-party auditing, the public has limited ability to assess how widespread the practice has become beyond academic research samples. This opacity affects trust in specific publications and potentially in digital information more broadly.

Regulatory landscape and policy gaps

Regulators and policymakers are beginning to address AI-generated content, though approaches remain fragmented. The NIST AI Risk Management Framework and OECD AI principles establish governance concepts for transparency and accountability, but implementation at the content level remains underdeveloped. Some jurisdictions are exploring disclosure requirements for AI-generated text in specific contexts like advertising or government communications, but no comprehensive framework exists for the broader web. The speed of adoption documented by Pew outpaces the development of policy responses.

Timeline of AI-written text detection growth across new English webpages following ChatGPT public launch. Image: SUPERBASH_.
Timeline of AI-written text detection growth across new English webpages following ChatGPT public launch. Image: SUPERBASH_.

The gap between technological deployment and regulatory clarity has real consequences. Content creators face uncertainty about best practices. Publishers making editorial decisions lack consistent guidance. Readers cannot easily determine whether text they are reading was produced by humans, machines, or some combination. This opacity complicates the work of researchers, archivists, and systems like Common Crawl that attempt to preserve and understand the web as a historical record. The variation between commercial and institutional sites suggests that market incentives and institutional norms are driving divergent outcomes. Whether voluntary industry standards, regulatory requirements, or some combination of both will emerge as the primary mechanism for governing this content remains to be determined. The final point can be checked against OECD AI principles.

Topics: AI, content authenticity, internet governance, media accountability, digital ethics