One in 10 English-language webpages now exhibits significant signs of artificial intelligence authorship, according to a Pew Research Center study published Aug. 22. Among pages published after ChatGPT's November 2022 launch, that share rises to 35 percent.
Researchers analyzed approximately 490,000 pages from the Common Crawl web archive spanning January 2021 through July 2026, using Open Pangram, an AI detection model developed by Pangram Labs, to assess statistical patterns in text.
The disparity across top-level domains is stark. Commercial.com domains show AI indicators at roughly 10 times the rate of.edu or.gov domains, which each register near 1 percent.org sites fall between, at 4.6 percent. In 2021, all four categories showed nearly identical AI presence. By January 2026,.com sites reached 9.35 percent—a tenfold increase over five years—while.edu and.gov remained flat.
The shift correlates with measurable linguistic changes. Since 2023, em dashes have doubled in frequency, Oxford commas increased 63 percent, and words like "delve," "interplay," and "testament" more than doubled. Negative parallelism—the construction "it is not just X, it is Y"—tripled, though Pew noted it remains rare overall.
Pew Research acknowledged detection limitations. AI models can misclassify individual pages and cannot distinguish full machine generation from human text enhanced by AI assistance. Open Pangram has also flagged AI content in roughly 9 percent of U.S. newspaper articles, including opinion pages from outlets such as The New York Times, indicating the phenomenon extends beyond commercial websites.
The.com surge reflects structural publishing economics. Academic and government sites typically undergo rigorous editorial review and institutional sign-off on slower timelines. Commercial domains span newsrooms to affiliate-marketing farms—operations generating vast page volumes that human editors cannot sustain without AI assistance.
Future detection may improve through technology advances rather than statistical inference alone. Anthropic and other major AI companies are developing model-level text fingerprinting to make AI-generated content identifiable and reduce misclassification errors.
