Site Logo
Commentary on News

AI Authorship on the Web: Qualitative Research Guide

Evidano6 min read

Primary keyword: AI-authored web content. According to the Pew Research Center and reported by Ibtimes.com.au, detectable signs of AI authorship now appear across a large share of recently published web pages. Qualitative researchers, UX teams, and content analysts need practical methods for sampling, detection, and thematic synthesis because changes in authorship affect how we interpret online narratives, discourse trends, and user-facing messaging.

Key Takeaways

According to the Pew Research Center and reported by Ibtimes.com.au, signs of AI authorship appear on 35% of English-language web pages published after ChatGPT's November 2022 release, and a July 2026 random sample of 10, 000 pages showed about 10% with "significant signs of AI authorship."

  • Pew Research Center analyzed nearly half a million English-language pages using the Common Crawl archive and an AI detector called Open Pangram, according to Ibtimes.com.au in August 2026.
  • In a July 2026 random sample of 10, 000 pages, Pew Research Center reported roughly 10% showed "significant signs of AI authorship, " per Ibtimes.com.au.
  • When restricting to pages published after ChatGPT's November 2022 launch, Pew Research Center found signs of AI authorship on 35% of pages in the July 2026 snapshot, according to Ibtimes.com.au.
  • Pew Research Center reported domain differences: .edu and.gov domains registered about 1% AI-signals, .org about 4.6%, and.com domains roughly 10 times the.edu/.gov rate, according to Ibtimes.com.au.

What Happened and How Pew Measured It

Answer: Pew Research Center measured AI authorship signals by analyzing nearly half a million English-language pages from the Common Crawl archive and running text through an AI detector, as reported by Ibtimes.com.au.

Pew Research Center used the Common Crawl web archive to collect a corpus spanning roughly five years, then applied the Open Pangram AI-detection tool to that text, according to Ibtimes.com.au.

Pew Research Center reported that the overall July 2026 snapshot of 10, 000 random pages had about 10% showing "significant signs of AI authorship, " but after excluding pages published before ChatGPT's November 2022 release, the share rose to 35% for post-ChatGPT content, according to Ibtimes.com.au.

Pew Research Center also identified stylistic markers that increased over time, including the use of em dashes, Oxford commas, and characteristic phrasing patterns; Pew wrote, "In the July 2026 snapshot, signs of AI authorship can be found in over one-third of pages published after ChatGPT was released, " according to Ibtimes.com.au.

Findings Snapshot

DateMetricValueImplication
July 2026Random sample size10, 000 pagesBaseline used by Pew Research Center to estimate AI signals, per Ibtimes.com.au.
July 2026Pages with "significant signs of AI authorship" (overall sample)≈10%Indicates detectable AI patterns are present across the sampled web, per Ibtimes.com.au.
Post-November 2022Pages with AI signals (restricted to post-ChatGPT content)35%Pew Research Center reported a much higher AI-signal share when focusing only on content published after ChatGPT's November 2022 launch, per Ibtimes.com.au.
July 2026AI-signal by top-level domain.edu/.gov ≈1%, .org 4.6%, .com ≈10%Domain type strongly correlates with AI-signal prevalence, per Ibtimes.com.au.

Implications for Qualitative Researchers

Answer: Qualitative researchers should treat authorship signals as a sampling and validity variable, not a binary quality label, because Pew Research Center's methodology shows AI contributions are common but often mixed with human editing, as reported by Ibtimes.com.au.

Pew Research Center's finding that 35% of post-ChatGPT pages show AI signals in July 2026 implies researchers must record publication date, domain, and detected AI-signal strength as metadata during web sampling to avoid biased thematic conclusions.

Pew Research Center's method relied on the Open Pangram detector and Common Crawl data, so qualitative teams should triangulate detector outputs with manual coding and source checks because TechCrunch cautioned that detection tools "are not infallible" and may misclassify human writing, according to Ibtimes.com.au.

Pew Research Center's domain-level differences suggest researchers should stratify samples by top-level domain (for example, .com, .org, .edu, .gov) and by publication window (pre- and post-November 2022) to capture adoption dynamics reported by Ibtimes.com.au.

How Evidano Helps

What Evidano is and why it matters here

Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents.

Evidano ingests webpages, transcripts, and spreadsheets, then produces thematic, frequency, and cross-segment analyses that make it practical to track AI-authorship signals as metadata alongside themes and sentiment.

Problem: No scalable way to tag authorship signals in web corpora

Solution: Evidano can ingest Common Crawl extracts or scraped pages and attach detector outputs as item-level metadata for thematic coding and cross-segment comparisons.

Evidano integrates web scraping and document ingestion so teams can combine detector scores with human coding and export codebooks for audit trails; see Evidano features for relevant capabilities.

Problem: Detector uncertainty and mixed human/AI content

Solution: Evidano supports human-over-AI workflows where coders review flagged pages, add contextual notes, and train custom dictionaries so detection outputs feed into reproducible analytic pipelines.

Evidano's AI chat and visualization tools help teams surface recurring AI-associated stylistic markers reported by Pew Research Center, then quantify their prevalence by segment and date.

Problem: Need for secure, auditable research data

Solution: Evidano encrypts data and provides audit-ready exports so researchers can document methods, detector versions, and sampling decisions when publishing replication materials; see Evidano data security.

FAQ: AI-authored web content

How reliable are AI-detection tools for web content?

Answer: AI-detection tools are useful but imperfect and should be interpreted as probabilistic signals rather than definitive labels.

Pew Research Center used the Open Pangram detector and acknowledged limits by relying on stylistic markers, and TechCrunch cautioned that detectors "can occasionally misclassify genuinely human-written pages, " per Ibtimes.com.au.

Best practice is to combine detector scores with manual checks and document detector version and thresholds.

Does the 35% figure mean AI wrote one-third of recent web pages completely?

Answer: No, 35% indicates detectable AI signals, not exclusive AI-only authorship.

Pew Research Center and Digital Trends emphasized that pages often blend human and AI contributions, for example by using AI for drafting, editing, or SEO rewrites, as reported by Ibtimes.com.au.

How should researchers sample the web to study AI influence?

Answer: Researchers should stratify by publication date, top-level domain, and site type, and include both detector scores and manual validation.

Pew Research Center's approach used Common Crawl and a July 2026 random sample of 10, 000 pages, then compared pre- and post-November 2022 content to isolate ChatGPT-era changes, per Ibtimes.com.au.

Can AI-authored pages be included in qualitative analysis?

Answer: Yes, AI-authored or AI-assisted pages can be analytically valuable if researchers code authorship signal as metadata and interpret themes accordingly.

Pew Research Center's findings show AI influence is widespread, so excluding AI-flagged content outright risks creating biased samples; instead, report AI-signal prevalence and analyze differences by signal strength, as suggested by Ibtimes.com.au.

Conclusion & Next Steps

Pew Research Center's analysis, as reported by Ibtimes.com.au, shows that detectable AI signals now appear on a large share of post-November 2022 web pages, with 35% flagged in the July 2026 post-ChatGPT snapshot and domain-level differences that matter for sampling.

Qualitative researchers should record detector outputs as metadata, stratify samples by date and domain, and combine automated detection with manual coding to interpret themes responsibly.

Evidano helps teams operationalize those steps by ingesting web content, attaching detector scores as metadata, supporting human review, and producing reproducible thematic and cross-segment analyses; learn more on our features page.

If you want to pilot an AI-enabled workflow for studying AI-authored web content, Try Evidano for free.

Topics

  • AI-authored web content
  • AI authorship web content
  • AI detection web content
  • qualitative analysis AI content

Keep reading

Browse all articles