Site Logo
All articles
Commentary on News

Faster Insights: qualitative analysis of AI health content

Evidano7 min read

This introduction summarizes the June 4, 2026 PLOS ONE study that compared ChatGPT (GPT-5.2), Gemini 3 Flash and Perplexity (Sonar-4) on myofascial pain syndrome queries and found 18 queries produced 54 AI responses that were all above a 6th-grade reading level (p < 0.001). The PLOS ONE authors used Google Trends (March 1, 2026) to select the keywords, ran each query in clean sessions against three free LLM versions, and archived the first responses for analysis. If you run qualitative analysis of AI health content, the PLOS ONE paper shows exactly what to measure and why: readability indices (FRES, FKGL, GFOG, CLI, ARI, SMOG) and quality/reliability scales (GQS, EQIP, Modified DISCERN, JAMA). The open dataset is available on Figshare and the study is published in PLOS ONE.

Key Takeaways

Evidano is an AI-powered qualitative data analysis platform that helps researchers reproduce and scale the PLOS ONE checks on AI health content, from ingestion through readability and quality scoring to stakeholder reports.

The PLOS ONE study (June 4, 2026) found that 18 queries produced 54 AI responses that were all above a 6th-grade reading level (p < 0.001), with Perplexity scoring highest on quality metrics and ChatGPT producing the clearest text.

  • The PLOS ONE corpus: 18 queries produced 54 responses (18 × 3 models), derived from a Google Trends snapshot on March 1, 2026.
  • Readability risk: all model outputs were significantly more complex than a 6th-grade reading level (p < 0.001) according to multiple indices.
  • Quality vs readability: Perplexity led on JAMA, DISCERN, GQS and EQIP scores, while ChatGPT led on linguistic clarity, demonstrating a readability‑vs‑reliability trade-off.
  • Reproducibility: two independent raters produced inter-rater ICCs between 0.761 and 0.999, indicating good to excellent agreement.

Findings snapshot

Date / MetricValueSourceImplication
PublishedJune 4, 2026PLOS ONERecent comparative evaluation of major LLMs on medical content
Queries / Outputs18 unique English queries → 54 AI responsesPLOS ONEManageable corpus size for reproducible qualitative coding
Models evaluatedChatGPT (GPT-5.2), Gemini 3 Flash, Perplexity (Sonar-4)PLOS ONECompare readability vs. reliability trade-offs
Key quantitative resultAll responses > 6th-grade reading level (p < 0.001)Readability tests: FRES, FKGL, GFOG, CLI, ARI, SMOGAccessibility risk for patients with low e-health literacy
Quality & reliabilityPerplexity scored highest on JAMA, DISCERN, GQS, EQIPPLOS ONESource-linked validation boosts academic reliability
Readability leaderChatGPT produced the clearest text (lowest linguistic complexity)PLOS ONEFluent text can create fluency bias if accuracy is low
Open datasetRaw prompts and responses available on FigshareFigshareDirectly import into Evidano for replication

What the study did (plain English)

The PLOS ONE study selected 18 high-relevance English keywords about myofascial pain syndrome using Google Trends as of March 1, 2026, ran each query in clean sessions against three free LLM versions and saved the first response for analysis.

The PLOS ONE authors measured readability with six indices (FRES, FKGL, GFOG, CLI, ARI, SMOG) and measured quality and reliability with GQS, EQIP, Modified DISCERN and JAMA scores; two independent raters scored reliability and quality and the study reported Kruskal–Wallis, Mann–Whitney U and Wilcoxon tests plus ICC for rater agreement.

  • Corpus: 54 responses (18 × 3 models).
  • Readability: All models produced outputs significantly more complex than 6th-grade (p < 0.001).
  • Quality vs readability: Perplexity > ChatGPT on quality metrics; ChatGPT > Perplexity on readability.
  • Inter-rater reliability: ICCs ranged from 0.761 to 0.999 (good to excellent).

So what for researchers, UX and policy teams

For qualitative researchers & UX teams

Qualitative researchers and UX teams should not assume fluent prose equals accuracy: high readability can amplify misinformation through fluency bias.

Qualitative researchers and UX teams should run reproducible checks that combine readability indices and citation/reliability scales alongside thematic coding.

Qualitative researchers and UX teams should segment comparisons by model and query type (symptom, diagnosis, treatment) to spot systematic gaps.

For clinicians & health policy analysts

Clinicians and health policy analysts should treat LLM outputs as secondary advisory material requiring clinician-in-the-loop validation, as recommended by the PLOS ONE paper.

Clinicians and health policy analysts should monitor accessibility and plan plain-language edits if patient-facing AI content is above a 6th-grade level.

Clinicians and health policy analysts should require transparent citations and disclaimers in any AI-assisted patient materials.

Do more, faster with Evidano

Ingest the raw corpus (easy replication)

Evidano ingests the PLOS ONE study's Figshare dump or saved chatbot responses directly so teams can reproduce the 18×3 dataset in minutes.

Import the PLOS ONE open dataset from Figshare or paste saved chatbot logs into Evidano as documents or spreadsheets.

Automate the readability & quality layer

Evidano computes FRES, FKGL, Gunning Fog, SMOG and other indices across a corpus and visualizes distributions by model and query type.

Evidano maps GQS, DISCERN, JAMA and EQIP as coded variables so teams can compare models using cross-segment analysis.

Thematic and reliability synthesis

Evidano's AI-assisted coding extracts themes (trigger points, treatments, diagnostic language) and nests them into hierarchies for fast memoing.

Evidano runs co-occurrence networks and quote matrices to show where high readability coexists with poor citation practices.

Stakeholder outputs and governance

Evidano generates slide-ready reports and exportable dashboards (word clouds, hierarchical codes→subcodes) for clinicians, regulators, and UX stakeholders.

Evidano stores data with encryption and does not use customer data to train external models, supporting clinician-in-the-loop governance as advised in the PLOS ONE paper.

Checklist: 7-step workflow to reproduce the study and extend it

This seven-step checklist reproduces the PLOS ONE study and extends it to larger corpora and languages.

Step 1: Pull keywords via Google Trends and freeze the list, the PLOS ONE paper used a March 1, 2026 snapshot.

Step 2: Capture first-response outputs from each model in clean sessions and document the model versions (GPT-5.2, Gemini 3 Flash, Perplexity Sonar-4).

Step 3: Import responses into Evidano as documents or CSV files.

Step 4: Run an automated readability suite (FRES, FKGL, GFOG, CLI, ARI, SMOG) and export summary tables.

Step 5: Apply reliability and quality codes (DISCERN, JAMA, GQS, EQIP) as structured annotations inside Evidano and compute ICC between raters.

Step 6: Perform thematic coding and co-occurrence analysis to map misinformation risk areas, for example treatment advice without citation.

Step 7: Produce a stakeholder brief with recommended clinician-in-the-loop edits and plain-language rewrites, and publish the reproducible dataset and methods as the PLOS ONE authors did.

FAQ: qualitative analysis of AI health content

Can I trust readability scores alone?

No, readability scores alone are not sufficient because readability measures accessibility but not factual accuracy.

The PLOS ONE study demonstrates that ChatGPT produced the clearest text but was not the highest on reliability metrics, so pair readability indices with reliability scales (DISCERN, JAMA) and source checks.

How do I compare segments (model × query type)?

Tag each response by model and query category and then run cross-segment analyses to detect patterns.

In Evidano, tag responses by model and query type and run frequency and thematic analyses to compare treatment queries, symptom queries and other segments for systematic differences.

Is this ethical for patient-facing use?

No, AI outputs should be treated as research or advisory content and must be validated by clinicians before patient-facing use.

The PLOS ONE authors recommend clinician oversight, disclaimers and transparent citations to reduce safety risk when AI-generated text is used in health contexts.

Wrapping up: next steps

The next steps are to replicate the PLOS ONE checks (June 4, 2026) or scale them to larger corpora and languages using Evidano's pipeline from ingestion to stakeholder-ready reports.

  • Import the open Figshare dataset or your own chatbot logs and run the seven-step workflow above.
  • Use Evidano to operationalize clinician-in-the-loop validation and produce reproducible reports that meet audit needs.
  • Security note: Evidano encrypts customer data and does not use it to train third-party models, which is important for sensitive health research.

Ready to move from anecdote to evidence? Try Evidano for free or import the PLOS ONE Figshare dump to reproduce the study in hours, not weeks.

Company
About
Newsletter

Product updates, research, and tips — straight to your inbox.

© Evidano, All Rights Reserved.