The PLOS ONE study published July 20, 2026 (n=245) reported that asking adults with chronic pain to recount a validating consultation increased odds of incidental recall for audio health messages by about 18.5%. Full paper: PLOS ONE. If you run qualitative research on consultations, transcripts, or patient narratives, this result matters for coding decisions, sampling, and intervention design.
Key Takeaways
The PLOS ONE study (Jul 20, 2026; n=245) found that asking people with chronic pain to recount a validating consultation increased odds of incidental recall of audio health messages by about 18.5%.
Evidano is an AI-powered qualitative data analysis platform that helps reproduce this mixed-methods workflow by automating transcription, QA, AI-assisted coding, and exporting item-level matrices for GLMMs.
- The study found approximately an 18.5% increase in odds of recalling a given audio item after a validating versus invalidating consultation (PLOS ONE, Jul 20, 2026).
- The analysed sample was n=245 (initial plan n=300), data collected Oct 8–30, 2024 across 16 countries with a majority in the UK.
- Quality control matters: 12 responses flagged as likely AI-generated were removed and manual validation of automated recall scoring showed 97.5% agreement.
- Reproducible mixed-methods workflows should combine item-level binomial GLMMs with predefined qualitative codebooks and transparent QA logs (OSF: OSF).
Findings snapshot
| Date | Sample | Primary outcome | Key result | Notes / Source |
|---|---|---|---|---|
| Jul 20, 2026 | Adults with chronic pain (n=245) | Incidental recall of 20 audio health messages | +18.5% odds of recall after describing a validating vs. invalidating consultation | PLOS ONE |
| Data collection Oct 8–30, 2024 | International (16 countries; majority UK) | Secondary: pain intensity & pain-related fear | Effect not mediated by change in pain-related fear | Binomial GLMM; buildmer; OSF data: OSF |
| Preprocessing | Initial n planned 300; final n=245 | QA checks | 12 responses flagged as AI-generated and removed; manual validation of automated recall scoring (97.5% agreement) | Methodological note relevant for AI-era studies |
What the study did (plain English)
The study quasi-randomly assigned people with chronic pain to write about either a validating or an invalidating past consultation, then played 20 short audio health tips and tested incidental recall.
Researchers modelled recall at the item level using binomial GLMMs, controlling for pain duration, immersion, age, sex and pain measures.
- Design: between-subjects (validation vs. invalidation) with within-subject measures (pre/post PASS-20).
- Primary DV: whether each of 20 audio messages was recalled (1/0).
- Key method wins: item-level GLMMs, simulated power analysis, and manual validation of automated scoring (97.5% agreement).
Why this matters for qualitative researchers
1) Measurement choices change what you can claim
Measurement choices change what researchers can claim about narrative effects on downstream cognition.
The study combined autobiographical texts (participant-written memories) with an audio-based recall task, which raises cross-modality challenges: transcript coding of memories versus scoring of short audio-message recall.
For qualitative teams, predefine how narrative content maps to downstream cognitive outcomes, for example codes for perceived validation, emotional tone, and immersion before analysis.
2) Quality control in the AI era
Quality control in the AI era requires layered automated and human review to protect data integrity.
The PLOS team flagged 12 likely AI-generated descriptions and removed them; online recruitment should therefore build AI-detection and manual review into the pipeline with automated filters, attention checks, and human adjudication.
Evidano can ingest raw transcripts and surface provenance flags and human-in-the-loop review logs to help trace and remove low-quality or AI-like responses before coding.
3) Effect sizes are real but modest
Effect sizes in this context are detectable but modest at the individual level.
An approximately 19% increase in odds of recalling a single item is meaningful at scale but small per patient, suggesting validating behaviours are detectable as shifts in narrative content and retention but should not be over-claimed.
Mixed-methods are recommended: use qualitative themes to define mechanisms and quantitative item-level models to estimate effect magnitude.
Do more, faster with Evidano
Ingest & QA
Ingesting mixed inputs requires an integrated pipeline for audio, transcripts, and survey scales.
Evidano can ingest audio and transcripts, apply automated transcription with a custom dictionary for medical terms and clinician phrases, and run PII redaction before coding.
Detect low-quality / AI-like responses
Detecting low-quality or AI-like responses needs automated pattern flags plus human adjudication.
Evidano can flag patterned text for human review, surface attention-check failures, and export adjudication logs for audit trails so teams can replicate QA transparency like the PLOS study.
Code & synthesize at scale
Scaling coding without losing consistency requires AI-assisted suggestions plus targeted human review.
Evidano allows importing a codebook, running AI-assisted coding, and reviewing suggested codes, then producing thematic frequency tables, co-occurrence networks, and hierarchical code-to-subcode maps to show how 'validation' clusters with emotion, trust, and action orientation.
Cross-segment and item-level analysis
Retaining item-level granularity permits direct modelling of recall and avoids loss of signal from aggregation.
Evidano can compute frequency and cross-segment analyses (for example validation vs. invalidation, pain-duration buckets) and export item-level matrices that feed directly into GLMMs or mixed-methods reports.
Secure research-grade environment
Research-grade data governance requires encrypted storage and controls against third-party model training.
Evidano provides encrypted storage and proprietary models tuned for qualitative research; the platform does not use your data to train third-party models, which aligns with academic best-practices for data governance and reproducibility.
Checklist: Reproduce this study as mixed-methods in 7 steps
Follow these seven steps to reproduce the PLOS workflow with rigorous QA and AI assistance.
- 1) Design: predefine sampling, attention checks, and item-level outcomes (20 messages → binary recall per item).
- 2) Collect: audio-record messages; collect autobiographical texts with minimum word/character limits and enforced time-on-task.
- 3) Transcribe: use AI transcription with a custom medical dictionary; run PII redaction.
- 4) QA: run automated flags for AI-generated text, attention checks, and manual adjudication (retain logs).
- 5) Code: import a codebook and run AI-assisted coding; manually review a random 10% for agreement (target ≥95%).
- 6) Analyse: export item-level recall matrices and covariates; run binomial GLMMs (random intercepts for participant & item) and complementary qualitative theme prevalence analysis.
- 7) Report: include effect sizes, confidence intervals, QA flow (n recruited → n analysed → exclusions), and share code/data when possible (OSF: OSF).
FAQ: Qualitative analysis of clinical validation
How do I code 'validation' reliably?
Define 'validation' with explicit behavioral and affective indicators and example-driven codebook entries.
Create example-driven codebook entries that include behavioral indicators (for example explicit acknowledgment, belief statements, empathic summaries) and run double-coding on a subset to set agreement thresholds before full coding.
Can AI help without introducing bias?
AI can help speed transcription and suggest codes, but humans must make final labeling decisions on sensitive concepts.
Use AI to suggest codes and to speed transcription, keep humans in the loop for final labels (especially for concepts like validation versus invalidation), and maintain audit logs for each AI suggestion to trace decisions.
Is participant narrative authenticity verifiable?
Participant narrative authenticity is verifiable only with layered QA including attention checks, timing filters, manual review, and AI-pattern flags.
Replicate the PLOS team's transparency by removing flagged cases and documenting the QA pipeline, including attention-check outcomes and manual adjudication records.
How do I combine qualitative themes with GLMMs?
Combine qualitative themes with GLMMs by transforming theme presence into predictors or moderators at the item or participant level.
Transform theme presence into binary or frequency predictors and include them as fixed effects or moderators in GLMMs, mirroring the PLOS item-level modelling approach.
Wrapping up: two immediate moves
If you analyse consultation narratives or patient advice retention, treat 'validation' as a measurable code and plan item-level analyses for downstream outcomes.
Try a small pilot ingest in Evidano to automate transcription, apply codebooks, and produce thematic plus cross-segment reports you can feed into GLMMs, and Try Evidano for free.
