Researchers and care teams struggle to know which consultation behaviours actually help patients remember advice. A July 20, 2026 PLOS ONE study (n = 245) found that recounting a validating consultation raised odds of recalling short health messages by ~18.5 (≈19%) per message. This is a practical signal for anyone doing qualitative analysis of clinical consultations: perceived clinician validation is linked to measurable gains in short-term retention. Read on to learn practical steps, how the study was run, where automated coding succeeded, and how to reproduce and scale the same analysis with AI-enabled tools.
Key Takeaways
Perceived clinician validation predicts measurable short-term recall benefits: participants who described a validating consultation had higher odds of recalling brief audio health messages.
- A July 20, 2026 PLOS ONE proof-of-concept experiment (Lee et al.) found ~18.5% higher odds of recalling each of 20 short audio health messages after describing a validating versus an invalidating consultation.
- The study analysed a final sample of n = 245 adults with chronic pain, with data collected Oct 8–30, 2024 and the paper published in PLOS ONE.
- Automated scoring plus human QA worked at scale: a rule-based keyword matcher scored recall with 97.53% agreement versus manual checks, and 12 suspected AI-generated descriptions were removed during cleaning.
Findings snapshot
| Metric | Value | Note / Source |
|---|---|---|
| Published | 20 July 2026 | PLOS ONE |
| Sample size (final) | n = 245 | Data collection Oct 8–30, 2024 |
| Primary outcome | Incidental recall of 20 audio health messages | Scored by rule-based keyword matcher; 97.53% agreement vs manual check |
| Effect size | ~18.5% higher odds per message | Retained after GLMM with participant & item random intercepts |
| Data & code | Open at OSF | OSF |
| Notable QA step | 12 suspected AI-generated descriptions removed | Shows need for AI-detection and manual review |
What happened, methods in plain English
The methods used an autobiographical recall writing task, immediate audio exposure, and an incidental free-recall test to measure downstream memory for short health messages.
The team recruited participants via Prolific and asked them to write about a prior validating or invalidating healthcare consultation, then immediately played 20 short audio health tips and later ran an incidental free-recall test.
Recall was scored per item (1 or 0) using a rule-based keyword algorithm implemented in R, and a 10% random subset was manually double-checked producing 97.53% agreement before full scoring.
The primary inferential model was a binomial generalized linear mixed-effects model (GLMM) with random intercepts for participant and message, and the final model kept condition and pain duration as predictors.
Teams doing qualitative analysis should note the replicable pipeline used here: autobiographical prompts, audio stimuli, automated scoring plus manual validation, and mixed-effects modelling to handle nested item-level responses.
So what for qualitative researchers & UX/health teams?
Design implications
Perceived validation in consultations should be coded and treated as a potential predictor of downstream outcomes such as recall and adherence.
The authors used audio delivery to mirror clinical encounters, therefore capture and analyse audio where possible rather than relying only on transcripts.
Credibility & QA
Combine automated coding with human validation to achieve high agreement, as the study used a rule-based matcher plus a 10% double-check yielding 97.5% agreement.
Screen for low-quality or AI-generated responses early, the study removed 12 suspected AI-generated descriptions during data cleaning.
Analysis & reporting
Model at the item level when outcomes are multiple items per participant, the study used GLMMs to account for participant and message random effects.
Report direct effects and mediation tests, the study found a direct effect of perceived validation on recall and did not find mediation via pain-related fear.
Do more, faster with Evidano
Problem: multimodal, messy inputs
Evidano is an AI-powered qualitative data analysis platform that ingests audio, transcripts, and survey text so teams can work in one secure workspace.
Clinical research mixes audio, transcripts, and open-text survey responses, merging these sources manually is slow and error-prone.
Evidano solution
Evidano ingests audio, transcripts, and spreadsheets in one workspace, offering transcription with custom dictionaries and PII redaction.
Evidano translates and normalises free text across geographies with custom dictionaries so a single codebook applies consistently.
Problem: inconsistent coding & slow validation
Manual coding is costly and introduces drift across coders and waves, which slows reproducible workflows.
Evidano solution
Evidano imports your codebook, runs AI-assisted coding to auto-tag excerpts, and supports targeted human checks (for example, 10% samples) to reproduce the 97%+ agreement workflow used in the paper.
Evidano generates co-occurrence networks and hierarchical code trees to surface validation versus invalidation patterns across segments such as age, pain duration, and country.
Problem: measuring downstream effects
Linking conversational features to outcomes such as recall and adherence requires cross-segment frequency tables and exports ready for mixed-model analysis.
Evidano solution
Evidano runs thematic, frequency, and cross-segment analyses and exports item-level matrices prepared for GLMM or SEM without extra wrangling.
Evidano keeps data encrypted and assures stakeholders that the platform does not use your data to train third-party models.
Checklist: reproduce this study as an AI-enabled qualitative workflow
This checklist lists practical steps to reproduce the study using an AI-enabled qualitative workflow.
- 1) Collect audio of consultations and participant autobiographical prompts, store data securely.
- 2) Transcribe audio with a medical or custom dictionary and run PII redaction.
- 3) Preprocess free text (lowercase, punctuation removal) and build a keyword scaffold or classifier for the 20 target messages.
- 4) Auto-score recall with a rule-based matcher and validate on a random 10% human-coded subset, iterate until agreement exceeds 95%.
- 5) Tag each consultation for perceived validation or invalidation using a mixed AI and human codebook workflow.
- 6) Export item-level data with one row per message per participant and fit binomial GLMMs to estimate odds ratios while accounting for participant and message random effects.
- 7) Visualise code co-occurrence, segment differences, and quote-level evidence for reports or training materials.
FAQ: qualitative analysis of clinical consultations
What did the PLOS ONE study find about validation and recall?
The PLOS ONE study found that perceived clinician validation increased the odds of recalling short audio health messages by about 18.5% per message.
The study used an autobiographical writing task followed by 20 audio tips and measured incidental free recall, with results robust in GLMMs controlling for participant and message effects.
How was recall measured and validated?
Recall was measured per item using a rule-based keyword matcher and validated against manual coding, showing 97.53% agreement on a 10% random subset.
The rule-based matcher in R scored each of the 20 messages as recalled or not recalled and manual checks were used to confirm automated scoring accuracy.
What quality assurance steps did the study use?
The study combined automated scoring with manual validation and removed 12 suspected AI-generated descriptions during cleaning.
The QA process included a 10% human double-check of automated scores and screening for low-quality or AI-generated submissions.
How can teams reproduce the analysis pipeline?
Teams can reproduce the pipeline by collecting audio and autobiographical prompts, transcribing with a custom dictionary, auto-scoring recall, validating a 10% sample, tagging consultations for validation, and exporting item-level data for GLMMs.
The provided checklist lists step-by-step actions from secure data collection through GLMM-ready export and visualisation.
Where can I access the study data and code?
The study data and code are openly available at OSF.
See OSF for the repository linked in the paper.
Conclusion, next steps & CTA
The PLOS ONE study shows perceived clinician validation predicts better short-term recall, about a 19% increase in odds per message.
For teams running qualitative analysis of clinical consultations, combining robust QA, item-level modelling, and multimodal ingestion is the fastest path from data to decision.
Ready to operationalise this pipeline? Try Evidano for free.
If you want a reproducible starter, import your transcripts and run a 2-week pilot using the checklist above; Evidano speeds coding, validation, and export so mixed models and stakeholder reports are ready sooner.
