Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents. According to HealthsystemCIO.com on August 19, 2026, Penn Medicine’s Susan Harkness Regli, PhD, describes a staged human factors approach that caught a risky AI behavior during beta testing. This post explains the practical steps Penn used, the measurable signals a qualitative researcher should watch for, and how AI-enabled qualitative research tools can accelerate safe adoption. The primary keyword for this post is AI usability testing in healthcare and the intended audience is clinical researchers, usability teams, and clinical IT leaders who must evaluate AI tools quickly and safely.
Key Takeaways
According to HealthsystemCIO.com on August 19, 2026, Penn Medicine used staged alpha and beta exposure plus active monitoring to catch an AI-generated patient-message workflow that some clinicians accepted without edits.
- Penn ran a beta group of 25 to 50 users per role in its staged testing, according to HealthsystemCIO.com on August 19, 2026.
- Penn discovered the risk during beta monitoring in August 2026 and prevented the tool from reaching the wider clinician population, according to HealthsystemCIO.com on August 19, 2026.
- Penn is consolidating seven hospitals onto a single EHR instance, which the organization used to surface duplicative workflows, according to HealthsystemCIO.com on August 19, 2026.
What Happened: AI usability testing in healthcare
The direct answer: Penn Medicine staged AI exposure and monitored usage data to identify usability risk before broad rollout.
According to HealthsystemCIO.com on August 19, 2026, Penn first exposed new AI tools to an alpha group of clinical and informatics champions, then to a beta cohort, before wider deployment.
According to HealthsystemCIO.com on August 19, 2026, the beta cohort size used at Penn is 25 to 50 people per role, with surveys between stages and continuous usage data collection.
According to Susan Harkness Regli, PhD, on the HealthsystemCIO.com podcast, "If we stop every time a new tool comes out and do a major evaluation of it, we’re not going to keep up, " which summarizes Penn’s lightweight, staged evaluation approach.
Findings Snapshot
| Date | Metric | Value | Implication |
|---|---|---|---|
| August 19, 2026 | Article publication | HealthsystemCIO.com coverage | Describes Penn’s staged AI evaluation approach |
| August 2026 | Beta cohort size | 25 to 50 users per role | Small, role-specific testing surfaces behavioral patterns |
| August 2026 | EHR consolidation | 7 hospitals onto one instance | Migration reveals duplicative workflows and integration issues |
Implications for Clinical IT and Research Teams
The practical answer: teams must instrument and observe AI tools in small, role-specific cohorts and require vendor data access before approval.
According to HealthsystemCIO.com on August 19, 2026, Penn now requires vendors to provide a stable test environment and expose sufficient telemetry so the health system can monitor usage and safety signals.
According to Susan Harkness Regli on HealthsystemCIO.com, Penn treats an absence of user edits as a finding rather than a success, because users can carry assumptions from the EHR into an AI tool that uses different premises.
Clinical IT leaders should set a return-review date at approval and name which data streams will be checked at that time, a practice Penn recommends on HealthsystemCIO.com on August 19, 2026.
How Evidano Helps
Problem: Slow synthesis of staged usability feedback
Answer: Rapid thematic and frequency analysis is required to turn alpha and beta notes into actionable fixes.
Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents.
Evidano feature mapping: ingest stage surveys, think-aloud transcripts, and vendor logs, then produce thematic summaries, frequency counts, and exemplar quotes to prioritize fixes quickly. See the features page for capabilities and workflows.
Problem: Missing or inconsistent user edit signals
Answer: Qualitative signals like 'no edits' require context that combines usage telemetry with user explanations.
Evidano feature mapping: merge usage spreadsheets with open-text survey responses to cross-segment where roles accept outputs unchanged, then surface co-occurrence networks and representative verbatim quotes for designers and governance.
Problem: Vendor data access and traceability
Answer: Governance needs reproducible evidence to enforce vendor requirements and post-approval reviews.
Evidano feature mapping: ingest vendor test-environment transcripts and repeated simulation results, redact PII automatically during transcription, and create timestamped reports that link governance checkpoints to the underlying evidence.
FAQ: AI usability testing in healthcare
How large should a beta cohort be for AI usability testing in healthcare?
Answer: A practical beta cohort size is 25 to 50 users per role.
According to HealthsystemCIO.com on August 19, 2026, Penn Medicine uses a beta group of 25 to 50 people per role as part of its staged evaluation.
Choose cohort size to balance statistical signal in usage metrics with the agility to iterate on interface and permission changes.
What signals indicate an AI tool is unsafe or unusable in clinical practice?
Answer: Signals include zero edits from a role, unexpected workarounds, and mismatches between tool assumptions and clinician mental models.
According to HealthsystemCIO.com on August 19, 2026, Penn found a role that edited nothing in a patient-message tool, prompting targeted changes to permissions and interface design.
Pair these qualitative signals with telemetry and follow-up interviews to determine root cause.
How often should governance revisit approved AI tools?
Answer: Governance should set a named return date at approval and specify which data will be reviewed at that time.
According to HealthsystemCIO.com on August 19, 2026, Penn’s approach includes setting a return review and asking in advance 'what data are we going to look at in six months, ' a question Susan Harkness Regli described on the HealthsystemCIO.com podcast.
Regular, data-driven reviews detect drift in usage, safety signals, and conceptual integration problems.
Conclusion & Next Steps
Human factors work and staged exposure are practical ways to reduce AI usability risk in hospitals, as described by Penn Medicine on HealthsystemCIO.com on August 19, 2026.
Clinical teams should require vendor test environments, monitor small beta cohorts of 25 to 50 users per role, and set explicit post-approval review dates.
Operational researchers can accelerate synthesis by using AI-enabled qualitative tools that merge telemetry and transcripts into prioritized recommendations.
To try these workflows yourself, Try Evidano for free and see how staged-test feedback becomes actionable evidence.
Topics
- AI usability testing in healthcare
- human factors AI
- clinical AI usability
- AI tool evaluation healthcare
Keep reading
- Commentary on NewsFaster Photovoice Qualitative Analysis with AIHow AI accelerates photovoice qualitative analysis: extract themes, frequencies, and policy-ready insights from community photos and narratives. Learn methods and try Evidano.
- Commentary on NewsAI-assisted qualitative analysis for museum observationsHow AI-enabled qualitative analysis accelerates coding and insight from museum visitor observations. Learn methods and stats from PLOS ONE, and try Evidano.
- Commentary on NewsAI qualitative insights: epilepsy medication adherenceAI-enabled qualitative analysis of epilepsy medication adherence in Uganda: 277 patients, 66.5% adherence, and actionable insights for researchers and care teams.
