Site Logo
Commentary on News

Human Factors: AI Usability Testing in Healthcare

Evidano5 min read

Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents. According to HealthsystemCIO.com on August 19, 2026, Penn Medicine’s Susan Harkness Regli, PhD, describes a staged human factors approach that caught a risky AI behavior during beta testing. This post explains the practical steps Penn used, the measurable signals a qualitative researcher should watch for, and how AI-enabled qualitative research tools can accelerate safe adoption. The primary keyword for this post is AI usability testing in healthcare and the intended audience is clinical researchers, usability teams, and clinical IT leaders who must evaluate AI tools quickly and safely.

Key Takeaways

According to HealthsystemCIO.com on August 19, 2026, Penn Medicine used staged alpha and beta exposure plus active monitoring to catch an AI-generated patient-message workflow that some clinicians accepted without edits.

  • Penn ran a beta group of 25 to 50 users per role in its staged testing, according to HealthsystemCIO.com on August 19, 2026.
  • Penn discovered the risk during beta monitoring in August 2026 and prevented the tool from reaching the wider clinician population, according to HealthsystemCIO.com on August 19, 2026.
  • Penn is consolidating seven hospitals onto a single EHR instance, which the organization used to surface duplicative workflows, according to HealthsystemCIO.com on August 19, 2026.

What Happened: AI usability testing in healthcare

The direct answer: Penn Medicine staged AI exposure and monitored usage data to identify usability risk before broad rollout.

According to HealthsystemCIO.com on August 19, 2026, Penn first exposed new AI tools to an alpha group of clinical and informatics champions, then to a beta cohort, before wider deployment.

According to HealthsystemCIO.com on August 19, 2026, the beta cohort size used at Penn is 25 to 50 people per role, with surveys between stages and continuous usage data collection.

According to Susan Harkness Regli, PhD, on the HealthsystemCIO.com podcast, "If we stop every time a new tool comes out and do a major evaluation of it, we’re not going to keep up, " which summarizes Penn’s lightweight, staged evaluation approach.

Findings Snapshot

DateMetricValueImplication
August 19, 2026Article publicationHealthsystemCIO.com coverageDescribes Penn’s staged AI evaluation approach
August 2026Beta cohort size25 to 50 users per roleSmall, role-specific testing surfaces behavioral patterns
August 2026EHR consolidation7 hospitals onto one instanceMigration reveals duplicative workflows and integration issues

Implications for Clinical IT and Research Teams

The practical answer: teams must instrument and observe AI tools in small, role-specific cohorts and require vendor data access before approval.

According to HealthsystemCIO.com on August 19, 2026, Penn now requires vendors to provide a stable test environment and expose sufficient telemetry so the health system can monitor usage and safety signals.

According to Susan Harkness Regli on HealthsystemCIO.com, Penn treats an absence of user edits as a finding rather than a success, because users can carry assumptions from the EHR into an AI tool that uses different premises.

Clinical IT leaders should set a return-review date at approval and name which data streams will be checked at that time, a practice Penn recommends on HealthsystemCIO.com on August 19, 2026.

How Evidano Helps

Problem: Slow synthesis of staged usability feedback

Answer: Rapid thematic and frequency analysis is required to turn alpha and beta notes into actionable fixes.

Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents.

Evidano feature mapping: ingest stage surveys, think-aloud transcripts, and vendor logs, then produce thematic summaries, frequency counts, and exemplar quotes to prioritize fixes quickly. See the features page for capabilities and workflows.

Problem: Missing or inconsistent user edit signals

Answer: Qualitative signals like 'no edits' require context that combines usage telemetry with user explanations.

Evidano feature mapping: merge usage spreadsheets with open-text survey responses to cross-segment where roles accept outputs unchanged, then surface co-occurrence networks and representative verbatim quotes for designers and governance.

Problem: Vendor data access and traceability

Answer: Governance needs reproducible evidence to enforce vendor requirements and post-approval reviews.

Evidano feature mapping: ingest vendor test-environment transcripts and repeated simulation results, redact PII automatically during transcription, and create timestamped reports that link governance checkpoints to the underlying evidence.

FAQ: AI usability testing in healthcare

How large should a beta cohort be for AI usability testing in healthcare?

Answer: A practical beta cohort size is 25 to 50 users per role.

According to HealthsystemCIO.com on August 19, 2026, Penn Medicine uses a beta group of 25 to 50 people per role as part of its staged evaluation.

Choose cohort size to balance statistical signal in usage metrics with the agility to iterate on interface and permission changes.

What signals indicate an AI tool is unsafe or unusable in clinical practice?

Answer: Signals include zero edits from a role, unexpected workarounds, and mismatches between tool assumptions and clinician mental models.

According to HealthsystemCIO.com on August 19, 2026, Penn found a role that edited nothing in a patient-message tool, prompting targeted changes to permissions and interface design.

Pair these qualitative signals with telemetry and follow-up interviews to determine root cause.

How often should governance revisit approved AI tools?

Answer: Governance should set a named return date at approval and specify which data will be reviewed at that time.

According to HealthsystemCIO.com on August 19, 2026, Penn’s approach includes setting a return review and asking in advance 'what data are we going to look at in six months, ' a question Susan Harkness Regli described on the HealthsystemCIO.com podcast.

Regular, data-driven reviews detect drift in usage, safety signals, and conceptual integration problems.

Conclusion & Next Steps

Human factors work and staged exposure are practical ways to reduce AI usability risk in hospitals, as described by Penn Medicine on HealthsystemCIO.com on August 19, 2026.

Clinical teams should require vendor test environments, monitor small beta cohorts of 25 to 50 users per role, and set explicit post-approval review dates.

Operational researchers can accelerate synthesis by using AI-enabled qualitative tools that merge telemetry and transcripts into prioritized recommendations.

To try these workflows yourself, Try Evidano for free and see how staged-test feedback becomes actionable evidence.

Topics

  • AI usability testing in healthcare
  • human factors AI
  • clinical AI usability
  • AI tool evaluation healthcare

Keep reading

Browse all articles