Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents. Researchers and HR teams face a pressing question: do apparently objective selection tests reduce or reproduce hiring barriers for autistic applicants? According to Whelpley et al., 2026 in PLOS One, commonly used tests including general mental ability (GMA) measures, Five-Factor personality inventories, and situational judgment tests (SJTs) produced subgroup differences unfavorable to autistic respondents. This post refracts the PLOS One findings through the lens of AI-enabled qualitative research so assessment designers and hiring teams can synthesize evidence, surface mechanisms, and prototype fairer selection systems.
Key Takeaways
According to PLOS One, Whelpley et al., 2026, a comparison of 99 neurotypical and 107 autistic adults found that multiple common selection tools produced subgroup differences unfavorable to autistic applicants (PLOS One).
- The study collected responses from June 6 to June 9, 2022, and was published on August 18, 2026, in PLOS One (Whelpley et al., 2026).
- In the sample, general mental ability (GMA) means were 5.96 (neurotypical, n = 99) and 4.73 (autistic, n = 107) with Cohen’s d = 0.52, reported by Whelpley et al., 2026.
- Situational judgment test (SJT) means were 4.24 (neurotypical) and 0.57 (autistic) with Cohen’s d = 0.84, and the SJT difference remained significant after controlling for GMA, personality, age, and gender (Whelpley et al., 2026).
- Whelpley et al., 2026 caution that “These findings do not imply a lack of job-relevant capability among autistic applicants; rather, they highlight how commonly used assessment tools may disadvantage qualified individuals when their embedded demands are misaligned with how applicants perceive and respond to evaluative situations.”
- The PLOS One article also quotes corporate hiring initiatives noting that “innovation comes from the edges, ” a framing used by SAP in practitioner discussions cited in the paper (Whelpley et al., 2026).
What happened: study design and how differences were measured
The PLOS One study by Whelpley et al., 2026 compared three common selection tools (GMA test, Five-Factor personality inventory, and an SJT) using samples recruited via MTurk, answering whether these tools reduce interview-related adverse impact.
According to Whelpley et al., 2026, the study recruited 99 neurotypical respondents (mean age 37.8, 47.5% female) and 107 autistic respondents (mean age 35.1, 41.1% female), with data collected June 6–9, 2022 and later analyzed for mean differences and effect sizes.
According to Whelpley et al., 2026, instruments included the ICAR-16 for GMA (reliabilities 0.58 neurotypical, 0.51 autistic), the 50-item IPIP for Five-Factor personality (noting low reliabilities in the autistic sample), and a 12-stem proprietary SJT validated for customer service contexts.
The authors used standardized mean comparisons, Cohen’s d effect sizes, correlation matrices, and multivariate regression (controlling for gender, age, GMA, and personality) to test whether being autistic predicted lower SJT scores independent of cognitive ability and personality (Whelpley et al., 2026).
Findings snapshot
| Date | Metric | Value | Implication |
|---|---|---|---|
| June 6–9, 2022 | Sample sizes | Neurotypical n = 99; Autistic n = 107 | Study approximated an applicant pool by requiring employment history (Whelpley et al., 2026) |
| 2026-08-18 | GMA mean scores | Neurotypical 5.96; Autistic 4.73 (Cohen’s d = 0.52) | GMA differences in this sample were meaningful but may be sample-specific (Whelpley et al., 2026) |
| 2026-08-18 | SJT mean scores | Neurotypical 4.24; Autistic 0.57 (Cohen’s d = 0.84) | SJT differences persisted after controlling for GMA and personality (Whelpley et al., 2026) |
| 2026-08-18 | Personality reliabilities | Autistic sample reliabilities as low as 0.10 for some facets | Standard personality inventories may not function equivalently for neurodiverse respondents (Whelpley et al., 2026) |
Implications for HR researchers and hiring teams
Organizations cannot assume that replacing interviews with standard tests automatically improves equity: Whelpley et al., 2026 show that SJTs, GMA tests, and personality inventories each produced subgroup differences unfavorable to autistic applicants.
- Assessment design matters: according to Whelpley et al., 2026, SJTs that rely on social judgment can disadvantage autistic applicants even when cognitive ability is controlled for.
- Measurement validity must be re-evaluated: according to Whelpley et al., 2026, low reliabilities in personality scales for autistic respondents suggest standard personality items may be misinterpreted or non-equivalent.
- Prefer job-relevant, task-based measures: Whelpley et al., 2026 recommend work samples or low-fidelity simulations that foreground task performance rather than interpersonal inference.
- Pilot and analyze subgroup effects: HR teams should collect demographic and neurodiversity markers, then use statistical and qualitative analyses to detect adverse impact as recommended by Whelpley et al., 2026.
How Evidano helps teams translate these findings into inclusive hiring design
Problem: Standard selection tools produce subgroup differences
Solution: Use Evidano to ingest assessment transcripts, open-ended SJT responses, and work-sample feedback to run thematic and content analyses that reveal whether items depend on social inference rather than task skill.
Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents; see Evidano features for capabilities relevant to instrument redesign.
Problem: Personality items show low reliability for autistic respondents
Solution: Use Evidano’s text analytics to cluster problematic items, extract respondent interpretations, and surface item-level language that confounds autistic respondents, enabling informed revisions or replacements with behaviorally anchored items.
Problem: SJTs may embed social norms that disadvantage applicants
Solution: Use Evidano to run cross-segment analyses (e.g., autistic vs neurotypical), link SJT free-text rationales to score patterns, and visualize co-occurrence networks that pinpoint which scenario features create construct contamination.
Problem: Pilots produce mixed quantitative signals
Solution: Combine Evidano’s thematic coding with frequency and cross-segment analyses to triangulate qualitative patterns with quantitative metrics such as Cohen’s d, improving decisions about which assessments to scale.
FAQ: AI qualitative analysis for hiring autistic applicants
Can SJTs disadvantage autistic applicants even after controlling for cognitive ability?
Yes: Whelpley et al., 2026 found SJT mean differences (neurotypical 4.24 vs autistic 0.57) and reported that being autistic remained a significant negative predictor of SJT score after controlling for GMA and personality.
Use qualitative analysis of SJT response rationales to detect whether items ask for social inference rather than job-relevant judgment, then redesign scenarios accordingly.
Are standard personality inventories reliable for autistic respondents?
Not always: Whelpley et al., 2026 reported unusually low reliabilities on several Five-Factor facets in their autistic sample (for example, agreeableness reliability = 0.10 in that sample).
Researchers should test measurement equivalence and use open-ended probes analyzed with AI-enabled qualitative tools to understand how items are interpreted.
What practical alternatives reduce adverse impact for autistic applicants?
Work samples and task-based simulations are promising: Whelpley et al., 2026 recommend prioritizing job-relevant demonstrations of skill over assessments that require social inference.
AI-enabled qualitative research can rapidly synthesize pilot feedback, surface patterns in free-text responses, and support iterative redesign of assessment content.
How should teams pilot inclusive assessments?
Start with mixed-method pilots that collect scores, open-ended rationales, and participant feedback, and analyze both quantitatively and qualitatively as Whelpley et al., 2026 suggest.
Evidano’s combined thematic and cross-segment analyses can accelerate that process by revealing why items produce subgroup differences and where redesign will have the most impact.
Conclusion & Next Steps
Whelpley et al., 2026 in PLOS One provide clear evidence that commonly used selection tools can reproduce adverse outcomes for autistic applicants, and that SJTs in particular showed large subgroup differences in their sample published on August 18, 2026.
Teams designing inclusive hiring systems should pair quantitative metrics (means, Cohen’s d, reliabilities) with AI-enabled qualitative analyses to understand the mechanisms behind subgroup differences and to redesign assessments accordingly.
If you want to pilot mixed-method selection research, use AI to accelerate transcript coding, cross-segment comparisons, and item-level diagnostics, then iterate with work-sample–based alternatives.
To explore Evidano’s platform and start a pilot, Try Evidano for free.
Topics
- AI qualitative analysis for hiring autistic applicants
- hiring autistic applicants
- neurodiversity hiring assessments
- AI-enabled qualitative research
Keep reading
- Commentary on NewsAI for Qualitative Analysis of Museum VisitorsAI turns observation notes into insights: qualitative analysis of museum visitors that finds themes, age segments, and exhibit recommendations for UX teams
- Commentary on NewsBoost Exhibit Design: AI Qualitative Analysis for MuseumsUse AI qualitative analysis for museums to convert observational studies into targeted exhibit changes. Learn from PLOS ONE findings and try Evidano to accelerate insight.
- Commentary on NewsImprove Museum Visitor Research: AI Qualitative AnalysisAI qualitative analysis for museums: convert visitor observations into themes, counts, and actionable design changes. Learn from a PLOS ONE study and see how Evidano helps.
