Site Logo
All articles
Commentary on News

Inclusive Hiring: AI Qualitative Analysis of Selection Tests

Evidano7 min read

Primary keyword: ai qualitative analysis of hiring tests. HR researchers and qualitative teams must understand whether commonly used selection tools reproduce hiring disparities for autistic applicants. According to the PLOS One study by Whelpley et al., 2026, autistic respondents in a sample scored lower than neurotypical respondents on personality inventories, general mental ability (GMA) tests, and situational judgment tests (SJTs). According to PLOS One (Whelpley et al., 2026), these differences persisted for SJTs even after controlling for GMA, personality, gender, and age, which suggests selection-tool demand mismatch rather than lack of job-relevant capability.

Key Takeaways

According to PLOS One (Whelpley et al., 2026), commonly used selection tests produced subgroup differences unfavorable to autistic applicants; the original article is available at PLOS One.

  • Sample sizes: the study used 99 neurotypical respondents and 107 autistic respondents, with data collected June 6–9, 2022, and published August 18, 2026 in PLOS One.
  • GMA result: according to PLOS One (Whelpley et al., 2026), mean ICAR-16 scores were 5.96 for neurotypical respondents and 4.73 for autistic respondents (Cohen’s d = 0.52).
  • SJT result: according to PLOS One (Whelpley et al., 2026), mean SJT scores were 4.24 for neurotypical respondents and 0.57 for autistic respondents (Cohen’s d = 0.84), and being autistic remained a significant negative predictor of SJT score after multivariate controls.
  • Practical reading: according to PLOS One (Whelpley et al., 2026), shifting from interviews to standard tests does not automatically eliminate adverse impact; organizations should validate fit between assessment demands and job requirements.

What happened: study design and core findings

This section answers what the PLOS One study measured and what it found.

According to PLOS One (Whelpley et al., 2026), the authors recruited two samples on Amazon Mechanical Turk: 99 neurotypical respondents (average age 37.8) and 107 self-identified autistic respondents (average age 35.1), with attention checks removing 14 participants in total.

According to PLOS One (Whelpley et al., 2026), the study administered a 50-item Five-Factor personality inventory (IPIP-50), the 16-item ICAR GMA measure, and a 12-stem, job-validated customer service SJT scored across 24 responses.

According to PLOS One (Whelpley et al., 2026), the personality reliabilities were lower in the autistic sample (for example, agreeableness alpha = 0.10 in the autistic sample versus 0.82 in the neurotypical sample), which the authors flagged as a limitation for interpreting personality differences.

According to PLOS One (Whelpley et al., 2026), autistic respondents scored lower on average across personality facets (notably agreeableness, conscientiousness, and emotional stability), GMA, and SJT, with SJT differences persisting after controlling for GMA and personality in regression models.

Findings snapshot

Date / SourceMetricValue (neurotypical vs autistic)Implication
June 6–9, 2022 / PLOS OneSample size99 neurotypical; 107 autisticApproximate applicant-like samples for comparison
August 18, 2026 / PLOS OneGMA mean (ICAR-16)5.96 vs 4.73 (Cohen’s d = 0.52)Moderate-large group difference in this sample
August 18, 2026 / PLOS OneSJT mean (customer service SJT)4.24 vs 0.57 (Cohen’s d = 0.84)Large group difference; SJT may draw on social inference
August 18, 2026 / PLOS OnePersonality reliabilities (alpha)Neurotypical e.g., agreeableness 0.82; Autistic agreeableness 0.10Standard personality measures may not function equivalently for autistic respondents

Implications for HR researchers and qualitative teams

What should HR researchers change in validation studies?

Answer: validate selection tools for neurodiverse subgroups explicitly and report subgroup reliabilities.

According to PLOS One (Whelpley et al., 2026), the study found extremely low alpha values for some personality scales in the autistic sample, and the authors recommend examining measurement equivalence before assuming a tool is fair.

How should qualitative teams study applicants’ experiences?

Answer: use mixed-methods inquiry to surface how test items are interpreted by autistic applicants.

According to PLOS One (Whelpley et al., 2026), the likely mechanism for SJT differences is a mismatch between social-interpretive demands in the test and autistic respondents’ information processing, which qualitative probing can reveal.

What hiring changes reduce adverse impact for autistic applicants?

Answer: prioritize job-relevant work samples and task-based assessments over measures that rely on social inference or normative self-comparison.

According to PLOS One (Whelpley et al., 2026), practitioner programs like those at Microsoft and SAP favor project-based demonstrations and assessment centers, and the authors suggest these may better align with job requirements while reducing social-impression contamination.

How Evidano helps researchers translate PLOS One findings into inclusive hiring practice

Problem: ambiguous item interpretation across groups → Solution: targeted qualitative coding

Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents.

Evidano can ingest interview transcripts, open-ended feedback, and survey comments from applicants and then generate thematic and content analyses that highlight where autistic respondents interpret SJT or personality items differently; see Evidano features for relevant capabilities.

Problem: low reliability in subgroup measures → Solution: iterative item testing with AI-assisted tagging

Evidano supports transcription with PII redaction and custom dictionaries, enabling consistent capture of respondent phrasing for coding and measurement equivalence checks; see Evidano speech-to-text.

Evidano’s co-occurrence and subcode visualizations let teams see which item wordings correlate with misunderstanding among autistic respondents, accelerating item revision.

Problem: designing inclusive SJTs → Solution: rapid prototyping using mixed-methods bundles

Evidano can combine open responses to SJT stems, demographic segments, and performance scores to produce cross-segment analyses that show whether SJT content privileges normative social inference over procedural judgment.

Evidano’s AI chat-over-documents lets teams ask natural-language questions (for example, “Which SJT items triggered social-norm language for autistic respondents? ”) and receive extractable evidence for validation reports.

FAQ: ai qualitative analysis of hiring tests

Do SJTs disadvantage autistic applicants compared with neurotypical applicants?

Answer: In the PLOS One sample, SJTs were associated with large group differences unfavorable to autistic applicants.

According to PLOS One (Whelpley et al., 2026), mean SJT scores were 4.24 for neurotypical respondents versus 0.57 for autistic respondents (Cohen’s d = 0.84), and autism remained a significant negative predictor after controlling for GMA and personality.

Are personality inventories unreliable for autistic respondents?

Answer: Standard personality inventories may show low internal reliability in autistic samples and should be tested before use.

According to PLOS One (Whelpley et al., 2026), alpha coefficients in the autistic sample were markedly lower for several Big Five facets (for example, agreeableness alpha = 0.10), prompting caution about score interpretation.

Can AI qualitative analysis help redesign hiring assessments?

Answer: Yes, AI-enabled qualitative analysis can systematically surface interpretation patterns and recommend targeted item or process changes.

According to PLOS One (Whelpley et al., 2026), the likely mechanism for adverse outcomes is construct contamination by social inference, which qualitative methods can identify and help redesign into task-relevant formats.

How should research teams combine quantitative effect sizes with qualitative evidence?

Answer: Start with subgroup effect-size reporting and follow with thematic interviews that explain why items produce those effects.

According to PLOS One (Whelpley et al., 2026), reporting Cohen’s d for GMA (0.52) and SJT (0.84) alongside qualitative probes into item interpretation gives a more actionable validation strategy.

Conclusion & Next Steps

PLOS One (Whelpley et al., 2026) demonstrates that commonly used selection tests can produce subgroup differences unfavorable to autistic applicants, so HR teams should not assume objectivity without validation.

PLOS One (Whelpley et al., 2026) authors warn that "these findings do not imply a lack of job-relevant capability among autistic applicants; rather, they highlight how commonly used assessment tools may disadvantage qualified individuals" (Whelpley et al., 2026).

For next steps, combine the quantitative benchmarks reported in PLOS One with targeted qualitative probes to redesign items or substitute work samples where appropriate.

Start by extracting applicant open-text responses and interview transcripts into an AI-enabled qualitative workflow to generate thematic, frequency, and cross-segment analyses; learn more at Evidano features and Evidano speech-to-text.

Try Evidano for free: Try Evidano for free.

Topics

  • ai qualitative analysis of hiring tests
  • qualitative analysis of selection tests
  • AI-enabled qualitative research
  • neurodiversity hiring analysis

Keep reading

Browse all articles
Company
About
Newsletter

Product updates, research, and tips — straight to your inbox.

© Evidano, All Rights Reserved.