Site Logo
All articles
Commentary on News

AI-driven Qualitative Analysis: Hiring Tests for Autism

Evidano6 min read

This post explains how AI-enabled qualitative analysis can surface why widely used hiring tests may disadvantage autistic applicants, using the PLOS One study by Whelpley et al. (published August 18, 2026) as the empirical anchor. The primary keyword for this article is "qualitative analysis of hiring tests" and the audience is hiring teams, I-O researchers, and qualitative analysts who must translate test scores, interviews, and open responses into fair hiring designs. The PLOS One study compared personality inventories, general mental ability tests, and situational judgment tests across autistic and neurotypical samples and reported concrete effect sizes and sample statistics that a qualitative research workflow can help explain and mitigate.

Key Takeaways

According to PLOS One, a study published August 18, 2026, autistic respondents scored lower on personality inventories, GMA tests, and situational judgment tests than neurotypical respondents, and the authors warn that "commonly used selection tools... may nonetheless result in adverse impact for autistic applicants if used in hiring decisions, " (Whelpley et al., 2026).

  • The PLOS One study sampled 107 autistic and 99 neurotypical adults, with data collected June 6–9, 2022, and published August 18, 2026.
  • The neurotypical mean GMA score was 5.96 versus 4.73 for autistic respondents, and the study reported Cohen’s d = 0.52 for GMA differences, according to PLOS One.
  • The situational judgment test (SJT) means were 4.24 for neurotypical versus 0.57 for autistic respondents, with Cohen’s d = 0.84, and the difference on the SJT remained significant after controlling for GMA and personality, as reported in PLOS One.
  • The PLOS One authors conclude that "these preliminary findings suggest that commonly used selection tools... may nonetheless result in adverse impact for autistic applicants if used in hiring decisions, " (Whelpley et al., 2026).

What Happened and how the study measured subgroup differences

The PLOS One study compared autistic and neurotypical adults on three common selection tools and measured mean differences and effect sizes to assess adverse impact.

According to PLOS One, the authors recruited participants on Amazon Mechanical Turk, collected responses June 6–9, 2022, and after quality checks retained 107 autistic and 99 neurotypical respondents.

The PLOS One authors used the 50-item IPIP for personality, the ICAR-16 for general mental ability (GMA), and a 12-stem situational judgment test (SJT) developed for customer service, then reported reliabilities, mean scores, and Cohen’s d values for group comparisons.

The PLOS One study also ran multivariate regressions and reported that, even after controlling for age, gender, GMA, and personality, being autistic still predicted lower SJT scores (see Table 5 in PLOS One).

Findings snapshot

Date / SourceMetricValueImplication
June 6–9, 2022 (data collection), PLOS OneSample sizesAutistic = 107, Neurotypical = 99Study approximates applicant pools by using employed respondents
August 18, 2026, PLOS OneGMA meanNeurotypical = 5.96; Autistic = 4.73; Cohen’s d = 0.52Moderate to large group difference on cognitive test
August 18, 2026, PLOS OneSJT meanNeurotypical = 4.24; Autistic = 0.57; Cohen’s d = 0.84Large group difference that persisted after controls

Implications for hiring teams and researchers

Hiring teams should not assume objective tests automatically remove bias, because the PLOS One study shows tests can reproduce adverse impact.

According to PLOS One, all three selection methods (personality, GMA, SJT) produced subgroup differences unfavorable to autistic respondents, which implies that test content and format deserve scrutiny before being used as gatekeepers.

Researchers should treat validity as a design problem: the PLOS One authors recommend testing whether assessments predict job-relevant outcomes for autistic employees, not just whether they separate groups on average.

Practitioners should prioritize task-based work samples and assessment-center simulations for roles where interpersonal inference is not central, since the PLOS One discussion cites Microsoft and SAP project-based hiring as practitioner models worth studying.

How Evidano helps translate these findings into inclusive hiring practice

Problem: Tests show group differences but provide little qualitative context

Answer: Evidano extracts thematic explanations from interview transcripts, open responses, and test feedback to explain why groups differ.

Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents.

Evidano can ingest SJT response rationales, interview transcripts, and debrief notes to generate thematic codes and cross-segment frequency counts that reveal whether item wording, social framing, or response format drives subgroup gaps.

Problem: Low reliability and measurement non-equivalence in neurodiverse groups

Answer: Evidano supports mixed-methods workflows to diagnose measurement problems with qualitative evidence.

Evidano can pair statistical summaries (e.g., low Cronbach alpha reported by the PLOS One authors for several personality scales) with qualitative excerpts showing how autistic respondents interpret items differently, helping teams decide whether to adapt items or choose alternative measures.

Evidano integrations with speech-to-text enable accurate transcripts of live assessments which can be fed back into thematic analyses.

Problem: Designing valid, scalable alternatives to interviews

Answer: Evidano helps teams prototype and evaluate work samples and low-fidelity simulations with qualitative tagging and AI-assisted synthesis.

Evidano can compare performance narratives across demographic segments to surface whether a work sample measures job-relevant skills rather than social impression management, and the platform’s thematic reports can document design decisions for legal defensibility.

Contextual links

Learn more about capabilities on our features page.

FAQ: qualitative analysis of hiring tests

How did the PLOS One study measure subgroup differences between autistic and neurotypical respondents?

Answer: The PLOS One authors measured subgroup differences using mean comparisons and Cohen’s d for personality, GMA, and SJT scores and followed with multivariate regression to control for covariates.

According to PLOS One, the study reported means, effect sizes (Cohen’s d = 0.52 for GMA; Cohen’s d = 0.84 for SJT), and regression models showing autism remained a negative predictor of SJT after controls.

Are the PLOS One findings definitive for all hiring contexts?

Answer: No, the PLOS One authors describe the results as preliminary and sample-specific.

According to PLOS One, limitations include MTurk sampling, SJT content validated on neurotypical workers, and low personality reliabilities for the autistic sample, so replication with field applicants is needed.

How can qualitative analysis help reduce adverse impact identified by the PLOS One study?

Answer: Qualitative analysis identifies the content features and respondent interpretations that create mismatch between assessments and applicant experience.

The PLOS One paper recommends design-focused solutions such as using work samples and task-based assessments; qualitative coding of open responses can show whether item wording, social inference, or normative language is the root cause.

What direct quotations from the study should hiring teams cite?

Answer: Use the authors’ cautionary language to justify testing and redesign rather than outright rejection of standard tools.

For example, Whelpley et al. (2026) state, "these preliminary findings suggest that commonly used selection tools... may nonetheless result in adverse impact for autistic applicants if used in hiring decisions, " as reported in PLOS One.

Conclusion & Next Steps

The PLOS One study (Whelpley et al., 2026) shows that personality inventories, GMA tests, and SJTs can produce subgroup differences unfavorable to autistic applicants and that SJT differences persisted after controlling for several covariates.

Organizations and researchers should pair quantitative gap analysis with AI-enabled qualitative workflows to discover whether test content, item phrasing, or format drives observed disparities.

Evidano helps teams run those mixed-methods analyses, synthesize themes that explain subgroup gaps, and document redesign decisions so hiring tools measure job-relevant skills equitably.

If you want to pilot an AI-assisted qualitative workflow that maps test items to respondent interpretations and produces evidence-backed design changes, Try Evidano for free.

Topics

  • qualitative analysis of hiring tests
  • hiring tests for autistic applicants
  • AI-enabled qualitative research
  • situational judgment tests autism

Keep reading

Browse all articles
Company
About
Newsletter

Product updates, research, and tips — straight to your inbox.

© Evidano, All Rights Reserved.