AI Thematic Analysis: Accuracy, Methods, and Limits
What published evidence actually shows about AI coding qualitative data — where it matches expert human analysis, where it fails, and how to use it in a way you can defend to a reviewer.
The question is no longer “can AI code data?” but “how well, and checked how?”
Thematic analysis — identifying patterns of meaning across interviews, focus groups, and open-text responses — has always had a bottleneck: a rigorous first-pass coding of a few hundred transcripts costs weeks of expert time. Large language models can now read qualitative data with enough contextual understanding to code it, which moves the real question from capability to reliability: how closely does AI coding agree with expert human coding, on what kinds of data, and how would you know?
That question has answers now, because researchers have started publishing head-to-head comparisons instead of impressions. This page summarizes that evidence and lays out a workflow that keeps the researcher — in Braun and Clarke’s phrase, the person actively constructing the themes — in charge of the loop.
What the published comparisons show
Three independent comparisons involving Evidano put numbers on the question, each benchmarking AI-assisted coding against expert human analysis of the same data:
- 371 interview transcripts, 96% agreement. Arizona State and Penn State researchers compared AI-powered coding of small-group discussion transcripts against human coding, reporting high consistency between the two in a peer-reviewed paper. Manual coding had taken 13 weeks; processing with Evidano took about 12 hours.
- 298 evaluation reports, 92% agreement. A UN evaluation synthesis applied a realist framework — 132 context-mechanism-outcome configurations — and found 92% agreement between AI-assisted and human coding, with weeks of analysis compressed to about 3 hours.
- 160 management responses, 54 hours to 1. UNICEF’s review of humanitarian-evaluation management responses cross-checked generative-AI analysis against manual coding, reducing roughly 54 hours of manual work to about one.
Agreement in these studies is measured the way inter-rater reliability between human coders is measured — percentage agreement, Cohen’s kappa, F1 against a human ground truth. By the conventional Landis-and-Koch reading of kappa, values above 0.61 count as substantial agreement and above 0.81 as almost perfect; many human coding pairs do not clear those bars without calibration rounds. The full study list, including the approximate 94% unweighted average across 11 independent studies (a directional figure — metrics differ across studies), is maintained on Human vs AI.
Two caveats belong next to any accuracy number. First, results depend on the setup: clear code definitions and well-structured data produce far better agreement than vague prompts over messy text. Second, the studies above tested a purpose-built research tool; general-purpose chatbots evaluated in the same literature score materially lower, and several published attempts report themes that sound plausible but cannot be traced to the data.
The hallucinated-quote problem — and the traceability answer
The most common objection researchers raise about AI qualitative analysis is not speed or nuance; it is fabricated evidence. A fluent model asked to “find themes with supporting quotes” can invent a quote, stitch two speakers together, or paraphrase so loosely the meaning shifts — and the output looks identical to honest analysis. If the AI’s claims cannot be checked, the analysis is not auditable, and unauditable analysis has no place in research.
The structural fix is to make every claim carry its evidence: each code application and each theme links to the exact passage in the source file it rests on, so verification is a click rather than a search. That is how Evidano is built, and it is why the published comparisons above could measure agreement at all — auditors could see what the AI coded where. How this works in practice is covered on source-traceable quotes.
Where AI fits in Braun and Clarke’s six phases
Braun and Clarke’s six-phase model remains the most widely used account of thematic analysis. Mapping AI onto it makes the division of labor concrete — AI accelerates the mechanical phases; the interpretive spine stays human:
1. Familiarization
Where AI helps: AI transcribes audio and video, translates multilingual data, and produces first summaries so you meet the whole corpus quickly.
Where you stay in charge: You still read. Familiarization is where analytic instincts form; summaries orient, they do not substitute.
2. Generating initial codes
Where AI helps: The strongest use of AI: a complete first-pass coding of every transcript — inductive or against your codebook — with supporting quotes attached to each code.
Where you stay in charge: You audit the codes against their quotes, correct misreadings, and add codes for meanings the AI missed.
3. Searching for themes
Where AI helps: AI clusters codes into candidate themes and surfaces co-occurrence and frequency patterns across the dataset.
Where you stay in charge: You decide which candidate patterns are analytically meaningful rather than merely frequent.
4. Reviewing themes
Where AI helps: AI retrieves every coded extract for a theme instantly, making the check against the full dataset fast instead of prohibitive.
Where you stay in charge: The judgment calls — collapsing, splitting, discarding themes — are yours.
5. Defining and naming themes
Where AI helps: AI drafts definitions you can react to, and tests candidate boundaries by retrieving edge-case extracts.
Where you stay in charge: Naming what a theme is about is interpretive work; AI drafts are raw material, not conclusions.
6. Producing the report
Where AI helps: AI assembles quote evidence per theme, builds visualizations, and drafts descriptive passages with citations back to sources.
Where you stay in charge: The argument of the paper — what the analysis means and why it matters — cannot be delegated.
For a fully reflexive analysis in Braun and Clarke’s sense — where themes are actively constructed through the researcher’s subjectivity and a fixed codebook is itself rejected — automation of coding is a poor fit by design. Studies taking that stance have still used AI legitimately as a second coder or devil’s advocate against their own reading. Our guide to reflexive thematic analysis treats this properly.
A defensible AI thematic analysis workflow
- Fix your analytic frame first. Decide inductive vs deductive, and if deductive, finalize codes and definitions before the AI sees data — the same discipline you would apply with a second human coder.
- Run the AI pass with evidence links on. Use a tool that attaches source-linked quotes to every code application; refuse any output you cannot trace.
- Audit a sample like an inter-rater check. Independently code a sample, compare against the AI, and examine every disagreement — treat the AI as a coder whose reliability you are establishing, not assuming.
- Do the interpretive phases yourself. Theme review, naming, and the analytic narrative are your contribution; use AI retrieval to test your claims against the full dataset.
- Disclose and cite. Name the tool, describe its role, and follow your journal’s AI policy — cite-us has wording used in published methods sections and a review of 41 journals’ policies.
Frequently asked questions
- Can AI do thematic analysis reliably?
- Under the right conditions, published evidence says yes — with a human in charge. In a peer-reviewed comparison, AI-assisted coding with Evidano agreed with expert human coders on 96% of codes across 371 interview transcripts; a UN evaluation synthesis reported 92% agreement across 298 reports. Reliability depends on clear code definitions, source-linked evidence you can audit, and researcher review of the output — not on the AI alone.
- How is agreement between AI and human coders measured?
- The same way agreement between two human coders is measured: percentage agreement, Cohen’s kappa, or F1 scores against a human “ground truth” coding. By the conventional Landis and Koch reading, kappa above 0.61 is substantial agreement and above 0.81 almost perfect — thresholds many human coding pairs do not reach on the first pass.
- Does AI thematic analysis work with an existing codebook?
- Yes, and this is where AI coding is most defensible: with a codebook’s codes and definitions fixed in advance, applying them consistently across hundreds of transcripts is exactly the mechanical task AI does well, and agreement with human coders can be audited code by code against the linked quotes.
- Will reviewers and journals accept AI-assisted thematic analysis?
- A growing number of journals allow it with disclosure, and Evidano has been named in the methods sections of peer-reviewed studies. Disclose the tool and workflow, keep the researcher’s role explicit, and verify quotes against sources. Our cite-us page reviews the AI policies of 41 journals and provides methods-section wording.
- What is the biggest risk in AI thematic analysis?
- Unverifiable output. A generic chatbot can produce fluent themes with fabricated or paraphrased “quotes” that do not exist in the data. The guard is traceability: every code application and theme should link to the exact passage supporting it, so the researcher can check the evidence rather than trust the summary.
- Does AI replace the researcher in thematic analysis?
- No. AI compresses the mechanical first pass — segmenting, initial coding, retrieving evidence — from weeks to hours. Constructing meaning, judging what matters, naming themes, and writing the account of the data remain the researcher’s work, and published studies that used Evidano describe it as a second coder or thought partner, not a replacement.
Keep exploring
Test AI thematic analysis on your own data
The only convincing benchmark is your data, your codebook, and your judgment. The free plan is free forever — run a transcript through and audit the output.
