Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents. According to the PLOS One study published on August 4, 2026, transcribing a human codebook directly into an LLM prompt reaches 82–88% agreement with human annotators across three experimental-economics classification tasks (PLOS One). According to the PLOS One study, the experiments included 703, 253 API calls conducted between March 24 and April 6, 2026, and used 1, 494 promise items, 493 level-k items, and 851 level-k items respectively, giving concrete scale to the claims. This post explains what the PLOS One results mean for AI-enabled qualitative research and shows how researchers can operationalize the study's recommendation to invest in codebook quality rather than surface-level prompt tricks.
Key Takeaways
According to the PLOS One study published August 4, 2026, faithfully transcribing a human codebook into an LLM prompt (including category definitions and worked examples) produced 82–88% agreement with human-agreed labels on three experimental classification tasks (PLOS One).
- The PLOS One study reports that 703, 253 classification requests were submitted between March 24 and April 6, 2026, with 702, 809 successful calls and a 0.06% error rate.
- The PLOS One study measured 1, 494 items for the promise task, 493 items for level-k task I, and 851 items for level-k task II, giving task-level sample sizes for replication.
- The PLOS One study finds that model choice explains most variation on recognition-heavy tasks, while the level of informational detail in the prompt drives performance on learning-heavy tasks.
- The PLOS One authors conclude: "preparing instructions as one would for human annotators, with detailed context, category definitions, and examples, " is the recommended approach (Çelebi & Penczynski, PLOS One, August 4, 2026).
What Happened and how the experiment was measured
Answer: The PLOS One team converted three published human codebooks into prompt variants and tested them across four LLMs and six experimental manipulations to measure classification accuracy.
According to the PLOS One study published August 4, 2026, the researchers translated each codebook into prompt components (experiment context, theory context, classification instructions, worked examples) and ran roughly 703, 253 API calls between March 24 and April 6, 2026 across GPT 5.4, GPT 5.4-nano, Qwen 3.5 397B, and Qwen 3.5 9B (PLOS One).
According to the PLOS One study, ground truth was defined by items where two human annotators agreed, and accuracy was the proportion of items where the model matched that agreed label; the study reports 95% Wilson confidence intervals for all accuracy estimates.
Findings snapshot
| Date | Metric | Value | Implication |
|---|---|---|---|
| August 4, 2026 | Reported overall agreement range | 82–88% agreement with human-agreed labels | A verbatim codebook prompt plus examples yields human-level agreement on the tested tasks (PLOS One). |
| March 24–April 6, 2026 | API requests | 703, 253 calls (702, 809 successful, 0.06% error) | Large factorial design supports robust paired comparisons in the PLOS One experiments. |
| Dataset details (reported in PLOS One) | Item counts | Promise P = 1, 494; L I = 493; L II = 851 | Per-task sample sizes are adequate for significance testing and McNemar paired comparisons. |
| August 4, 2026 | Model sensitivity | Larger models tolerate surface variation; smaller models are more brittle | Choose model scale according to task complexity and desired robustness (PLOS One). |
Implications for qualitative researchers and coding teams
Answer: Invest time improving the codebook (definitions, boundary cases, worked examples) and pair it with a capable model rather than iterating on formatting and persona prompts.
According to the PLOS One study published August 4, 2026, detailed classification instructions and worked examples drive accuracy on learning-heavy tasks, while model choice dominates recognition-heavy tasks such as promise detection.
According to the PLOS One study, surface choices like markdown formatting, role persona framing, or minor phrasing changes had negligible effects when the codebook content was complete and verbatim.
According to the PLOS One study, "The practical recommendation is to pair a frontier model with a well-defined codebook, " which means teams should prioritize clear category definitions and examples over superficial prompt engineering (Çelebi & Penczynski, PLOS One, August 4, 2026).
How Evidano helps you apply codebook prompting
Problem: Manual transcription and noisy codebook text
Answer: Evidano automates ingestion and cleaning of codebooks and source documents so you can produce verbatim prompts reliably.
Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents, and Evidano can ingest codebooks, filter procedural notes, and output a cleaned verbatim prompt suitable for LLM classification.
Evidano feature pages explain the import and cleaning steps in detail: Evidano features.
Problem: Need to test prompt variants and measure robustness
Answer: Evidano runs systematic variant tests and returns paired comparisons and frequency tables so you can detect underspecification.
Evidano’s analysis workflows produce thematic and frequency analyses and let you run paired comparisons across prompt versions; these outputs let you replicate the McNemar-style paired tests used in the PLOS One study.
Problem: Data privacy when using proprietary LLMs
Answer: Evidano encrypts data and does not share customer data to third-party model training, supporting safe codebook-to-prompt workflows.
Evidano’s data handling and encryption practices are described on the security page: Evidano data security.
FAQ: codebook prompting for LLM classification
Can I just paraphrase my codebook and expect the same accuracy?
Answer: No, paraphrasing can change outcomes if the prompt becomes underspecified; fidelity matters.
According to the PLOS One study published August 4, 2026, paraphrasing produced effects below 1 percentage point for frontier models but caused larger and inconsistent changes for smaller models when examples were absent.
Do formatting and role personas matter once I include a full codebook?
Answer: Not reliably; surface choices are negligible on well-specified prompts for capable models.
According to the PLOS One study, formatting and framing had almost no significant effect on verbatim (full-content) prompts for GPT 5.4 and only marginal effects for other large models.
Can model reasoning (chain-of-thought) replace missing codebook detail?
Answer: No, model reasoning cannot reliably substitute for missing domain detail.
According to the PLOS One study, enabling reasoning did not compensate for missing classification instructions on learning-heavy tasks and in some cases worsened performance for GPT-family models.
What sample sizes and tests did the PLOS One study use?
Answer: The PLOS One study used 1, 494 promise items, 493 L I items, 851 L II items and paired McNemar tests to control for item difficulty.
According to the PLOS One study, the experiments covered 703, 253 classification requests across models and conditions, and all paired comparisons were evaluated at p < 0.001 with 95% Wilson confidence intervals.
How can I tell if my prompt is well-specified or underspecified?
Answer: Test surface robustness: if paraphrases, formatting, or small edits change outputs, your prompt is likely underspecified.
According to the PLOS One study, robustness to surface variation is the diagnostic signal: well-specified prompts produce consistent classifications regardless of formatting or phrasing.
Conclusion & Next Steps
Answer: Follow the PLOS One recommendation: transcribe your human codebook verbatim, include worked examples, and pair it with a capable model for robust LLM classification.
According to the PLOS One study published August 4, 2026, this approach produced 82–88% agreement with human-agreed labels across three tasks and 703, 253 total classification requests during their experimental window.
If your team needs a platform to ingest codebooks, run prompt variants, and produce paired analyses and visualizations, Evidano automates those steps and preserves data security.
Get started and Try Evidano for free.
