The PLOS study (published July 6, 2026) pinpoints a pragmatic mismatch in Chinese scholars’ spoken English: heavy use of We-based hedges (We see) and underuse of I-based hedges (I think). This post shows how to run a practical qualitative analysis of lexical hedges on your transcripts, reproduce the study’s diagnostics, and convert results into classroom-ready Data-Driven Learning (DDL) materials using Evidano. If you work in EAP, research training, or spoken discourse analysis, learn a reproducible pipeline you can run on your own corpus with www.evidano.com.
Key Takeaways
Evidano is an AI-powered qualitative data analysis platform that lets you reproduce the July 6, 2026 PLOS One finding that CASEC favors We V (that) while MICASE favors I V (that), and convert corpus diagnostics into classroom DDL materials.
The PLOS One diagnosis (July 6, 2026) shows a diagnostic I vs We pragmatic split across CASEC and MICASE, and the same pipeline (automatic transcription, concordance extraction, phrase-level coding) reproduces those diagnostics at scale.
- The PLOS study compared CASEC (696, 009 words; 279 events) with MICASE (1, 848, 364 words; 152 events) and found opposing phrasal preferences.
- The phrasal split: MICASE I V (that) = 2, 126 pmw vs CASEC I V (that) = 76 pmw; CASEC We V (that) = 889 pmw vs MICASE We V (that) = 87 pmw.
- Overall lexical-hedge frequency was CASEC 3, 592 pmw vs MICASE 3, 823 pmw (LL = -7.27, p < 0.01), a macro-level difference the study documents.
Fast take: what the PLOS study found
The PLOS One study (published July 6, 2026) found a striking I vs We split: MICASE heavily favors I V (that) while CASEC overuses We V (that).
Wang et al. compared a 696, 009-word Chinese Scholars’ Academic Spoken English Corpus (CASEC; 279 events) with MICASE (1, 848, 364 words; 152 events) and documented the phrasal mismatch that can affect reception in Anglo-American gatekeeping venues.
- Full study: PLOS One article (July 6, 2026).
- Why it matters: the phrasal mismatch can affect reception in Anglo-American gatekeeping venues; below we show how to reproduce the qualitative analysis and operationalize DDL with AI-enabled tools.
Findings snapshot
| Metric | Value | Source | Note |
|---|---|---|---|
| CASEC size | 696, 009 words; 279 events | PLOS One article (July 6, 2026) | Spoken presentations (2018–2025) |
| MICASE size | 1, 848, 364 words; 152 events | PLOS One article (July 6, 2026) | Anglo‑American reference |
| Overall lexical-hedge frequency | CASEC: 3, 592 pmw; MICASE: 3, 823 pmw (LL = -7.27, p < 0.01) | PLOS study (July 6, 2026) | Macro underuse in CASEC |
| think (pmw) | MICASE: 2, 803 pmw; CASEC: 1, 247 pmw | PLOS study (Table 8) | MICASE preference for mentalistic hedges |
| see (pmw) | CASEC: 1, 027 pmw; MICASE: 310 pmw | PLOS study (Table 8) | CASEC preference suggests visual/objective framing |
| Phrasal patterns (I V vs We V) | I V (that): MICASE 2, 126 pmw vs CASEC 76 pmw; We V (that): CASEC 889 pmw vs MICASE 87 pmw | PLOS study (Table 9) | Core diagnostic: I vs We split |
What the authors did (methods, short)
The authors compiled CASEC (696k words, 279 events; 2018–2025) and compared it to MICASE (1.85M words) and targeted 10 high-frequency lexical verbs (see, think, assume, consider, etc.).
- Automatic ASR then manual correction: the study used OpenAI Whisper large-v3 for initial transcripts, manual correction in ELAN, and final verification, yielding word-level agreement 98.4%.
- Concordance retrieval (AntConc) produced 9, 567 lines; the authors manually coded hedge function with Cohen’s κ = 0.85.
- Normalization to pmw and log-likelihood testing determined significance; phrasal pattern extraction focused on subject type × complement (I V that, We V that, It V that, N V that).
Implications for EAP instructors, presenters, and corpus researchers
EAP instructors
EAP instructors should target phrasal patterns, not just single words, because I think and We see function differently in pragmatic terms.
Use concordance-led noticing tasks (DDL) to shift pragmatic awareness before drilling production.
Presenters & research teams
Presenters aiming for Anglo‑American gatekeeping venues should explicitly signal authorial stance with I when asserting ownership of claims.
Balance We-based consensus language for data presentation with I-based statements when asserting novel claims.
Corpus & qualitative analysts
Corpus and qualitative analysts should prioritize phrase-level analysis because subject + verb + complement reveals pragmatic patterns that token counts miss.
Triangulate frequency with concordance-driven qualitative coding and dispersion checks to avoid misleading claims.
Do more, faster with Evidano (map to this use case)
Ingest, transcribe, translate
Evidano ingests recordings, papers, or mixed corpora and auto-transcribes with custom dictionary support and PII redaction, reproducing the study’s Whisper→manual pipeline but at scale.
If you need multilingual materials, use Evidano to run translation with a custom dictionary to keep technical terms consistent.
Phraseology, concordances & coded themes
Evidano extracts concordances and identifies phrasal patterns (subject type × complement) automatically, producing the I V vs We V tables researchers need.
Evidano generates hierarchical codes and subcodes for hedges, stance markers, and pragmatic functions and visualizes co-occurrence networks.
Frequency & cross-segment analysis
Evidano runs normalized frequency comparisons (pmw), log-likelihood tests, and cross-segment contrasts (discipline, genre, year) in one workspace.
Evidano exports tables and charts that match the structure used in the PLOS paper.
Turn diagnosis into DDL materials
Evidano automatically generates worksheets (concordance selections, guided prompts) and role-play scripts and uses AI chat over your uploaded corpus to draft instructor guides.
Evidano can optionally run AI avatar interviewers to simulate Q&A practice sessions for presenters.
Security & compliance
Evidano encrypts data and does not use your data to train third-party models, important when working with prepublication talks or sensitive defense recordings.
You control exports and redaction settings in Evidano.
Two-week pilot workflow (checklist)
This two-week checklist reproduces the PLOS-style qualitative analysis on your corpus using the same diagnostic steps.
- 1) Collect 20–50 recorded talks or 50–200 transcripts and upload to Evidano.
- 2) Auto-transcribe with a custom dictionary, then review edits in the editor; enable PII redaction if needed.
- 3) Extract target lexical verbs and generate concordances; export 9–10k lines like the study’s sample.
- 4) Auto-tag subject types (I, we, it, inanimate) and complement frames; produce pmw tables and run LL tests.
- 5) Curate 20 contrastive concordance lines for I vs We worksheets and let Evidano generate guided questions.
- 6) Run AI avatar Q&A role-plays for presenters to practice switching I/We strategies and capture recordings for feedback.
- 7) Produce a one-page pedagogical brief and export slides/visuals for stakeholders.
FAQ: common questions from researchers
How do I reliably label hedge vs non-hedge instances?
Use concordance context windows and a small manually coded seed to label hedge vs non-hedge instances, the most reliable approach. Evidence-based practice is to seed with about 500 lines and then validate with inter-coder checks.
Evidano’s AI-assisted coding can suggest labels, and you should validate suggestions with inter-coder checks targeting Cohen’s κ ≥ 0.8.
Can I compare cohorts (disciplines or years)?
Yes, you can compare cohorts by running normalized pmw metrics and cross-segment frequency tests. Evidano supports segmentation by metadata such as discipline, year, and genre.
Is the ASR reliable for accented academic speech?
ASR accuracy improves when you add custom lexicons and perform manual correction, that hybrid workflow yields high agreement. The PLOS study combined Whisper large-v3 with manual verification and reported 98.4% word-level agreement; Evidano provides the same hybrid workflow.
Wrapping up & next steps
The PLOS One diagnosis (July 6, 2026) of a clear I vs We pragmatic split is a concrete example of how corpus-driven qualitative analysis uncovers high-leverage pedagogical targets.
Replicate the study’s diagnostic steps in days, not months, by combining automated transcription, concordance extraction, phrase-level coding, and AI-assisted worksheet generation. Start a pilot at Try Evidano for free to ingest your corpus, run the qualitative analysis, and produce classroom-ready DDL materials.
