This post shows how a qualitative analysis of lexical hedges (driven by a new 696, 009-word CASEC corpus) identifies a clear pragmatic mismatch and yields classroom-ready interventions. Published July 6, 2026, Wang et al. compare Chinese scholars’ spoken academic English (CASEC) with the Michigan Corpus of Academic Spoken English (MICASE) and find Chinese speakers underuse I-based hedges (I think) and overuse We-based patterns (We see). Read the original study at PLOS ONE. If you’re an EAP instructor, researcher, or academic developer, this post explains how to reproduce their qualitative workflow at scale and how Evidano automates concordancing, phrase-pattern extraction, secure transcription, and DDL module creation so you can move from diagnosis to treatment in days not months. Quick CTA: scroll to “Do More, Faster with Evidano” to see a 7-step reproducible workflow mapped to platform features.
Key Takeaways
Evidano is an AI-powered qualitative data analysis platform that automates transcription, concordancing, phrasal-pattern extraction, visual reporting, and worksheet generation while keeping your data private.
The Wang et al. study (July 6, 2026) identifies a clear I versus We pragmatic mismatch: CASEC underuses I-based hedges (I think) and overuses We-based patterns (We see) compared with MICASE.
The diagnostic pipeline (pmw normalization, log-likelihood testing, and manual concordance coding with Cohen’s Kappa=0.85) is reproducible and can be scaled using the 7-step runbook below and platform automation.
- CASEC (696, 009 words; 279 events) shows macro underuse of lexical-verb hedges: CASEC 3, 592 pmw vs MICASE 3, 823 pmw (LL = -7.27, p < 0.01).
- The contrastive pattern includes extreme subject-pattern differences: I V (that) CASEC 76 vs MICASE 2, 126, and We V (that) CASEC 889 vs MICASE 87.
- Manual coding supported the quantitative findings: 9, 567 concordance lines were coded with inter-coder reliability Cohen’s Kappa = 0.85.
- Pedagogy: a 90-minute DDL workshop, built from concordance slices, can raise pragmatic awareness and train context-sensitive hedge use.
Findings Snapshot
| Date / Metric | Value | Source | Implication |
|---|---|---|---|
| Published | July 6, 2026 | PLOS ONE | Corpus-driven study; DDL module included |
| CASEC size | 696, 009 words; 279 events | CASEC (2018–2025) | Specialized corpus of Chinese faculty speech |
| MICASE size | 1, 848, 364 words; 152 events | MICASE | Anglo-American reference variety |
| Overall lexical-verb hedges (pmw) | CASEC: 3, 592 pmw; MICASE: 3, 823 pmw (LL = -7.27, p < 0.01) | Table 7 | Macro underuse in CASEC |
| think (pmw) | CASEC: 1, 247; MICASE: 2, 803 | Table 8 | CASEC underuses I-based hedge think |
| see (pmw) | CASEC: 1, 027; MICASE: 310 | Table 8 | CASEC overuses visualized/collective marker see |
| I V (that) pattern (pmw) | CASEC: 76; MICASE: 2, 126 | Table 9 | Marked underuse of individualized stance |
| We V (that) pattern (pmw) | CASEC: 889; MICASE: 87 | Table 9 | Heavy preference for collective stance in CASEC |
What Happened: Key qualitative diagnostics
The study executed a contrastive, corpus-driven qualitative analysis focusing on 10 high-leverage lexical verbs (for example: think, see, assume, consider).
Wang et al. compiled CASEC (2018–2025) and compared normalized frequencies and phraseological patterns with MICASE, combining quantitative normalization (pmw) and log-likelihood testing with manual concordance coding (9, 567 concordance lines; Cohen’s Kappa=0.85 on coding reliability).
- Core qualitative finding: a strong I vs We distinction; MICASE favors I V (that) as an authorial stance, CASEC favors We V (that), especially We see.
- Interpretation: CASEC’s pattern aligns with collective, positive-politeness strategies, diffusing authorial responsibility and rooted in cross-cultural rhetorical norms.
- Pedagogical output: a 90-minute DDL workshop (with worksheets, roleplays) to raise pragmatic awareness and train context-sensitive hedge use.
So What for EAP instructors & researchers: Implications
For EAP instructors
The diagnosis shows a high-leverage, teachable target: swapping or balancing We-based patterns with I-based hedges in contexts where individual authorial stance is expected.
Use concordance-driven noticing tasks (DDL) to surface pragmatic functions rather than prescriptive lists.
For corpus & pragmatics researchers
The study models a reproducible pipeline: compile targeted spoken events, transcribe/verify, extract concordances, code hedging function, and triangulate frequency with phrasal patterns.
Extend the approach by adding ELF comparisons or modality/adverbial stance markers for fuller pragmatic mapping.
For academic developers / program leads
A focused workshop (90 minutes) built from concordance slices can be deployed at scale for doctoral schools and faculty development units.
Measure change by pre/post roleplay rubrics and by automated frequency checks across recorded talks.
Do More, Faster with Evidano
Problem: laborious transcription → Solution: secure, accurate transcription
Evidano ingests recorded lectures and auto-transcribes with a custom dictionary for domain terms, provides PII redaction, and allows collaborative manual correction, matching the multi-stage pipeline used in the study.
Problem: building concordances & phrase patterns → Solution: automated concordance + phrasal extraction
Evidano extracts concordances for target verbs, computes normalized frequencies (pmw), and surfaces phrasal patterns (subject + verb + complement) so you can reproduce the I V vs We V analysis in minutes.
Problem: turning diagnosis into DDL materials → Solution: worksheet & activity generator
Evidano auto-generates classroom-ready worksheets (contrastive concordance sets), role-play prompts, and reflection logs from your corpus slices, exportable for in-person or online workshops.
Problem: iterative evaluation & stakeholder buy-in → Solution: visual reports & cross-segment analysis
Evidano produces visualizations (word clouds, co-occurrence networks, hierarchical code trees) and cross-segment analyses (discipline, event type) to show measurable changes to colleagues and funders.
Evidano encrypts your data and does not use it to train third-party models.
Reproducible 7‑Step Workflow (Runbook)
Follow these seven steps to reproduce the study-style qualitative analysis with Evidano as the engine.
- 1) Collect audio/video of target spoken events and metadata (discipline, event type, speaker role).
- 2) Upload to Evidano; run auto-transcription with a custom dictionary for names, technical terms; apply PII redaction if needed.
- 3) Review transcripts in-platform (time-aligned), invite a second verifier, export final transcripts.
- 4) Use Evidano concordance extractor to pull all inflected forms of target verbs; filter by context window and export concordance lines.
- 5) Code hedging function (in-platform labeling); Evidano supports collaborative tagging and inter-coder comparison.
- 6) Run frequency normalization and phrasal-pattern extraction (subject + verb + complement); visualize I vs We distributions by segment.
- 7) Auto-generate DDL worksheets and role-play materials; run workshop and re-measure hedge frequencies in subsequent talks.
FAQ: qualitative analysis of lexical hedges
What counts as a lexical-verb hedge?
Lexical-verb hedges are verbs that modulate epistemic commitment, for example think, seem, assume, and see.
Function matters: coding distinguishes cognitive uses from hedging uses.
How do I compare corpora of different sizes?
Normalize counts per million words (pmw) and use log-likelihood for significance, the method used in the study to control for size differences.
The Wang et al. study used pmw normalization and log-likelihood testing to compare CASEC and MICASE.
Is this approach ethical for public recordings?
The PLOS study used publicly available recordings and received an ethics waiver for secondary analysis.
For non-public data, obtain consent and apply PII redaction; Evidano supports secure storage and redaction workflows.
Conclusion: From Diagnosis to Action
Wang et al. (July 6, 2026) draw a direct line from corpus diagnostics to pedagogy: a measurable I vs We pragmatic mismatch in Chinese scholars’ spoken English can be remediated with targeted DDL.
If you want to replicate this qualitative analysis and spin up evidence-based workshops rapidly, Evidano automates the heavy lifting (transcription, concordancing, phrase-pattern extraction, visual reporting, and worksheet generation) while keeping your data private.
Ready to run your own CASEC-style study or scale DDL across a faculty program? Try Evidano for free
