Site Logo
All articles
Commentary on News

Teach Authorial Voice: Qualitative Analysis of Lexical Hedges

Evidano7 min read

Fast payoff: a July 6, 2026 PLOS ONE study comparing a new Chinese Scholars’ Academic Spoken English Corpus (CASEC; 696, 009 words, 279 events) with MICASE shows a striking pragmatic mismatch: Chinese faculty underuse I-based hedges (I think) and overuse We-based and visual verbs (We see, see). Read this if you design EAP curricula, run UX/qual studies with multilingual presenters, or need a reproducible Data-Driven Learning (DDL) workflow. We summarize the findings, give a quick numbers snapshot, and show how to reproduce the paper’s diagnosis-to-treatment pipeline using AI-enabled qualitative research tools (including automated transcription, concordance extraction, phrase-pattern counts, and worksheet export). Try it on your own corpus with Evidano.

Key Takeaways

Evidano is an AI-powered qualitative data analysis platform that auto-transcribes academic speech, extracts concordances, and exports DDL worksheets.

Wang et al. (Published July 6, 2026) compared CASEC (compiled 2018–2025) with MICASE and found a pronounced I versus We authorial split: MICASE heavily uses I-based hedges while CASEC favors We and visual/collective verbs.

The study provides a reproducible pipeline and classroom-ready DDL supplements that you can apply to multilingual presentation corpora.

  • Wang et al. built CASEC from 279 open-access events (696, 009 words) and compared lexical verb hedges against MICASE using normalized frequencies (pmw) and log-likelihood tests.
  • MICASE shows much higher I-based hedging, for example the I V(that) pattern at 2, 126 pmw versus 76 pmw in CASEC; CASEC favors We patterns (We V(that) 889 pmw vs. 87 pmw in MICASE).
  • The transcription pipeline used OpenAI Whisper large-v3 auto drafts with manual correction and independent verification (mean word agreement 98.4%), and manual coding achieved Cohen’s κ = 0.85 on a 20% sample.
  • The paper includes reproducible materials: concordances and DDL classroom worksheets that fast-track replication and intervention design.

Fast Take + Source

The core finding: Wang et al. (Published July 6, 2026) diagnosed a massive I vs. We split in lexical verb hedging between CASEC and MICASE, with MICASE relying on individual-stancer tokens like I think and CASEC favoring We see/We assume.

Why it matters: pragmatic markers like I think function as authorial stance in Anglo-American gatekeeping contexts, and their underuse can affect perceived authorial ownership.

Payoff for practitioners: the study supplies an actionable DDL module and concordance worksheets as supplements, and the pipeline is reproducible with AI tools for rapid corpus diagnostics.

Full article: PLOS ONE.

Findings snapshot

MetricCASEC (Chinese scholars)MICASE (Reference)Note / Implication
Corpus size (words)696, 009 (279 events; compiled 2018–2025)1, 848, 364 (152 events)CASEC designed to represent Chinese faculty-level spoken English
Overall lexical verb hedges (pmw)3, 5923, 823CASEC slightly lower overall; distribution differs by verb
think (pmw)1, 2472, 803MICASE reliance on 'think' (individual stance) much higher
see (pmw)1, 027310CASEC overuses 'see' (visual/collective framing)
I V (that) pattern (pmw)762, 126MICASE uses I-based hedges as default authorial stance
We V (that) pattern (pmw)88987CASEC strongly favors collective stance markers (We)

What happened, methods in plain English

The authors compiled CASEC from open-access lectures, conference talks, and PhD defenses (2018–2025), transcribed 279 events, and compared lexical verb hedges against MICASE using normalized frequencies (pmw) and log-likelihood tests.

Target verbs (see, think, assume, consider, etc.) produced 9, 567 concordance lines, and manual coding (Cohen’s κ = 0.85 on a 20% sample) separated hedging uses from other senses.

The transcription pipeline used OpenAI Whisper large-v3 auto drafts, followed by manual correction and independent verification yielding mean word agreement 98.4%, and key dates are: Received March 17, 2026; Accepted June 14, 2026; Published July 6, 2026.

  • Units of analysis were lexical verb plus subject plus complement (I V(that), We V(that), It V(that), inanimate N V(that)).
  • Analytical steps were automatic retrieval (AntConc) then manual disambiguation, phrasal-pattern counts, qualitative interpretation, and DDL module design.

Implications for EAP instructors, researchers, and teams

For EAP instructors

EAP instructors should target the I versus We contrast explicitly in workshops to raise awareness and practice context-sensitive stance choices.

The paper supplies a DDL module and concordance worksheets so instructors can guide noticing, hypothesis formation from concordance lines, and role-play practice.

Measure learning outcomes by tracking increases in I V(that) use in simulated Q&A tasks and reductions in ambiguous We usage where individual stance is required.

For UX / qualitative research teams

UX and qualitative research teams should add phrasal-pattern counts (subject type plus complement) to thematic coding to detect pragmatic stance transfer across segments.

Teams can use cross-segment comparisons (discipline, native language, presentation genre) to identify where pragmatic mismatches are likely to surface in stakeholder-facing communications.

For corpus linguists & program leads

Corpus linguists and program leads can reproduce the pipeline: compile recordings, auto-transcribe, manually verify, run concordance, code hedging tokens, and export worksheets for DDL.

The study’s supplements include concordances and classroom materials to fast-track replication, and next steps include triangulating CASEC–ELFA–MICASE comparisons and testing the DDL module empirically.

Do more, faster with Evidano

From raw recording to concordance (minutes → hours)

Evidano speeds the pipeline from raw recording to concordance by combining auto-transcription tuned for academic speech with in-app correction and time-aligned text export.

Auto-transcription can use a custom dictionary for author names and technical terms and supports PII redaction to preserve research ethics, and Evidano lets you correct transcripts in-app and export annotated concordance lines for target verbs (e.g., see, think, assume).

Automate phrasal-pattern counts and cross-segment tests

Evidano can compute normalized frequency comparisons (pmw) across segments such as discipline, genre, and speaker seniority, and it produces tables similar to the paper’s tables.

Evidano generates phraseology outputs (I V(that), We V(that)) and exports classroom-ready concordance worksheets to reproduce the paper’s DDL materials.

Move from diagnosis to pedagogy

Evidano helps teams move from diagnosis to pedagogy by assembling example concordances, creating worksheet PDFs, and attaching role-play prompts in one workflow.

You can deploy AI avatar interviewers for practice and measure hedge usage over time, and security and compliance features include data encryption and assurances that data is not used to train third-party models; learn more or start a trial at Evidano.

Checklist: Reproduce this study’s diagnosis-to-treatment in 2 weeks

This checklist provides a step-by-step runbook to reproduce the study’s pipeline using AI-enabled qualitative research tools.

  • 1) Gather recordings (public talks, defenses). Record metadata: speaker role, discipline, event date.
  • 2) Auto-transcribe with a custom dictionary for names/terms; then run one-pass manual correction and a 10% verification sample.
  • 3) Extract target verbs via concordance and disambiguate hedging sense with a small human-coded sample to establish κ.
  • 4) Compute normalized frequencies (pmw) and log-likelihood tests across segments, then export tables.
  • 5) Extract phrasal patterns (subject type + complement) and assemble 20–40 contrastive concordance lines for DDL worksheets.
  • 6) Run a 90-minute DDL workshop, use role-play scenarios, collect pre/post samples, and compare hedge-pattern frequencies to measure transfer.

FAQ: Qualitative analysis of lexical hedges

Can automated tools reliably detect hedging uses?

No, automated tools cannot fully detect hedging uses without human oversight because the hedging function is context-dependent.

Automatic retrieval is an excellent first pass for finding candidate tokens, but manual disambiguation or human-in-the-loop models are required for high precision, as the paper demonstrates with manual coding and κ = 0.85.

Is the pedagogical goal to replace We with I?

No, the pedagogical goal is repertoire expansion, not replacement: teach scholars when I-based hedges signal individual authorial stance and when We-based phrasing suits consensus-building.

The paper’s DDL module trains context-sensitive choice rather than enforcing a wholesale substitution of We by I.

Are the study's materials and DDL worksheets available?

Yes, the study’s supplements include concordances and classroom materials to fast-track replication and DDL implementation.

Practitioners can use those supplements to assemble worksheets and to reproduce the paper’s diagnosis-to-treatment pathway in their own contexts.

What were the key dates and reliability measures reported?

The study’s manuscript was received March 17, 2026, accepted June 14, 2026, and published July 6, 2026, and manual coding yielded Cohen’s κ = 0.85 on a 20% sample.

The transcription pipeline reported mean word agreement of 98.4% after auto-draft plus manual correction and independent verification.

Wrapping up: next moves (with a CTA)

Bottom line: Wang et al. (Published July 6, 2026) provide a compact, reproducible diagnosis (the I vs. We split) and a ready-to-run DDL module you can adapt for EAP teaching or multilingual presentation analysis.

If you analyze multilingual presentation corpora or teach EAP for faculty, add phrasal-pattern counts to your analytic toolkit and operationalize the DDL worksheets to measure transfer.

  • Ready to try this on your own recordings? Upload audio, auto-transcribe, run concordances, and export DDL worksheets in one workflow with Try Evidano for free.
  • Reproduce the original paper and access the supplements here: PLOS ONE.
Company
About
Newsletter

Product updates, research, and tips — straight to your inbox.

© Evidano, All Rights Reserved.

Teach Authorial Voice: Qualitative Analysis of Lexical Hedges | Evidano