Evidano is an AI-powered qualitative data analysis platform that streamlines corpus ingestion, multilingual tokenization, AI-assisted coding, and reproducible exports. Fast, repeatable qualitative analysis is essential when researchers need to compare how institutions frame science across countries or platforms. A new PLoS ONE study (published July 9, 2026) builds two corpora (Weibo (China) and X (U.S.)) and uses KH Coder to show systematic differences in topics, addressing terms, and co-occurrence patterns. This post explains what the paper measured (Apr 1–Oct 1, 2023), why those metrics matter for researchers and communicators, and a practical AI-enabled workflow you can run in Evidano (Evidano) to reproduce and extend the study on your own datasets.
Key Takeaways
Chen et al. (PLoS ONE, Jul 9, 2026) compared state science agency posts on Weibo and X and found systematic differences in topics, addressing-term mixes, and co-occurrence networks.
This post summarizes the study window (Apr 1–Oct 1, 2023), the methods used (KH Coder, Jaccard co-occurrence, addressing-term frequencies), and a practical workflow to reproduce the analyses using Evidano.
- Study corpora: Chinese = 895 posts (24, 769 tokens); English = 1, 204 posts (25, 063 tokens).
- Chinese accounts mixed national pride, personable address, and operational updates; U.S. accounts emphasized professional expertise, mission execution, and audience engagement.
- Methods included token frequencies, co-occurrence networks (Jaccard), and addressing-term frequency analysis, with a minimum term-frequency filter of ≥30 and post-level analysis units.
- You can reproduce and extend these analyses in days using the seven-step runbook and Evidano’s ingestion, multilingual tokenization, AI-assisted coding, and export features.
Fast take: What the PLoS ONE study found (July 9, 2026)
Chen et al. (published July 9, 2026 in PLoS ONE) compared state science agency posts on Weibo and X using KH Coder and a corpus-assisted discourse framework.
Read the original paper on PLoS ONE.
- Scope: posts from Apr 1 to Oct 1, 2023; Chinese corpus = 895 posts / 24, 769 tokens; English corpus = 1, 204 posts / 25, 063 tokens.
- Key contrasts: Chinese accounts mix national pride, personable address, and operational updates; U.S. accounts emphasize professional expertise, mission execution, and audience engagement.
- Methods: co-occurrence networks, addressing-term frequencies, and word-association (Jaccard coefficient) via KH Coder.
Findings snapshot
| Metric | Chinese corpus (Weibo) | English corpus (X) | Study window | Source |
|---|---|---|---|---|
| Number of posts | 895 | 1, 204 | 1 Apr – 1 Oct 2023 | PLoS ONE |
| Token count | 24, 769 | 25, 063 | , | PLoS ONE (Jul 9, 2026) |
| Top themes | Milestones, personnel, national identity, greetings, operations | Mission narratives, audience engagement, operations, events | , | PLoS ONE (Jul 9, 2026) |
| Addressing-term mix | Self 35.8% | Audience 18.7% | Nation 45.6% | Self 73.8% | Audience 22.8% | Nation 3.4% | , | PLoS ONE (Jul 9, 2026) |
What the paper did (plain English)
The paper built two cleaned corpora and used KH Coder to run quantitative content analysis, including token frequencies, co-occurrence networks, and word-association around self-addressing terms.
The authors then read concordances and interpreted patterns through positioning theory and critical discourse analysis.
- Why it matters: combining corpus linguistics with critical discourse analysis increases reproducibility while preserving contextual interpretation.
- Key methodological choices to note: manual segmentation checks, minimum term-frequency filter (≥30), and analysis unit = post-level segments.
Implications for researchers and communicators
For academic researchers
This study shows how measurable linguistic signals, such as addressing-term distributions and co-occurrence clusters, map to institutional positioning.
Researchers should use the same lens to test cross-cultural hypotheses or to validate qualitative claims with frequency and network evidence.
Reproducibility tip: keep raw tokens, document preprocessing, and report frequency thresholds, the paper used a 30-term filter.
For communication teams and UX researchers
Communication teams should track addressing-term mixes and co-occurrence shifts when A/B testing message frames to gauge personalization or institutional labeling effects on engagement and trust.
Operational tip: map theme clusters to campaign KPIs (trust, signups, event attendance) rather than relying on impressions alone.
For policy & security analysts
Policy and security analysts can use linguistic framing signals as indicators of soft power and legitimacy strategies, measurable through frequency and co-occurrence metrics.
Analytic tip: run time-series keyword frequency and co-occurrence to spot rapid reframing after major events.
Do more, faster with Evidano (mapped to this study)
Problem: building bilingual corpora & cleaning tokens
Evidano ingests scraped posts and spreadsheets, applies multilingual tokenization and lemmatization, and supports custom dictionaries to preserve domain tokens (e.g., ‘Artemis II’, 航小科).
Problem: inconsistent coding and low reproducibility
Evidano imports and exports codebooks, runs AI-assisted coding across the corpus, and records coder agreements to produce frequency-stable code counts and versioned exports for audits.
Problem: mapping co-occurrence networks and themes quickly
Evidano generates one-click co-occurrence networks and hierarchical theme visualizations with cluster coloring and edge strength by Jaccard, and provides downloadable node and edge tables for replication.
Problem: multilingual concordance and stakeholder reports
Evidano provides translation with custom dictionaries and AI chat over your documents so analysts can query concordances, pull exemplar quotes, and create executive briefs with visuals.
Security & compliance
Evidano provides enterprise-grade encryption, PII redaction in transcripts, and a strict no-third-party-model-training policy on customer data.
Runbook: 7 steps to reproduce & extend the study in two weeks
This runbook gives seven steps to reproduce and extend the study in two weeks using public post exports, a codebook, and Evidano’s ingestion and analysis features.
- Minimal inputs: public post exports (CSV/JSON) or direct scrape from platform handles, an optional codebook, and a target date window.
- 1) Import: upload Weibo/X exports or point the Evidano scraper at verified accounts.
- 2) Preprocess: enable Chinese/English tokenization, add project-specific dictionary entries (e.g., ‘长征’, ‘Artemis II’).
- 3) Auto-code: run AI-assisted thematic coding, then review a 200-item sample and accept or adjust codes.
- 4) Network: generate co-occurrence network (post-level unit) and export top 150 edges (Jaccard) for visualization.
- 5) Addressing-term analysis: define self/audience/nation groups and compute frequencies by segment.
- 6) Cross-segment: run cross-tab analysis (platform × theme × time) and flag significant shifts.
- 7) Report: use Evidano’s dashboard to export visuals, exemplar quotes, and a stakeholder one-page brief.
FAQ: Corpus-assisted qualitative analysis
What did Chen et al. compare in the PLoS ONE study?
Chen et al. compared state science agency posts on Weibo (China) and X (U.S.) to identify differences in topics, addressing-term mixes, and co-occurrence networks.
The study built two corpora spanning Apr 1–Oct 1, 2023, and used quantitative corpus methods plus qualitative interpretation to map institutional positioning.
What data and time window did the study use?
The study used posts from Apr 1 to Oct 1, 2023, with a Chinese corpus of 895 posts (24, 769 tokens) and an English corpus of 1, 204 posts (25, 063 tokens).
Those corpora were cleaned, tokenized, and analyzed at the post level using KH Coder tools and term-frequency thresholds.
Which methods did the authors apply to the corpora?
The authors applied token-frequency analysis, co-occurrence networks using the Jaccard coefficient, addressing-term frequency counts, and concordance reading guided by positioning theory and critical discourse analysis.
The analysis used a minimum term-frequency filter of ≥30 and manual segmentation checks as part of the preprocessing.
How can I reproduce the analyses quickly?
You can reproduce the analyses using the seven-step runbook: import exports or scrape verified accounts, preprocess with bilingual tokenization, run AI-assisted coding, generate co-occurrence networks, compute addressing-term frequencies, cross-tab segments, and export reports.
Evidano supports each step from ingestion to export so teams can reproduce and audit results in days.
Conclusion: turn corpus-assisted qualitative analysis into repeatable decisions
Chen et al.’s July 9, 2026 PLoS ONE paper demonstrates how corpus methods make discourse claims testable, scalable, and comparative across languages and platforms.
If teams need to move from anecdote to evidence, for example to compare framing across countries, measure shifts after a campaign, or produce audit-ready reports, Evidano streamlines the whole pipeline.
- Ready to try this on your own corpora? Start a secure trial and reproduce these analyses in days with Try Evidano for free.
