Evidano is an AI-powered qualitative data analysis platform that helps teams reproduce and scale corpus-assisted discourse studies. Researchers Chen et al. (Published 9 July 2026) compared how Chinese and U.S. government science accounts positioned themselves on Weibo and X using a corpus-assisted discourse approach. This post explains the study and how research and communications teams can reproduce, extend, and scale that kind of qualitative analysis with Evidano (www.evidano.com). You’ll get a one-page snapshot of the datasets, the key differences the paper found, and a practical two-week workflow to run comparable thematic and co-occurrence analyses on your own social media corpus.
Key Takeaways
Chen et al. found Chinese government space accounts frame space work with national pride and approachable self-addressing, while U.S. agency accounts foreground professional mission updates and standardized address, using posts sampled 1 April–1 October 2023.
Evidano supports reproducing that pipeline with multilingual preprocessing, co-occurrence networks, AI-assisted coding, and traceability back to source quotes.
- Corpus snapshot: Chinese corpus (Weibo) = 895 posts, U.S. corpus (X) = 1, 204 posts, collected 1 Apr–1 Oct 2023 (Chen et al., published 9 Jul 2026).
- Addressing distribution: Chinese corpus nation-addressing 45.58% vs U.S. corpus nation-addressing 3.42% (annotated addressing-term groups, Chen et al., 2026).
- Reproducibility choices reported include selected accounts, date range, token thresholds, and top-edge filters, enabling replication or extension.
Fast take + source
Fast take: a corpus-assisted discourse study found reproducible cross-cultural framing differences between Chinese and U.S. official space accounts.
TL; DR: A corpus-assisted discourse study (Chen et al., published 9 July 2026) built two corpora from Weibo and X (1 Apr–1 Oct 2023) and used KH Coder to show that Chinese agencies frame space work with national pride and approachable self-addressing, while U.S. agencies foreground professional mission updates and standardized address. Read the paper: PLOS ONE.
- Why it matters: method = CADS (critical discourse analysis + corpus linguistics); payoff = reproducible signals (themes, addressing patterns, co-occurrence networks).
- How AI helps: automated tokenization, cross-language alignment, rapid co-occurrence network generation and coded segment comparisons.
Qualitative analysis of space science communication: findings snapshot
| Metric | Chinese corpus (Weibo) | U.S. corpus (X) | Notes / Dates | Source |
|---|---|---|---|---|
| Posts (n) | 895 | 1, 204 | Collected 1 Apr–1 Oct 2023 | Chen et al., published 9 Jul 2026 |
| Tokens | 24, 769 (≈67, 926 Chinese characters) | 25, 063 (≈37, 785 characters) | Preprocessed with KH Coder (lemmatized / segmented) | Chen et al., 2026 |
| Addressing term distribution | Self 35.75% | Audience 18.67% | Nation 45.58% | Self 73.78% | Audience 22.80% | Nation 3.42% | Quantified from annotated addressing-term groups | Chen et al., 2026 |
| Primary themes | Milestones, human/infrastructure, national pride, personalization, operations | Mission execution, audience engagement, astronauts, celestial events, partnerships | Derived from co-occurrence networks (Jaccard index) | Chen et al., 2026 |
What the study did (plain English)
The study built two language-specific corpora from three verified official accounts per country, sampled posts from 1 April to 1 October 2023, preprocessed text, and ran co-occurrence network and word-association analyses in KH Coder.
- Unit of analysis: post-level segments; token frequency threshold = 30; top 150 co-occurrence edges by Jaccard coefficient.
- Analytic focus: (1) topics/themes via co-occurrence clusters; (2) addressing terms (self, public, nation); (3) words co-occurring with self-addressing terms to infer positioning strategies.
- Key reproducibility choices are reported (accounts selected, dates, thresholds), so the pipeline can be replicated or extended.
Implications for researchers & comms teams
At a glance
This section summarizes implications for UX researchers, policy analysts, and communications teams based on Chen et al.'s findings.
Use theme extraction plus addressing-term annotation to map institutional relation to audiences, and use co-occurrence networks to surface clusters that manual coding might miss.
UX / qualitative researchers
UX and qualitative researchers should combine theme extraction with addressing-term annotation to map how institutions relate to audiences.
If your goal is to map how institutions relate to audiences, combine theme extraction with addressing-term annotation (as Chen et al. did). Use co-occurrence networks to surface clusters you might not see in manual coding.
Compare segments (platform, language, campaign) quantitatively: report frequencies plus representative concordance lines for each theme to avoid cherry-picking.
Policy & security analysts
Policy and security analysts can use quantified differences to diagnose semiotic alignment with national narratives and to track shifts over time.
Quantified differences (e.g., nation-addressing 45.6% vs 3.4%) are diagnostic: they reveal semiotic alignment with national narratives. Track shifts over time to detect strategic repositioning after policy events.
Cross-check with platform moderation or partner mentions to interpret operational transparency versus legitimating discourse.
Communications teams
Communications teams should test addressing frames because language choices materially change perceived proximity and credibility.
Language choices (colloquial self-addressing vs standardized agency labels) materially change perceived proximity and credibility. Test alternative address frames in A/B content experiments and measure engagement by segment.
Use co-occurrence clusters to craft messages that align subject matter (mission details) with the tone your audience expects (patriotic vs professional).
Do more, faster with Evidano (mapped to this use case)
Overview
Evidano supports ingestion, multilingual preprocessing, thematic and network analysis, AI-assisted coding, and traceability for this use case.
Import scraped social posts or upload CSVs, run multilingual segmentation, generate co-occurrence networks, and trace results back to source quotes within the platform.
Ingest & multilingual prep
Evidano ingests scraped social posts or CSVs and automatically segments Chinese and English text while supporting custom dictionaries.
Import scraped social posts or upload CSVs of posts and metadata; Evidano automatically segments Chinese and English text and supports custom dictionaries for domain terms (e.g., “长征”, “Artemis II”).
Optional: apply translation with a custom dictionary to align tokens across languages before cross-corpus comparison.
Automated thematic and co-occurrence analysis
Evidano generates thematic clusters and co-occurrence networks quickly and exports top edges with representative concordance lines.
Generate thematic clusters and co-occurrence networks (Jaccard or PMI) in minutes; export top edges and representative concordance lines to support CDA-style interpretation.
Visualize hierarchical codes, subcodes, and network clusters for stakeholder-ready figures that replicate KH Coder outputs with interactive filters.
Reliable coding & cross-segment stats
Evidano runs AI-assisted coding with traceability to reduce inter-coder drift and supports cross-segment frequency and significance tests.
Import or build codebooks and run AI-assisted coding to apply labels consistently; review and accept suggestions to reduce inter-coder drift.
Run cross-segment frequency and significance tests (e.g., address-term distributions by country, platform, or time window) and export tables for methods appendices.
Traceability, chat & security
Evidano links every result back to source quotes and post metadata and ensures data encryption without training third-party models.
Every result links back to source quotes and post metadata for transparent CDA work. Use AI chat over your corpus to iterate hypotheses (for example, “show me posts where we and China co-occur with ‘成功’”).
Data is encrypted and never used to train third-party models, important when handling state-related content or sensitive research datasets.
Two‑week workflow: reproduce & extend Chen et al.’s CADS
This two-week workflow reproduces and extends Chen et al.'s corpus-assisted discourse study (CADS).
- Day 1–2: Define accounts and scrape posts (1 Apr–1 Oct 2023 in the paper). Export raw JSON/CSV.
- Day 3–4: Clean, add custom dictionary entries (project names, mascots), and align Chinese↔English tokens via translation.
- Day 5–6: Run automated tokenization and frequency lists; correct segmentation errors interactively.
- Day 7–9: Generate co-occurrence networks (filter: top 150 edges; threshold=30) and extract cluster keywords.
- Day 10: Annotate addressing-term categories (self/audience/nation) and run frequency comparisons.
- Day 11–12: Inspect representative concordance lines; write analytic memos tying clusters to socio-political context.
- Day 13–14: Produce shareable visuals, export tables for appendix, and prepare a stakeholder brief.
Wrapping up & next steps
Wrapping up: Chen et al. (published 9 Jul 2026) demonstrate that corpus-assisted discourse methods yield reproducible, interpretable signals about institutional positioning across languages and platforms.
If you need to scale that approach (multilingual corpora, consistent coding, co-occurrence networks, and secure research workflows) try a hands-on pilot in Evidano: Try Evidano for free. Start with the two-week workflow above and iterate using the platform’s chat-over-corpus and visualization exports.
FAQ: space science communication
What did Chen et al. study and when was it published?
Chen et al. studied how Chinese and U.S. government science accounts position themselves on social media, and the paper was published 9 July 2026.
The authors built two corpora from Weibo and X covering posts from 1 April to 1 October 2023 and used corpus-assisted discourse methods with KH Coder to analyze themes and addressing-term distributions.
How were the corpora built and what date range did the study use?
The study built language-specific corpora from three verified official accounts per country and sampled posts from 1 April to 1 October 2023.
The authors preprocessed text with segmentation and lemmatization, applied a token frequency threshold of 30, and focused on the top 150 co-occurrence edges by Jaccard coefficient.
What were the main communication differences between Chinese and U.S. space accounts?
Chinese accounts emphasized national pride and more nation-addressing, while U.S. accounts emphasized professional mission updates and more self-addressing as agency labels.
Quantified examples include nation-addressing at 45.58% in the Chinese corpus versus 3.42% in the U.S. corpus, and different theme clusters such as milestones and personalization versus mission execution and partnerships.
How can a research team reproduce this study quickly?
A compact two-week plan can reproduce the study by scraping the same accounts and date range, preprocessing text, generating co-occurrence networks, annotating addressing terms, and exporting concordance lines and tables.
Follow the provided day-by-day workflow to define accounts, scrape posts, clean and align tokens, run tokenization and network generation, annotate addressing categories, and prepare visuals and appendices.
