Site Logo
Research MethodsContent and Framework Approaches

Summative Content Analysis: from word counts to meaning

Evidano6 min read

Summative content analysis begins by counting and refuses to end there. The first stage quantifies manifest content — occurrences of chosen words, phrases, or content markers across a corpus, often broken down by source, speaker, or period. The second stage interprets the counted: reading each occurrence in context to understand how the term is used, what it euphemises, when it appears and when it is conspicuously replaced. The design suits questions where usage itself is the phenomenon — how a diagnosis label spreads, how “safety” talk differs between documents and meetings — and its integrity lives entirely in stage two. Counting alone is a concordance, not an analysis.

The two stages and what each claims

Stage one — manifest quantification. Identify the target terms (including synonyms, inflections, and euphemisms decided in advance or discovered iteratively), count occurrences, and map their distribution: by document type, speaker role, time period, section. The claims here are descriptive and checkable: who says this word, where, how often.

Stage two — latent interpretation. Examine occurrences in their contexts — the keyword-in-context discipline — to characterise usage: senses, functions (assertion, hedge, irony, quotation), collocations, and the situations that trigger substitution or silence. The claims here are interpretive: what the pattern of use means.

The link between stages is the method’s spine. Every interpretive claim should be traceable to counted, locatable instances; every count should eventually be read. Studies that decouple them fail in opposite directions — numerology or impressionism.

Source and positioning

The approach is the third of Hsieh and Shannon’s Three approaches to qualitative content analysis: where conventional analysis derives categories from data and directed analysis from theory, summative analysis starts from manifest terms and proceeds to their latent use.

Its intellectual kin are corpus linguistics (concordances, collocation) and the manifest/latent vocabulary of the content-analysis tradition; what the summative approach adds is a qualitative research frame — purposive corpora, contextual interpretation, and trustworthiness standards rather than purely statistical ones.

It is also the approach most often performed unknowingly: any study that searches transcripts for mentions of X and then discusses “how X came up” is doing informal summative analysis, usually without the discipline either stage deserves.

Questions the design fits

  • Terminology as tracer: how “burnout” migrated from private talk to board papers; whether “consent” appears in training materials and how it is glossed.
  • Comparative usage: the same term across outlets, professions, or periods — where distribution differences are themselves findings.
  • Euphemism and avoidance: what replaces the direct term, and in which contexts the direct term becomes sayable.
  • Policy–practice language gaps: terms mandated in documents versus their presence and sense in recorded practice.
  • Poor fits: phenomena not tied to identifiable surface forms (experiences, reasoning), and any question where paraphrase variety would make term lists arbitrary — those need conventional or thematic analysis.

Doing both stages properly

Build the term set transparently

Start from the research question’s target terms; expand with synonyms, inflections, and discovered variants (stage-two reading feeds back new candidates). Log every addition with its rationale — the term set is the instrument, and reviewers should see its construction.

Count with context captured

Extract every occurrence with enough surrounding text to interpret later, tagged by source metadata (speaker, document, date, section). Report zero-counts too: the units where the term never appears are part of the distribution.

Read every occurrence — or a defensible sample

Small corpora: read all instances. Large ones: read all instances of rare terms and a stratified sample of frequent ones, stated as such. Code each occurrence for sense and function; the coding scheme here is usually small and emerges quickly.

Analyse the joint pattern

Put distribution and usage together: not “the word appeared 340 times” but “in documents it labels a procedure; in meetings it appears almost solely inside distancing quotation marks, and vanishes when seniors are present”. Discrepant instances get read closely, not discarded.

Report so both stages are auditable

Publish the term set, the counting rules, the distribution tables, and exemplar-in-context quotations for each claimed sense. The reader should be able to re-run stage one and challenge stage two.

Worked example: “restraint” in care-home documentation

A safeguarding study examined how physical restraint was documented across 12 care homes: 3,400 incident reports, care plans, and policy documents. The term set began with “restraint/restrain” and grew, via stage-two reading, to include “guided”, “redirected”, “supported to the floor”, and “comfort hold” — each addition logged.

Stage one showed the headline distribution: “restraint” was frequent in policies, rare in incident reports — while the euphemism set ran the other way, and one home accounted for a third of all “comfort hold” usage. Stage two read every incident-report occurrence: “restraint” appeared almost exclusively in negations (“no restraint was used”), while the euphemisms carried the descriptive load for physically identical acts, as revealed by the surrounding narrative detail.

The joint finding — the mandated term functioned as a category to be avoided rather than applied, with home-level euphemism dialects — was checkable at every step: term set published, counts by home tabled, and each claimed sense evidenced with quoted contexts. The regulator’s response targeted documentation training, and the follow-up audit reused the same instrument, which is the reproducibility this design buys.

Common mistakes

  • Stopping at stage one. Frequency tables presented as findings — the concordance fallacy.
  • Unstated term sets. Counts of an instrument nobody can inspect; synonym choices silently drive results.
  • Ignoring negation and quotation. “No restraint was used” counts as an occurrence only if the analysis knows what it is doing with polarity and voice.
  • Interpreting unread counts. Assigning meaning to distributions whose instances were never examined in context.
  • Corpus convenience. Counting whatever text was easy to get, then generalising to the practice the texts unevenly record.
  • Treating usage change as attitude change. Language shifts under mandates and fashions; the inference from words to beliefs needs argument, not assertion.

Limitations

The method is chained to surface forms: phenomena expressed without stable vocabulary escape it, and speakers’ paraphrase creativity is a permanent leak in any term set. Iterative expansion narrows the leak; nothing seals it.

Its corpora are records, with all the selection and production biases records carry — what gets written, by whom, under what incentives, is upstream of every count.

And the design’s comparative power depends on comparable units: counting across documents of wildly different lengths, genres, and authorship conventions requires normalisation choices that are themselves contestable and must be reported.

Where software helps

Both stages are tool-shaped. Stage one — term-set extraction with context, across thousands of documents, tabulated by source and period — is mechanical and should be automated; Evidano’s frequency and content analysis does the counting with every instance linked to its passage. Stage two’s occurrence coding (sense, function, polarity) is a small codebook applied at scale, the same auditable-coding workflow, leaving the analyst a complete keyword-in-context evidence base instead of a sampling hope.

The interpretation of the joint pattern — what the distribution of use means — is stage two’s human remainder, and the write-up’s worth stands or falls on it.

Topics

  • summative content analysis
  • word frequency analysis
  • manifest content
  • latent content interpretation
  • keyword in context
  • qualitative content analysis approaches

Other methods in content and framework approaches

Written guides are linked directly; the rest have a reference entry in the methodology directory.

Published research using these methods

Studies and evaluations where this family of method was applied with Evidano — the work, not the claim.

Keep reading

Browse all articles