Site Logo
All articles
Research MethodsContent and Framework Approaches

Conventional Qualitative Content Analysis: categories from the data

Evidano7 min read

Conventional qualitative content analysis is the workhorse of health and nursing research, and it is chosen for a good reason: when a field has little existing theory about a phenomenon, a method that derives its categories from the data without importing a framework is exactly right. Its weakness is the same as its strength. It is deliberately descriptive, it does not claim interpretive depth, and studies that use it while wanting the depth of thematic analysis end up producing neither well.

One of three approaches, and the distinction matters

Hsieh and Shannon's Three Approaches to Qualitative Content Analysis is one of the most cited methods papers in health research, and its central contribution is a distinction people routinely collapse.

Conventional content analysis derives coding categories directly from the data. It is used when existing theory or literature on the phenomenon is limited, and it avoids preconceived categories on purpose.

Directed content analysis starts from existing theory or prior research, which supplies initial codes that the analysis then extends or challenges. It is used to validate or extend a theoretical framework.

Summative content analysis counts occurrences of words or content, then interprets the underlying context of their use. It is the closest of the three to quantitative content analysis and the furthest from thematic work.

Saying which one you are doing is the first requirement of reporting. A paper that cites Hsieh and Shannon without specifying the approach has not told the reader what was done.

The procedural literature

Alongside Hsieh and Shannon, the other standard reference is Elo and Kyngäs's The qualitative content analysis process, which sets out the inductive and deductive routes in more procedural detail and is the paper most nursing studies follow.

Elo and colleagues later addressed the field's trustworthiness problem in Qualitative Content Analysis: A Focus on Trustworthiness, which is the more useful of the two for anyone preparing a study for review — it identifies the specific reporting failures that recur.

The method's roots are in mid-twentieth-century communication research, which is why its vocabulary (units of analysis, coding units, categories) reads more systematically than the interpretive traditions.

The three phases

Preparation

Select the unit of analysis — a whole interview, a passage, a sentence — and be explicit about it. Different units produce different analyses and the choice is rarely justified in published work.

Decide whether latent content (what is implied) will be analysed alongside manifest content (what is stated). Conventional content analysis can do both, but the study must say which, since a claim about implied meaning needs different evidence.

Then immerse: read the material repeatedly to obtain a sense of the whole before coding anything.

Organisation

Code openly, deriving labels from the words of the data itself rather than from an existing framework. Where the data supplies a good label, use it.

Group codes into subcategories, and subcategories into categories, on the basis of similarity and difference. The hierarchy should be built upward from the data, not imposed downward.

Where the analysis goes further, group categories into main categories or a theme that abstracts across them. This is the point at which content analysis is doing something close to thematic analysis, and where the two are most often confused.

Construct a category scheme and check it against the data: does every unit have a place, are the categories mutually distinguishable, and does the scheme cover the material?

Reporting

Report the analysis process and the resulting categories, with enough detail that a reader can follow how categories were formed. A category table with definitions and example units is standard and expected.

Frequencies may be reported where useful, but should not become the finding — that would be summative content analysis, which is a different approach.

Conventional content analysis against reflexive thematic analysis

Conventional content analysisReflexive thematic analysis
PurposeDescribe a phenomenon where theory is thinInterpret patterned meaning
Categories or themesCategories, built bottom-up from codesThemes with a central organising concept
Latent meaningOptional, must be declaredExpected, at least in part
Researcher subjectivityMinimised through procedureTreated as a resource
Frequency reportingAcceptable and commonDiscouraged
Multiple codersCommon, often with agreement checksNot for agreement
Typical outputA category scheme with definitionsAn argument built from themes

Worked example: a category scheme built upward

A study of how patients experienced a newly introduced remote-monitoring service used conventional content analysis, appropriately — nothing had been published on this service and there was no framework to test.

The unit of analysis was the meaning unit: a passage expressing one idea, which could run from a clause to a paragraph. That choice was stated, because a sentence-level unit would have fragmented accounts that only made sense across several sentences.

Open coding produced 187 codes, many in participants' own words: "the machine tells them before I do", "phoning about a number", "no one to ask about the number". Grouping produced subcategories — anticipated contact, unexplained readings, absent interpretation — and then three categories: being monitored rather than cared for, data without meaning, and the disappearance of the appointment.

The analysis stopped there, and stopping was the right call. The categories describe the experience clearly and are directly usable by the service. What the study did not do was claim an interpretive account of what remote monitoring means for the patient relationship — that would have required a different method and a different design, and asserting it on this analysis would have overreached.

The reported scheme gave each category a definition, its subcategories, and two example meaning units, which is what allows a reader to judge whether the categories hold.

Common mistakes

  • Not naming the approach. Citing Hsieh and Shannon without saying conventional, directed or summative leaves the method unspecified.
  • Unstated unit of analysis. It determines what the codes can be, and it is omitted more often than not.
  • Categories that are topics from the interview guide. If the categories reproduce the questions, the data has been sorted rather than analysed.
  • Silent drift into directed analysis. Starting from literature-derived codes while claiming a conventional approach misdescribes the study.
  • Frequency as finding. Counting is summative content analysis; presenting counts as the result of a conventional analysis conflates the two.
  • Claiming interpretive depth. Conventional content analysis is descriptive by design, and a discussion section that reads as latent interpretation needs the analysis to have supported it.
  • No trustworthiness account. Elo and colleagues specify what to report; most studies still do not.

How quality is judged

The expected reporting is more explicit than in interpretive traditions: state the approach, the unit of analysis, whether latent content was included, how categories were formed, and how the scheme was checked against the data.

A category table with definitions and exemplar units is the standard evidence, and its absence is a reasonable reason for a reviewer to ask for revision. Where multiple coders were used, say what their role was — consistency checking is coherent within this method in a way it is not within reflexive TA.

Limitations

The method is descriptive and does not generate theory. Studies wanting theory should use grounded theory; studies wanting interpretive depth should use thematic or phenomenological approaches, and choosing content analysis for its procedural comfort while wanting those outputs is the field's recurring mistake.

It is also weak on context. Segmenting into meaning units and grouping them across participants loses the shape of the individual account, and there is no case axis to recover it from.

And its apparent objectivity is partly illusory. Deriving categories "from the data" still requires judgement at every grouping, and the procedural vocabulary can obscure how much interpretation went into a category scheme.

Where software helps

The organisation phase is hierarchical grouping of a large code set, revised repeatedly — 187 codes into subcategories into categories, with the whole scheme checked back against the material each time it changes. That is mechanical work that a tool absorbs entirely.

The reporting requirement helps too: a category table with definitions and exemplar units is a by-product of coding in a platform and a separate chore otherwise. Evidano supports conventional qualitative content analysis as a named methodology and derives categories from the material rather than applying a stored framework. Choosing the unit of analysis, and deciding whether the analysis stops at description, remain the researcher's calls — and the second one is where most studies go wrong.

Topics

  • qualitative content analysis
  • conventional content analysis
  • inductive categories
  • coding
  • nursing research
  • health research
  • directed content analysis

Other methods in content and framework approaches

Written guides are linked directly; the rest have a reference entry in the methodology directory.

Published research using these methods

Studies and evaluations where this family of method was applied with Evidano — the work, not the claim.

Keep reading

Browse all articles
Company
About
Newsletter

Product updates, research, and tips — straight to your inbox.

© Evidano, All Rights Reserved.