Site Logo
Research MethodsThematic Analysis

Semantic Coding: staying at the surface on purpose

Evidano6 min read

Semantic coding is thematic coding that stays with what was actually said: codes capture the explicit, surface meaning of the data — participants’ stated experiences, opinions, and reports — without reaching for what might lie beneath. In Braun and Clarke’s vocabulary it is one of two levels at which themes can be built, the other being latent. Choosing the semantic level is not choosing the easy version; it is a commitment with its own discipline (stay tethered to the said), its own strengths (transparency, checkability, speed at scale), and its own characteristic failure (paraphrase mistaken for analysis). Most applied research — evaluations, needs assessments, service feedback — is semantic coding territory, and is better for owning it.

What “semantic” commits you to

A semantic code names something present in the data’s explicit content. If a participant says “I stopped going because the bus route changed”, semantic codes are available for transport barriers and discontinued attendance; a code for institutional abandonment is not — nothing in the utterance said it.

The commitment runs through the whole analysis: themes built from semantic codes summarise patterned explicit content, and the write-up’s claims stay at the level of what participants reported. Progression within the level is from description toward organised meaning — grouping, comparing, and interpreting the significance of what was said, not decoding what was meant.

The distinction with latent coding is a decision, not a discovery. The same dataset supports both; the research question dictates the level, and the methods section should state the choice and hold to it.

Anchors in the literature

The semantic/latent distinction is drawn in Braun and Clarke’s foundational Using thematic analysis in psychology, which remains the necessary citation and the clearest statement of what each level claims.

For a worked demonstration of the full analytic arc — including how semantic themes are developed and named — Byrne’s A worked example of Braun and Clarke’s approach to reflexive thematic analysis is the most useful single paper to imitate.

The neighbouring content-analysis tradition calls approximately the same surface “manifest content”; Graneheim and Lundman’s Qualitative content analysis in nursing research is the standard source for that vocabulary and its quality concepts.

When the semantic level is the right choice

  • When stakeholders need findings they can verify. Semantic themes trace visibly to quotes; a programme board can check the chain themselves.
  • When the question is about stated experience: reported barriers, expressed needs, described events. The said is the object, so the surface is the site.
  • When coding must be consistent across a team or at scale. Surface meaning supports shared definitions and agreement checks in a way latent interpretation resists.
  • When the corpus is large and heterogeneous — hundreds of open-ended survey responses reward systematic surface coding and punish deep-reading ambitions.
  • Not when the interesting content is unsayable or unsaid — stigma, ideology, identity work. Forcing those questions through semantic coding yields themes about euphemisms.

Doing semantic coding well

Define codes by inclusion rules, not vibes

Each code gets a name, a one-line definition, and an inclusion rule stated in terms of explicit content (“participant states a cost-related reason for non-attendance”). Rules are what make the surface level actually checkable.

Code exhaustively, close to the words

Work systematically through the corpus; label everything relevant to the question, staying near participants’ own terms. Resist premature abstraction — “transport barrier” is semantic; “structural exclusion” has left the level.

Collate and build candidate themes

Group codes into candidate themes that capture patterned explicit meaning. A semantic theme should be nameable in words participants could recognise — a good test of level discipline.

Review against the coded extracts and the whole

Check each candidate theme against its extracts (does the pattern hold?) and against the full dataset (what did it miss?). Negative instances — participants who explicitly said otherwise — are reported, not smoothed.

Interpret significance without switching levels

The discussion can and should say why the patterns matter — but as interpretation of reported content, flagged as such. If a latent reading becomes irresistible, add it as a labelled second pass, not a silent drift.

Worked example: why patients miss telehealth appointments

A clinic analysed 240 open-text survey responses and 18 short interviews about missed video appointments. The question — what do patients report as reasons — was explicitly semantic, and the codebook’s inclusion rules required stated content.

Coding produced 31 codes consolidating into four semantic themes: technology failures at join time (stated by 41% of coders’ units), appointment times colliding with caregiving, not seeing the point for check-ins (“nothing to examine”), and privacy at home (“nowhere in the flat to talk about this”). Each theme’s name stayed within the reported; each carried counts, verbatim exemplars, and its negative cases (six patients explicitly preferred video for privacy — reported alongside).

The level discipline mattered at write-up. A draft sentence reading “appointments were missed because telehealth erodes the felt legitimacy of care” was cut: nothing in the data said it, however plausible. The delivered version — patients state they skip appointments that feel purposeless without examination — supported the same service change (repurposing check-ins) while claiming only what 240 people had actually written.

Common mistakes

  • Paraphrase presented as themes. “Patients mentioned technology problems” restates data; a theme organises it — prevalence, variants, conditions, exceptions.
  • Level drift. Codes that start semantic and quietly become interpretive mid-corpus, so early and late transcripts are coded to different standards.
  • Topic coding mistaken for meaning coding. “Transport” as a bucket for everything mentioning buses mixes complaints, praise, and asides; semantic codes capture stated meaning, not vocabulary.
  • Counting without context. Semantic coding supports counts, but a frequency table with no analysis of variation within codes wastes the qualitative data.
  • Apologising for the level. Semantic analysis is not failed latent analysis; hedging it as “merely descriptive” undersells work that was, correctly, built to be checkable.

Limitations

The level’s honesty is its ceiling: semantic analysis cannot speak to what participants could not or would not say, and on sensitive topics the surface is systematically curated. The method reports the presented self.

Themes can also inherit the question’s framing — ask about barriers and the data will contain barriers — so instrument wording belongs in the interpretation.

And checkability invites false confidence: agreement on surface codes is achievable, but agreement is not validity, and a reliably applied shallow codebook yields reliable shallowness. The remedy is a codebook built from the data’s own range, revised as coding teaches.

Where software helps

Semantic coding is the level AI assists best, precisely because the standard is explicit content: rule-based codes over stated meaning are what Evidano applies across interview and survey corpora, with every application linked to its quote so the checkability the level promises is real — reviewers audit by clicking, and frequency and cross-segment comparisons come with the evidence attached. Published comparisons of exactly this kind of codebook-driven coding report agreement with expert human coders in the 92–96% range.

The researcher still owns the codebook’s categories, the theme construction, and the significance argument. The tool guarantees the tether to the said; what the said amounts to remains yours to argue.

Topics

  • semantic coding
  • thematic coding
  • manifest content
  • qualitative coding
  • surface-level coding
  • descriptive coding
  • coding transcripts

Other methods in thematic analysis

Written guides are linked directly; the rest have a reference entry in the methodology directory.

Published research using these methods

Studies and evaluations where this family of method was applied with Evidano — the work, not the claim.

Keep reading

Browse all articles