Realist evaluation replaces the question "does this programme work?" with a longer and more useful one: what is it about this programme that works, for whom, in what circumstances, in what respects, and why? The reframing is not rhetorical. It commits the evaluator to explaining outcomes through the reasoning and resources of the people involved, rather than treating the programme as a treatment applied to a passive population — and it commits them to expecting different results in different settings rather than treating that variation as noise to be averaged away.
The claim realist evaluation makes
Realist evaluation rests on a specific ontological position: programmes do not cause outcomes. Programmes offer resources, and outcomes follow from how people reason about and respond to those resources, in the circumstances they happen to be in. A microfinance loan does not cause enterprise growth; it offers capital, and what follows depends on whether the recipient reads it as opportunity or as debt exposure, which depends in turn on the market, the household, and what happened to the last person in the village who took one.
That position produces the method's signature unit of analysis, the context–mechanism–outcome configuration, usually written CMO. A CMO is a proposition: in context C, the programme resource triggers mechanism M, producing outcome O. Evaluation then means specifying, testing and refining a set of these propositions rather than estimating an average effect.
The practical consequence is that a realist evaluation does not conclude with a verdict. It concludes with a refined programme theory — a statement of the conditions under which the intervention does and does not fire, transferable to a new site in a way that an effect size is not.
Origins: Pawson and Tilley
The approach was set out by Ray Pawson and Nick Tilley in the 1990s, most fully in Realistic Evaluation (1997) and summarised in An Introduction to Scientific Realist Evaluation. Their target was the experimental orthodoxy in criminal-justice evaluation, where trials of the same intervention kept producing contradictory results across sites and the field responded by demanding larger trials.
Pawson and Tilley argued the contradiction was information, not error: the same CCTV scheme deters in a car park where offenders believe they are being watched and does nothing in one where they do not, and no sample size resolves that. Their answer was to borrow from Merton the idea of middle-range theory — abstract enough to travel between settings, concrete enough to be tested in one. Nick Tilley develops that lineage directly in The Middle-Range Methodology of Realist Evaluation.
Pawson later extended the programme in The Science of Evaluation: A Realist Manifesto, which is also the sharpest statement of what realists think is wrong with systematic review as ordinarily practised.
Choosing realist evaluation over its neighbours
Realist evaluation is expensive in analytic time and demands a theoretically confident evaluator. It is worth it under conditions that are reasonably easy to recognise.
- The same intervention has produced inconsistent results. This is the canonical realist case: the variation is the finding, and realist evaluation is built to explain it.
- The intervention is complex and human-mediated. Where outcomes run through judgement, trust, motivation or professional discretion, mechanisms are doing the work and are worth naming.
- Transferability is the real question. If the commissioner's next decision is whether to run this elsewhere, a CMO set answers it and an effect estimate does not.
- Prefer process tracing when the question is a single case. Realist evaluation compares configurations across contexts; process tracing establishes whether a specific mechanism operated in one case, and does it with a more explicit evidential test.
- Prefer contribution analysis when the deliverable is an accountability claim. Realist findings are conditional by design, which is exactly what a funder asking "did our money cause this?" does not want.
- Do not use it where the mechanism is not contested. If the intervention is a vaccine, the CMO apparatus adds vocabulary rather than knowledge.
Running a realist evaluation
Elicit the initial programme theory
Start from what designers, implementers and participants believe makes the intervention work, gathered through interviews, documents and any prior literature on similar programmes. Expect several competing theories rather than one; they become the hypotheses to be tested.
Write each as an explicit CMO proposition. A theory that cannot be written in that form is usually an activity description, not a theory.
Formalise the CMO configurations
Separate the three elements properly, which is harder than it sounds. Context is not "the setting" in general — it is the specific pre-existing condition that determines whether the mechanism can fire. Mechanism is not the programme activity — it is the change in reasoning or capacity the activity generates. Outcome is what follows, including outcomes nobody wanted.
The most common failure is a mechanism that is really an activity: "the training sessions" is an activity; "participants come to believe supervisors will act on what they report" is a mechanism.
Design the data collection to test them
Realist data collection is theory-driven and deliberately asymmetric: different respondents are asked about different parts of the configuration, because different people are in a position to know about context, mechanism and outcome respectively.
The realist interview is distinctive — the evaluator puts the theory to the respondent and invites them to correct it, rather than withholding it to avoid leading. Teacher-learner is the intended relationship, in both directions.
Test, refine and specify scope
Analysis asks which configurations held, which did not, and what the pattern of failure says about the conditions the mechanism depends on. Refinement is the normal result: a theory that survives untouched has usually not been tested hard enough.
The output should state the scope conditions plainly — the settings in which this programme theory is expected to hold, and those in which it is not.
Worked example: a hospital incident-reporting system
A hospital group introduced electronic incident reporting across eleven sites. Reporting rates rose sharply in four, marginally in five, and fell in two. A conventional evaluation would report an average increase and recommend wider rollout.
The realist evaluation elicited three candidate mechanisms from staff: the system was easier than paper (reduced effort), it allowed anonymous submission (reduced fear), and it produced visible feedback on what happened to reports (perceived efficacy). Interviews across sites then tested which was doing the work.
Reduced effort turned out to explain almost nothing — paper forms had not been the barrier. The operative configuration was: where staff had previously seen a colleague treated punitively after a report (C), anonymity plus visible non-punitive follow-up (M: it is safe to report) produced sustained increases in reporting (O). In the two sites where reporting fell, follow-up had been visible but had resulted in disciplinary action; the same system triggered the opposite mechanism, and staff correctly concluded the tool was a surveillance instrument.
The recommendation that followed was not "roll it out". It was that rollout be conditional on the site's disciplinary practice, because in a punitive culture the intervention makes things worse — a finding the average concealed entirely.
What counts as a mechanism
| Statement | Is it a mechanism? | Why |
|---|---|---|
| The programme delivered 40 workshops | No | An activity. Describes what was done, not why anyone responded. |
| Attendance was high | No | An outcome, and an intermediate one at that. |
| Participants gained confidence to challenge a supervisor | Yes | A change in reasoning triggered by a resource the programme offered. |
| The clinic was understaffed | No — this is context | A pre-existing condition that shapes whether a mechanism can fire. |
| Staff came to see the audit as a learning tool rather than a threat | Yes | A reinterpretation of the resource; the classic realist mechanism form. |
| Funding increased by 30% | No | A resource. It may trigger a mechanism; it is not one. |
Where realist evaluations go wrong
- Mechanism inflation. Every activity gets relabelled a mechanism and the CMO becomes a restatement of the logic model. Elaborating the Context-Mechanism-Outcome configuration is the clearest treatment of how the three elements come apart in practice.
- Context as scenery. Listing the country, sector and budget is not specifying context. Context is only the conditions that bear on whether the mechanism fires.
- Theory that cannot fail. Configurations pitched so abstractly that any finding confirms them. A realist proposition should identify what evidence would refute it.
- Retro-fitting. Building the CMO after the data is in, which converts the method from a test into a description with realist vocabulary.
- Counting configurations. Reporting that fourteen CMOs were "confirmed" mistakes the method for a tally. The output is an explanation, not a score.
Judging realist quality
Realist work is not assessed on sample size or inter-rater reliability. The relevant criteria are its own.
- Is the programme theory explicit and falsifiable? It should be possible to say what would have counted as disconfirmation.
- Do the configurations distinguish context from mechanism? This is the single most common point of failure and the first thing an informed reader checks.
- Was the theory refined? Evidence of the theory changing under data is evidence the test was real.
- Are scope conditions stated? A realist finding without stated limits has abandoned the method's central claim.
- Is the middle-range level right? Too abstract and it explains everything; too concrete and it transfers nowhere.
The RAMESES II reporting standards give a fuller checklist and are the usual reference point for realist reporting in health research.
Limitations worth stating
Realist evaluation does not produce effect sizes and cannot answer how much of an outcome the programme produced. Commissioners who need that number will not get it here, and pretending otherwise damages the method's standing.
It is also demanding in a way that is unevenly distributed: the quality of a realist evaluation depends heavily on the evaluator's theoretical range, and weak realist work is common precisely because the vocabulary is easy to adopt without the discipline. Mechanism identification remains partly interpretive, and two competent realists can specify different configurations from the same data — the method offers refinement and plausibility, not identification in the econometric sense.
Finally, the approach is time-hungry. A realist evaluation done properly needs iterative access to sites and respondents, which many evaluation contracts do not fund.
Where software helps, and where it does not
The analytic work in a realist evaluation — deciding what is context and what is mechanism — is not automatable, and a tool that claims to identify mechanisms is claiming something it cannot deliver.
What software does help with is the volume problem underneath. A realist evaluation across eleven sites generates a large interview corpus that has to be interrogated configuration by configuration: every passage bearing on one CMO, retrieved across all sites, compared. That is retrieval and structured coding at a scale where manual handling starts to lose things, and it is what a qualitative analysis platform is for. Evidano supports realist evaluation as a named methodology, so the analysis follows CMO structure rather than a generic coding pass — but the configurations remain the evaluator's to specify, defend and revise.
Topics
- realist evaluation
- context mechanism outcome
- programme theory
- evaluation
- realist synthesis
- process tracing
Other methods in realist and causal analysis approaches
Written guides are linked directly; the rest have a reference entry in the methodology directory.
Keep reading
- Research MethodsProcess Tracing: testing causal mechanisms in a single caseHow process tracing establishes causation without a comparison case: the four evidentiary tests, what counts as diagnostic evidence, and where the method is misapplied.
- Research MethodsContribution Analysis: a step-by-step guideHow to run a contribution analysis: build the contribution story, test it against evidence, and address rival explanations when no counterfactual exists.
- Research MethodsContribution Tracing: Bayesian confidence in contribution claimsContribution tracing combines process tracing with explicit Bayesian updating to put a defensible confidence level on a contribution claim. How it works, and its limits.
