Contribution tracing exists to fix a specific complaint about contribution analysis: it produces a plausible story and no way to say how confident anyone should be in it. The response, developed by Barbara Befani and Gavin Stedman-Bryce, is to borrow the evidentiary machinery of process tracing and make the probability judgements explicit — stating up front how likely a piece of evidence is to be found if the claim is true, and how likely it is to be found anyway. The output is a contribution claim with a defensible confidence level attached, and a clear account of what would have changed it.
What contribution tracing adds
Contribution analysis assembles a causal narrative and asks whether a reasonable person would find it convincing. That is a real method and it works, but the verdict is holistic and two competent evaluators can reach different conclusions from the same file without either being able to say where they diverged.
Contribution tracing localises the disagreement. Before evidence is collected, the team specifies for each candidate piece two quantities: sensitivity — the probability of finding this evidence if the contribution claim is true — and type I error, the probability of finding it even if the claim is false. Evidence with high sensitivity and low type I error is probative; evidence with high values on both is decoration.
Bayes' rule then does the arithmetic. A prior confidence is stated, evidence is gathered, and the posterior is calculated rather than asserted. The essential discipline is that the numbers are fixed before looking, which is what stops the exercise becoming a rationalisation of a conclusion already reached.
Origins
The method was set out by Barbara Befani and Gavin Stedman-Bryce in Process Tracing and Bayesian Updating for impact evaluation, developed through work with Oxfam on advocacy and governance programmes where conventional impact evaluation had nothing to offer.
Its immediate ancestor is Befani and John Mayne's Process Tracing and Contribution Analysis: A Combined Approach, which argued the two methods were complements: contribution analysis supplies the theory of change and the narrative frame, process tracing supplies the evidentiary rigour that theory-based evaluation had been criticised for lacking.
The underlying Bayesian apparatus comes from Fairfield and Charman's Explicit Bayesian Analysis for Process Tracing. The approach has since been applied well beyond development advocacy — see for example a Bayesian process-tracing analysis of Amazon deforestation policy.
The procedure
Write a single, precise contribution claim
One claim, one outcome, specific enough to be wrong. "The programme contributed to improved governance" cannot be traced; "the programme's district scorecards caused the provincial health office to reallocate the 2025 supervision budget toward the four lowest-scoring districts" can.
Contribution tracing handles one claim at a time. A portfolio evaluation traces the two or three claims that matter, not all of them.
Set the prior
State the confidence the team holds before new evidence, and say where it comes from — prior evaluations, the strength of the theory of change, base rates for this kind of influence. A prior of 0.5 is a legitimate statement of ignorance, not a default to reach for automatically.
Priors should be set collectively and recorded. Where team members differ, the range is informative and can be carried through the analysis.
Identify candidate evidence and score it
Brainstorm what evidence could exist. For each item, agree its sensitivity and its type I error, on the record, before collection.
This is where the method earns its keep. Teams routinely discover that most of their intended evidence — partner reports, staff testimony, activity records — has a type I error near one: it would exist whether or not the claim were true, and therefore proves nothing.
Prioritise and collect
Rank the candidate evidence by how much it would shift confidence and collect the top of the list first. This is a practical strength: it directs a limited fieldwork budget at the few items that can actually change the answer.
Record failures to obtain evidence explicitly. Predicted evidence that should exist and cannot be found is itself an update.
Update and report
Apply Bayes' rule to move from prior to posterior. Report the posterior with the full chain — prior, each item, its agreed probabilities, the resulting shift — so a sceptical reader can substitute their own numbers and see what happens.
A sensitivity check is standard practice: if plausible alternative probability assignments flip the conclusion, say so.
Scoring evidence before collecting it
| Candidate evidence | Sensitivity | Type I error | Worth collecting? |
|---|---|---|---|
| Programme staff say the scorecards were influential | High | High | No — they would say this either way |
| Scorecards appear in the provincial office document register | High | Moderate | Weakly — registers log much that is never read |
| Budget reallocation memo cites scorecard district rankings by name | Moderate | Very low | Yes — few rival explanations produce this |
| A finance officer with no programme relationship recalls the rankings driving the discussion | Moderate | Low | Yes |
| The reallocated districts match the four lowest-scoring ones exactly | High | Low | Yes — but check whether they are also the poorest districts |
| Partner newsletter claims credit | High | Very high | No |
Worked example: continuing the scorecard claim
The team set a prior of 0.4, reasoning that the provincial office had reallocated budgets before on other grounds and that two other agencies were also pressing for equity-based allocation.
Two pieces of evidence were prioritised. The budget memo was obtained and did cite the scorecard rankings by name — agreed sensitivity 0.6, type I error 0.05. That single item moved confidence from 0.4 to roughly 0.89.
The independent finance officer's recollection was sought and obtained, but with a qualification: the rankings had been discussed alongside a separate donor's equity index that pointed to three of the same four districts. The team had scored this item at sensitivity 0.5 and type I error 0.2; the qualification led them to revise type I error upward to 0.4 before applying it, and to record why. Posterior settled near 0.92.
A pre-agreed disconfirming check was also run: had the reallocation been drafted before the scorecards were delivered? The dated drafts showed it had not. Had that check failed, the team had agreed in advance it would have collapsed confidence below the prior.
The reported claim was: high confidence (≈0.9) that the scorecards were a substantive input to the reallocation, with an explicit note that a second agency's equity index pointed the same way for three of four districts and that the two influences cannot be separated on this evidence. The numbers are not the point; the auditable reasoning behind them is.
Common mistakes
- Scoring evidence after seeing it. This destroys the method entirely. If the probabilities are assigned once the document is in hand, the posterior is a decorated opinion.
- Claims too broad to trace. Anything at the level of "strengthened civil society" cannot be scored, because no specific evidence is diagnostic of it.
- Collecting the easy evidence. Programme documentation and staff testimony are cheap and have type I errors near one. The method exists to push a team past them.
- Treating the posterior as a measurement. It is a structured expression of judgement. Reporting 0.87 rather than "high confidence" implies a precision the inputs do not support.
- Ignoring disconfirmation. Specifying at least one piece of evidence whose absence would substantially reduce confidence is what separates this from advocacy.
- Skipping the sensitivity check. If the conclusion depends on one contested probability, the reader needs to know.
Quality criteria
A defensible contribution trace shows: a single precise claim; a prior with a stated basis; probability assignments recorded before collection with the names of who agreed them; at least one item capable of reducing confidence; the full updating chain; and an honest account of evidence sought but not obtained.
The strongest signal of quality is a report where the posterior is lower than the commissioning organisation hoped. The method is designed to be capable of that outcome, and a body of contribution traces that never produces one is not being run properly.
Limitations
The probabilities are judgements. Bayes' rule is arithmetic applied to numbers people made up, and the arithmetic cannot rescue bad inputs. What it does provide is transparency about which input a disagreement turns on — a real gain, but not the same as objectivity, and it should not be presented as one.
The method is also narrow by design: one claim at a time, at meaningful cost per claim. It is unsuited to portfolio-wide questions and to any programme whose value lies in the accumulation of many small effects rather than a few identifiable ones.
Finally, it depends on documentary access. Where records are closed, the high-sensitivity low-type-I-error evidence usually cannot be obtained, and the analysis is left updating on exactly the weak testimony it was built to avoid.
Where software helps
The updating arithmetic is a spreadsheet. The hard part is the evidence corpus: dozens of documents and transcripts, each linked to a specific claim, a specific probability assignment, and the date it was scored.
Keeping that auditable is a document-analysis problem, and it is where a qualitative platform earns its place — retrieving every passage bearing on one claim across a large corpus, and preserving the link between a scored item and the source text a reviewer will want to check. Evidano supports contribution tracing as a named methodology. The probability judgements are the team's, made in a workshop, on the record, before the evidence arrives.
Topics
- contribution tracing
- bayesian updating
- process tracing
- contribution analysis
- evaluation
- impact evaluation
Other methods in realist and causal analysis approaches
Written guides are linked directly; the rest have a reference entry in the methodology directory.
Keep reading
- Research MethodsContribution Analysis: a step-by-step guideHow to run a contribution analysis: build the contribution story, test it against evidence, and address rival explanations when no counterfactual exists.
- Research MethodsProcess Tracing: testing causal mechanisms in a single caseHow process tracing establishes causation without a comparison case: the four evidentiary tests, what counts as diagnostic evidence, and where the method is misapplied.
- Research MethodsRealist Evaluation: context, mechanism, outcomeWhat realist evaluation actually asks, how to build and test a CMO configuration, where the mechanism concept gets misused, and how realist quality is judged.
