Site Logo
All articles
Research MethodsDevelopmental and Utilization-Focused Evaluation

Developmental Evaluation: evaluating something still being invented

Evidano9 min read

Formative evaluation improves a model. Summative evaluation judges it. Developmental evaluation exists for the situation before either applies — when there is no model yet, the intervention is being invented in contact with a problem nobody has solved, and the useful question is not "is this working?" but "what are we learning, and what should we try next?". It is the most easily abused approach in the evaluation repertoire, because a method whose defining feature is adaptation is also a convenient cover for never being held to anything.

What makes it developmental

Developmental evaluation is defined by the conditions it is used in, not by its techniques. It uses interviews, observation, data dashboards and whatever else fits — the difference is the purpose and the timing. Its purpose is to support the development of an innovation in real time; its timing is continuous, feeding back within days rather than at the end of a cycle.

Three features follow from that. The evaluator is embedded in the team, present in design conversations rather than arriving afterwards to collect data. The evaluation questions change as the innovation changes, which is a feature rather than scope creep. And the deliverable is timely feedback, not a report — the primary output of a developmental evaluation may be a two-page memo circulated the same week a decision has to be made.

Crucially, the intervention itself is expected to change during the evaluation. In a conventional design that would be a fidelity failure. Here it is the point: the innovation is supposed to become something different from what it started as, and the evaluation's job is to make that change evidence-driven rather than reactive.

Origins in utilization-focused evaluation

Michael Quinn Patton named the approach in Developmental evaluation, published in Evaluation Practice in 1994, after work with community leaders who told him plainly that the improvement-then-judgement model did not describe what they were doing. His full treatment appeared as Developmental Evaluation: Applying Complexity Concepts to Enhance Innovation and Use (2011) — reviewed in the Canadian Journal of Program Evaluation and the Evaluation Journal of Australasia, both of which are useful on where the approach is strong and where it is vulnerable.

It grew directly out of Patton's utilization-focused evaluation, which insists that evaluations should be judged by whether intended users actually use them. Developmental evaluation is that principle applied to a setting where the intended use is continuous adaptation rather than a decision at the end.

Patton has continued to develop it, including the extension to principles-focused work discussed in Emergent Developmental Evaluation Developments.

When developmental evaluation is the right call

The approach is expensive in evaluator time and produces no summative verdict. It earns that under fairly specific circumstances.

  • The intervention is genuinely being invented. Not adapted from a known model — invented, because no established approach fits the problem.
  • The environment is volatile. Policy, funding or context shifts fast enough that a fixed design would be obsolete before it reported.
  • The team can actually act on feedback. Developmental evaluation is worthless where decisions sit elsewhere or the design is contractually locked.
  • Use formative evaluation instead when there is a model to improve. If the intervention exists and the question is how to run it better, that is formative work and it is cheaper.
  • Use summative evaluation when accountability is the question. DE cannot tell a funder whether to continue funding, and should not be asked to.
  • Do not use it as cover for a vague programme. "It is complex, so we are doing developmental evaluation" is the most common misuse in the field, and it is usually a programme that has not decided what it is doing.

What the evaluator actually does

Establish the relationship and the boundaries

The evaluator negotiates a place in the team's working rhythm — which meetings, what access, what they will and will not do. The boundary that matters most is that the evaluator brings evidence and asks questions; they do not make the design decisions.

This has to be explicit and written down at the start, because the pressure to blur it arrives quickly and from both sides.

Track what is being tried and why

The core documentation task is a running record of decisions: what was changed, on what reasoning, what was expected, what actually happened. Innovation teams do not keep this, and its absence is why so many innovations cannot say what they learned.

This record is also what makes a later summative evaluation possible at all — it is the only account of what the intervention actually was at each point.

Feed back fast, in usable form

Analysis is delivered on the team's decision clock. A finding that arrives after the decision is made has no value in this mode, however rigorous.

That imposes a real trade-off: developmental findings are provisional, and the evaluator must say so clearly rather than presenting fast analysis with the confidence of finished work.

Bring the awkward evidence

The distinctive value an embedded evaluator adds is disconfirmation. Teams inventing something develop attachment to it; the evaluator's job includes saying that the data does not support the thing everyone is excited about.

A developmental evaluation that never produces an unwelcome finding has become a documentation service.

Support the transition out

Innovations that stabilise stop needing developmental evaluation. Recognising that moment and handing over to formative or summative work is part of the job, not a failure of it.

Worked example: a homelessness prevention pilot

A city funded a team to reduce repeat presentations at homelessness services, with no prescribed model and an eighteen-month horizon. A developmental evaluator joined the team two days a week.

The team's starting theory was that repeat presentations reflected a gap in follow-up after housing placement, so they built a six-week check-in service. The evaluator tracked referrals, check-in completion and re-presentations, and interviewed clients who declined the service.

By month four the pattern was uncomfortable: check-in completion was high, satisfaction was high, and re-presentation was unchanged. The interviews with decliners explained why — the people re-presenting were not those who had lost contact, but those whose tenancies had failed for reasons no check-in could address, chiefly rent arrears accumulated before placement.

The team redesigned around arrears at the point of placement. That would have been a serious failure in a fidelity-focused evaluation and was the correct response here. The evaluator documented the original theory, the disconfirming evidence, the reasoning for the change and the new theory, so that the pivot was legible as learning rather than as drift.

At month fifteen the redesigned model had stabilised. The evaluator recommended ending the developmental phase and commissioning an independent summative evaluation — which was awarded to someone else, precisely because eighteen months of embedded work had made them the wrong person to judge it.

Developmental, formative and summative compared

DevelopmentalFormativeSummative
QuestionWhat should this become?How can this work better?Did this work?
State of the interventionBeing inventedExists, needs refiningFixed and stable
Evaluator positionEmbedded in the teamClose but externalIndependent
QuestionsChange as the work changesLargely fixedFixed in advance
OutputRapid feedback, running decision recordImprovement recommendationsJudgement of merit or worth
Intervention change during evaluationExpectedLimited and trackedThreatens validity
Primary audienceThe innovation teamProgramme managersFunders and decision-makers

The independence problem

Embedding an evaluator in a team creates an obvious conflict: they help shape decisions and then assess them. Developmental evaluation does not resolve this, and practitioners who claim it does should be treated with suspicion.

What responsible practice does instead is manage it explicitly.

  • Declare the role. Every output should state that the evaluator was embedded and what that means for the findings' status.
  • Keep the decision line. The evaluator supplies evidence and questions; the team decides. Written down at the outset, revisited when it slips.
  • Separate the summative evaluation. Whoever ran the developmental phase should not judge the result. This is the single most important safeguard.
  • Build in external challenge. A periodic review by someone outside the team catches the drift that embedded work is prone to.
  • Document disconfirmation. A record showing the evaluator regularly brought unwelcome findings is the best available evidence that capture did not occur.

Common mistakes

  • Using it to avoid accountability. The commonest abuse. Adaptation becomes the explanation for every unmet expectation.
  • No decision record. Without documented reasoning, adaptation and drift are indistinguishable afterwards.
  • Evaluator becomes team member. Gradual absorption into delivery is the failure mode the independence safeguards exist for.
  • Feedback too slow. Delivered on a conventional reporting cycle, developmental evaluation is just an expensive formative evaluation.
  • Never transitioning out. Interventions that have stabilised need a different kind of evaluation; continuing developmentally protects them from scrutiny.
  • Applying it to a stable programme. Where a model already exists, the approach adds cost and removes rigour.

Limitations

Developmental evaluation produces no summative judgement and cannot be substituted for one. It also produces findings that are provisional by construction, generated fast on incomplete data, and their usefulness depends on the evaluator being candid about that.

It depends heavily on the individual. The approach asks for someone who is analytically rigorous, comfortable in a team's internal politics, fast, and willing to be unpopular — and its quality varies more with the person than most methods do.

It is also poorly suited to accountability-heavy funding environments. A funder requiring predefined indicators and quarterly variance reporting has, in effect, ruled it out, and pretending otherwise sets the evaluation up to fail on criteria it was never designed to meet.

Where software helps

The distinctive data here is the decision record plus a continuous stream of interviews, observation notes and team reflections generated over months. Its value is longitudinal: what did we believe in month three, what changed it, what did we believe in month nine.

Answering that in month fifteen means searching and coding a corpus nobody had time to organise as it accumulated — which is exactly when it stops happening. Keeping the material in a qualitative platform from the start makes the retrospective account possible rather than notional. Evidano supports developmental evaluation as a named methodology. The judgement about when the evidence is strong enough to change the design is the team's, and it should stay with people who will live with the consequences.

Topics

  • developmental evaluation
  • complexity
  • adaptive management
  • utilization-focused evaluation
  • evaluation
  • innovation

Other methods in developmental and utilization-focused evaluation

Written guides are linked directly; the rest have a reference entry in the methodology directory.

Keep reading

Browse all articles
Company
About
Newsletter

Product updates, research, and tips — straight to your inbox.

© Evidano, All Rights Reserved.