Research MethodsDevelopmental and Utilization-Focused Evaluation

Formative Evaluation: evaluating to improve, while it still can

Evidano6 min read

Formative evaluation exists to make a programme better while it is still running — the counterpart to summative evaluation, which judges it after the fact. Scriven’s original distinction is usually glossed with the cook’s version: when the cook tastes the soup, that is formative; when the guests taste it, summative. The gloss hides the hard part. Tasting is easy; the method is in deciding what to taste, feeding the result back fast enough to matter, and keeping the improvement mission from corrupting the evidence. A formative evaluation that only finds encouragement is not formative — it is marketing with instruments.

What makes an evaluation formative

Purpose: the primary audience is the programme team, and the intended use is revision — of delivery, materials, targeting, or the theory itself. The report’s success measure is changes made, not judgements rendered.

Timing: cycles are short and scheduled against decision points. Findings that arrive after the design freeze are summative by accident.

Scope: formative work concentrates where improvement is possible — implementation quality, participant response, early outcome signals — rather than on impact claims the timeline cannot support.

The distinction with summative evaluation is a distinction of function, not method: the same interview can serve either. The origin is Scriven’s The Methodology of Evaluation (in Perspectives of Curriculum Evaluation, Rand McNally); the useful modern statements come from implementation science, where Stetler and colleagues’ The role of formative evaluation in implementation research distinguishes developmental, implementation-focused, progress-focused, and interpretive formative work across a project’s life.

When formative evaluation is the right investment

  • Pilots and first deployments, where the design is explicitly provisional and the cheapest failures are the early ones.
  • Complex or novel interventions, whose weak points cannot be predicted from the desk — the evaluation is the reconnaissance.
  • Scale-ups into new contexts, where “works there” meets “different here” and adaptation needs evidence, not improvisation.
  • Long programmes with real decision points — annual redesigns, curriculum revisions — that can consume findings on schedule.
  • Not when no one can change anything. A locked protocol, a fixed contract, an ending programme: formative findings without revision authority are documentation of regret.

Running formative cycles that actually form

Map the decision calendar first

List the moments the programme can change — staff training refresh, materials reprint, next cohort’s intake — and design each inquiry cycle to land evidence just before one. The calendar, not the methodology, sets the pace.

Choose few questions, tied to malleable things

Each cycle asks two or three questions about aspects the team can actually alter. “Is the referral form usable by frontline staff?” is formative gold; “does the programme reduce recidivism?” is a different study.

Use methods sized to the cycle

Short interviews, session observation, participant pulse feedback, delivery-log analysis — rigorous but rapid. Sampling favours variation (struggling sites as much as flagships) because improvement lives where the problems are.

Feed back in decisions, not reports

The unit of delivery is a working session with the team: findings, options, decisions minuted, owners assigned. A memo trail replaces the doorstop report; the full write-up can consolidate cycles later.

Log changes and re-test them

Every acted-on finding creates the next cycle’s question: did the fix work? The change log — finding, decision, revision, re-test — is the evaluation’s spine and, eventually, its summative gift: an account of what the programme learned.

Worked example: a benefits-navigation service in its pilot year

A charity piloted a benefits-navigation service with a twelve-month formative evaluation, cycles aligned to quarterly redesign meetings. Cycle one asked where clients stalled: session observation and case-note analysis found the stall was pre-service — the referral form’s consent language read as a fraud warning, and a third of referred clients never booked. The quarterly meeting rewrote the form; bookings from referral rose within six weeks, verified in cycle two’s logs.

Cycle two’s interviews surfaced a subtler problem: navigators were resolving urgent claims brilliantly and quietly dropping the “financial resilience” half of the model — not from overload alone, but because they judged it patronising as scripted. Rather than enforce fidelity, the team treated the adaptation as evidence about the design, rebuilt that module around client-set goals, and cycle three re-tested it: uptake of the rebuilt module tripled, and navigator notes stopped showing avoidance.

The year-end consolidation reported outcomes signals with proper humility — too early for impact claims — but delivered the change log: eleven findings, nine acted on, seven verified as improvements. The funder’s summative evaluation, commissioned for year three, inherited a programme whose weakest parts had already been found and fixed — which is what formative money buys.

Common mistakes

  • Findings after the decision. Cycles paced by researcher convenience rather than the programme’s calendar — accurate, late, useless.
  • Improvement bias. Sampling flagship sites, interviewing enthusiasts, reporting encouragement; the mission corrupts the evidence unless the design defends against it.
  • Everything questions. Ten simultaneous inquiry lines per cycle, none deep enough to act on.
  • Recommendations without owners. Feedback sessions that end in agreement and no assigned changes — the change log exists to expose this.
  • Formative drift into summative claims. Early-signal outcome data quoted as impact in the annual report; the evaluator’s job includes policing that boundary.
  • No re-test. Fixes assumed to work because they were plausible; the cheapest rigour in the method is checking.

Limitations

Formative evaluation’s closeness to the team is its engine and its exposure: the evaluator becomes part of the intervention, and independence claims should be made carefully. The honest framing is critical friendship with documented distance — variation sampling, negative-case reporting, an unedited change log.

Its evidence is provisional by design: small cycles, early signals, moving targets. That serves improvement and cannot serve accountability; commissioners wanting both need two designs, sequenced, not one evaluation asked to be soup-taster and dinner critic simultaneously.

And it consumes organisational attention. Programmes in crisis-mode delivery may be unable to absorb quarterly redesign, and a formative evaluation a team has no capacity to use is a cost with no mechanism.

Where software helps

Formative cycles live or die on turnaround: interviews and open-text feedback analysed in days, not months, so findings reach the quarterly meeting. Evidano compresses exactly that step — transcribe the cycle’s interviews, code them with quote-linked themes, compare against last cycle’s — and the change log’s “verified” column gets real evidence behind it.

The judgement calls — what is malleable, when adaptation is design evidence rather than drift, which findings deserve the meeting’s scarce attention — belong to the evaluator and the team. Speed is the tool’s contribution; discipline remains yours.

Topics

  • formative evaluation
  • formative vs summative evaluation
  • programme improvement
  • feedback loops
  • pilot evaluation
  • iterative improvement

Other methods in developmental and utilization-focused evaluation

Written guides are linked directly; the rest have a reference entry in the methodology directory.

Keep reading

Browse all articles