Site Logo
Research MethodsEvidence Synthesis

Systematic Literature Review: a step-by-step guide

Evidano8 min read

A systematic literature review answers a defined question by finding, screening, and synthesising every study that meets pre-stated criteria — and documents the process well enough that someone else could repeat it. That last clause is the whole difference between a systematic review and an ordinary literature review. An ordinary review is an argument supported by studies the author happened to select; a systematic review is a method, with a protocol written before the searching starts and an audit trail from database query to conclusion. Most reviews that fail peer review fail on process, not on reading: no protocol, an unreproducible search, or screening decisions nobody recorded.

What makes a review systematic

Three commitments separate systematic reviews from narrative ones. The question is fixed in advance — usually structured as PICO (population, intervention, comparator, outcome) or a variant like SPIDER for qualitative questions — so the evidence cannot be quietly re-scoped to fit the studies found.

The search is exhaustive and documented. Multiple databases, stated search strings, stated dates, plus grey literature and reference chasing. The test is reproducibility: another team running your strings on your dates should retrieve the same records.

Every screening and extraction decision is recorded. Titles and abstracts are screened against the inclusion criteria, ideally by two reviewers independently; disagreements are resolved and counted; the flow from records identified to studies included is reported in a PRISMA diagram. The reader can see exactly where the 4,212 records became 23 studies.

Where the method comes from

Systematic reviewing grew out of evidence-based medicine in the 1980s and 1990s, institutionalised by the Cochrane Collaboration, whose Cochrane Handbook for Systematic Reviews of Interventions (Wiley) remains the fullest procedural reference.

The reporting standard is PRISMA 2020, set out by Page and colleagues in The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. Journals increasingly require the checklist and the flow diagram, so it is easier to write toward PRISMA from day one than to retrofit it.

For orientation among the many review types — systematic, scoping, rapid, umbrella, and ten others — Grant and Booth’s A typology of reviews is the standard map, and quality of conduct (as opposed to reporting) is assessed with AMSTAR 2, published by Shea and colleagues in AMSTAR 2: a critical appraisal tool.

When a systematic review is the right choice

  • Use it when the question is answerable and the literature is mature. “Does mentoring reduce early-career teacher attrition?” is a systematic review question. “What is known about teacher wellbeing?” is a scoping review question.
  • Use it when the stakes justify the cost. A competent systematic review is months of work for a small team. Guidelines, policy decisions, and dissertations justify that; a background chapter usually does not.
  • Do not use it to showcase a position. If the conclusion is decided before the search, the systematic apparatus only decorates a narrative review.
  • Do not use it on a field too young to synthesise. Four heterogeneous studies produce a review that can only conclude “more research is needed” — a scoping review would have said more, for less.

The seven steps, in order

1. Write the protocol before searching

State the question, the inclusion and exclusion criteria, the databases, the appraisal tool, and the planned synthesis. Register it — PROSPERO for health topics, or the Open Science Framework elsewhere — so reviewers can check the review against its own plan.

The protocol is what protects you from yourself: without it, criteria drift toward the studies you wish you had found.

2. Build and run the search

Translate the question into search strings combining controlled vocabulary (MeSH or equivalent) with free-text synonyms, and run them across at least two or three databases relevant to the field. Record the exact string, database, and date for each search.

Supplement with backward and forward citation chasing on included studies and a defined grey-literature pass. A librarian or information specialist at this stage repays their time many times over.

3. Screen in two passes

De-duplicate, then screen titles and abstracts against the criteria, then screen the surviving full texts. Two independent screeners with a recorded agreement rate is the standard; a second reviewer on even a 20% sample is far better than none.

Keep a reasons log for full-text exclusions — PRISMA requires it, and it is the part reviewers check first.

4. Extract into a piloted form

Design an extraction table — study, setting, sample, design, measures, findings — and pilot it on three to five studies before committing. Piloting always changes the form, and changing the form after extracting thirty studies means re-extracting thirty studies.

5. Appraise study quality

Assess each included study with a tool matched to its design: RoB 2 for randomised trials, ROBINS-I for non-randomised studies, CASP checklists for qualitative work. Appraisal is not a gate for deletion; it feeds the synthesis, where strong and weak evidence should carry different weight.

6. Synthesise

Meta-analysis if the studies are similar enough in design, measures, and populations to pool; otherwise a structured narrative synthesis around the questions, with the SWiM guideline governing how non-pooled findings are reported. Declare which one you are doing and why — “we could not meta-analyse, so we describe” is a finding about the literature, not an apology.

7. Report against PRISMA

The flow diagram, the checklist, the search strings in an appendix, and the extraction table as supplementary material. A reader should be able to trace any sentence in the discussion back through the synthesis to named studies.

Worked example: school-feeding programmes and attendance

A two-person team reviewed the effect of school-feeding programmes on attendance in low- and middle-income countries. The protocol, registered on PROSPERO, fixed the population (primary-school pupils), the intervention (any on-site meal or take-home ration programme), the outcome (enrolment or attendance), and the designs admitted (experimental and quasi-experimental).

Searches in ERIC, EconLit, Scopus, and two regional databases returned 3,847 records; 62 survived to full text; 19 were included, with the reasons for the other 43 exclusions logged (most commonly: no attendance outcome, or no comparison group). Both reviewers screened independently and disagreed on 41 abstracts — almost all resolved by re-reading the criteria rather than by argument, which is the protocol doing its job.

Twelve studies were similar enough to pool, showing a modest average attendance gain concentrated in the poorest districts; the remaining seven entered a narrative synthesis that explained the pattern — programmes layered on top of fee abolition showed no additional effect. The headline finding was therefore conditional, and the review could say precisely on what.

Common mistakes

  • No protocol, or a protocol written after the search. Everything downstream inherits the suspicion.
  • A one-database search. Coverage differs sharply between databases; a Scopus-only search is a convenience sample of the literature.
  • Criteria that drift. If “we excluded studies without a comparison group” quietly becomes “except two that were interesting”, the review is no longer systematic.
  • Skipping appraisal, or appraising and then ignoring it. A synthesis that weights a weak pre-post study equally with a strong trial misleads politely.
  • Vote counting. “Six studies found an effect and four did not” is not synthesis; it discards effect size, precision, and quality in one move.
  • Writing the discussion from memory. Every claim should trace to the extraction table, not to the reviewer’s impression of the literature.

Limitations

A systematic review is only as good as the literature it inherits: publication bias, English-language dominance, and outcome switching in primary studies pass straight through unless explicitly assessed. Funnel plots and grey-literature searching mitigate; they do not cure.

The method is slow — median times run well over a year — and the literature keeps moving while you review it, which is the problem rapid reviews and living reviews exist to manage.

And the rigour is procedural, not interpretive. A review can be PRISMA-perfect and still ask a question nobody needed answered. The protocol disciplines the process; it cannot supply the judgement.

Where software helps

Screening and extraction are the mechanical heart of a systematic review, and they are exactly where tooling pays: de-duplication, blinded double screening with agreement tracking, and a piloted extraction form that every study passes through.

AI-assisted analysis earns its place at full-text stage. Evidano applies a codebook — your extraction fields — across a corpus of included PDFs and returns each field with the supporting quote linked to its source document, which turns extraction checking from re-reading into verification. Teams have used it to code hundreds of evaluation reports for synthesis at 92% agreement with human coders in a published UN comparison. The protocol, the criteria, and the appraisal judgements stay yours; no tool can decide what counts as evidence.

Topics

  • systematic literature review
  • systematic review
  • how to write a literature review
  • literature review example
  • PRISMA
  • inclusion and exclusion criteria
  • evidence synthesis
  • literature review vs systematic review

Other methods in evidence synthesis

Written guides are linked directly; the rest have a reference entry in the methodology directory.

Published research using these methods

Studies and evaluations where this family of method was applied with Evidano — the work, not the claim.

Keep reading

Browse all articles