How to Conduct a Systematic Review: A Step-by-Step Guide
A practical, step-by-step guide to conducting a systematic review — PICO, PROSPERO, search strategy, dual screening, risk of bias, and PRISMA 2020 reporting.
Step 1: Turn a Topic into an Answerable Question (PICO)
Most systematic reviews that stall do so in the first week, when a broad topic — "AI in radiology," "exercise for depression" — is mistaken for a research question. A systematic review needs a question precise enough that two independent reviewers, working from the same criteria, would make the same include-or-exclude decision on almost every record. The standard tool for getting there is PICO: Population, Intervention, Comparator, Outcome.
| Element | What to pin down | Illustrative example |
|---|---|---|
| Population | Condition, severity, age range, setting | Adults with treatment-resistant hypertension |
| Intervention | The treatment or exposure, with dose and duration where relevant | Renal denervation |
| Comparator | What the intervention is judged against | Sham procedure or usual care |
| Outcome | Prespecified, measurable endpoints | Change in 24-hour systolic blood pressure at 6 months |
Before committing, run scoping searches on PubMed to gauge the size of the literature, and check PROSPERO for registered reviews already underway on the same question. Finding a recent, well-conducted review on your exact question is not a disaster — it is a free answer. Finding an in-progress one lets you adjust scope before sinking months into duplication.
Step 2: Write a Protocol and Register It on PROSPERO
A protocol is what separates a systematic review from a literature review with good intentions. Written before screening begins, it locks eligibility criteria, outcomes, and the analysis plan in place — so decisions are driven by methods, not by results you have already glimpsed. The PRISMA-P extension of the PRISMA statement specifies what a protocol should contain: rationale, eligibility criteria, information sources, a full search strategy for at least one database, screening and extraction procedures, risk-of-bias assessment, and the planned synthesis.
Then register it on PROSPERO, the international prospective register of systematic reviews. Registration is free, creates a public timestamp of your planned methods, and signals to journals and readers that outcomes were not switched after the fact. Do it early — ideally before screening starts.
Step 3: Design a Search Strategy That Would Survive Peer Review
No single database indexes the whole biomedical literature. Each covers different journals, applies different controlled vocabularies, and indexes with different lag. A search limited to MEDLINE via PubMed will miss studies indexed only in Embase, and vice versa — which is why the Cochrane Handbook expects searches across multiple sources: typically MEDLINE, Embase, and Cochrane CENTRAL at a minimum, supplemented by trial registries such as ClinicalTrials.gov to surface completed-but-unpublished and ongoing studies.
The mechanics are consistent even where syntax differs. Build one concept block per PICO element you search on (usually population and intervention; outcomes are often left out of the search to preserve sensitivity), combine synonyms within a block with OR, and join blocks with AND. Within each block, pair controlled vocabulary — MeSH in PubMed, Emtree in Embase — with free-text title and abstract terms: controlled vocabulary catches consistently indexed records, while free text catches the newest papers that have not yet been indexed and the ones indexers classified differently.
- ·Involve a medical librarian or information specialist if you possibly can — search design is a specialist skill, and it is routinely underestimated.
- ·Save the exact search string, database, platform, date, and hit count for every search. PRISMA 2020 asks for the full strategy for each source.
- ·Have someone who did not draft the strategy review it line by line before you run the definitive search.
Steps 4–6: Deduplicate, Then Screen in Two Stages
Export all results into a reference manager and deduplicate before screening — the same trial will arrive from three databases with small formatting differences. Record how many records came from each source and how many duplicates you removed: those numbers are the top of your PRISMA flow diagram, and they are far easier to capture now than to reconstruct later.
Title and abstract screening
Two reviewers screen every record independently against the eligibility criteria, blind to each other's decisions, with disagreements resolved by discussion or a third reviewer. Before tackling the full set, calibrate: screen the same pilot batch of records, compare decisions, and sharpen the criteria wording until you understand why you disagree. Screening platforms such as AutoEvidence can prioritize likely-relevant records and keep an audit trail of every decision, but they support the dual-reviewer standard rather than replace it.
Full-text review
Retrieve full texts for everything that survives and screen again — this time recording a specific reason for every exclusion (wrong population, wrong comparator, conference abstract only, and so on) drawn from a predefined hierarchy. PRISMA 2020 requires these reasons in the flow diagram, and peer reviewers will ask about them. "Not relevant" is not a reason.
Steps 7–8: Extract Data and Assess Risk of Bias
Build a data extraction form and pilot it on a handful of included studies before extracting from all of them — the first version of any form is wrong in ways you only discover by using it. Capture study design, population characteristics, intervention details, outcome definitions, effect estimates with their measures of variance, and funding sources. The rigorous default is dual independent extraction; a common pragmatic compromise is one extractor plus a second reviewer verifying every field against the source.
Risk-of-bias assessment runs alongside extraction. For randomized trials, RoB 2 examines five domains: the randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. For non-randomized studies of interventions, ROBINS-I addresses the additional threats that come with the absence of randomization — confounding above all. Both tools are covered in depth in the Cochrane Handbook. Note that RoB 2 is applied per result, not per study: a trial can be low risk for mortality yet high risk for a subjective quality-of-life scale.
Step 9: Synthesis — When Meta-Analysis Helps and When It Misleads
Meta-analysis is a tool, not a milestone. Pooling is appropriate when the included studies ask a sufficiently similar question — comparable populations, interventions, comparators, and outcome definitions — that a combined estimate means something. When studies differ fundamentally in whom they enrolled or what they measured, pooling produces a precise-looking number that answers no actual question.
When you do pool, the choice of model matters. A fixed-effect model assumes every study estimates one identical true effect and differences are noise; a random-effects model assumes true effects genuinely vary across studies and estimates their distribution. In most clinical questions, populations and protocols really do differ, which is why random-effects is the more common default — but the choice belongs in the protocol, stated and justified, not made after seeing which result looks better. Examine heterogeneity, explore it with the subgroup and sensitivity analyses you prespecified, and resist inventing subgroups because they happen to be significant.
When meta-analysis is not defensible, say so and synthesize without it: structured tables grouped by population or intervention, effect directions, and a narrative that follows the structure laid out in the protocol. A well-reasoned narrative synthesis is a legitimate result. A forced meta-analysis is a methods flaw.
Step 10: Report with PRISMA 2020 — and Grade the Certainty
PRISMA 2020 (Page MJ et al., BMJ 2021;372:n71) is the reporting standard for systematic reviews: a 27-item checklist plus the flow diagram tracking every record from identification through screening to inclusion. Checklist and diagram templates are available at prisma-statement.org. Work through the checklist item by item before submission — many journals require it, and the flow diagram is impossible to reconstruct honestly if you did not keep counts along the way. For review types beyond intervention reviews, the EQUATOR Network indexes the relevant reporting guidelines and PRISMA extensions.
Reporting what studies found is not the same as saying how much to trust them. The GRADE approach rates certainty of evidence per outcome — high, moderate, low, or very low — starting from study design and downgrading for risk of bias, inconsistency, indirectness, imprecision, and publication bias. A summary-of-findings table built on GRADE is what turns a stack of pooled estimates into something a clinician or guideline panel can act on.
How Long It Really Takes — and What Goes Wrong
Plan in months, not weeks. The ranges below are planning heuristics for a small team — two reviewers with other commitments, a mid-sized clinical question — not benchmarks or promises. Screening effort scales with hit count: a search returning ten thousand records is a different project from one returning eight hundred.
| Phase | Realistic planning range | What stretches it |
|---|---|---|
| Question, protocol, PROSPERO registration | 4–8 weeks | Scope negotiation among co-authors |
| Search development and execution | 2–6 weeks | Librarian availability, strategy peer review |
| Title/abstract screening | 4–12 weeks | Hit volume, conflict rate |
| Full-text retrieval and review | 4–8 weeks | Paywalled papers, unresponsive authors |
| Extraction and risk of bias | 6–12 weeks | Number of included studies, outcome multiplicity |
| Synthesis, writing, PRISMA checks | 8–12 weeks | Analysis complexity, co-author review cycles |
Software compresses some of this — reference managers automate deduplication, and AI-assisted platforms such as AutoEvidence shorten the screening phase — but the protocol-driven, dual-reviewer core of the method is where a systematic review's credibility comes from, and it does not disappear. If months pass between your definitive search and submission, rerun the search and screen the new records.
Common mistakes worth avoiding
- ·Skipping the protocol and settling eligibility criteria while screening — criteria drift toward the studies you have already seen, and the review becomes unfalsifiable.
- ·Searching one database and assuming coverage. Complement PubMed with Embase, CENTRAL, and trial registries.
- ·Single-reviewer screening to save time. Every reviewer misses records; independent duplication is the error-correction mechanism.
- ·Vague exclusion logs. Record a specific, predefined reason for every full-text exclusion on the day you make the call.
- ·Pooling everything. Clinical heterogeneity is a reason not to meta-analyze, however tempting the forest plot.
- ·Treating risk of bias as decoration — assessed, tabulated, and never mentioned again in the discussion or certainty ratings.
- ·Reconstructing the flow diagram at the end. Track counts at every stage, starting from the first database export.