How to Write a Systematic Review Protocol (PRISMA-P Guide)
A PRISMA-P guide to writing a systematic review protocol: precise eligibility criteria, draft search strategy, synthesis plans, PROSPERO timing, and pitfalls.
Why Write a Protocol Before You Search
A systematic review protocol is a complete description of your methods, finalized before you have seen a single result. That timing is the entire point. Once you know which studies exist and roughly what they show, every methodological choice you make afterward — which outcomes to emphasize, which designs to admit, how to pool — is potentially shaped by that knowledge, consciously or not.
The protocol guards against two specific failure modes. The first is outcome switching: a review that quietly promotes a secondary outcome to primary position because the original primary outcome showed nothing interesting. The second is criteria drift: eligibility criteria that bend mid-screening to admit an appealing study or exclude an inconvenient one. A dated, publicly registered protocol makes both visible, because readers can compare what you planned against what you eventually report.
There is also a mundane, practical reason to invest in the protocol: it is the team's operating manual. Screeners consult its eligibility criteria hundreds of times, extractors work from its data items, and the statistician executes its synthesis plan. A vague protocol does not avoid those decisions — it just defers them to moments when they must be made ad hoc, under time pressure, with results in view.
What PRISMA-P Asks You to Specify
PRISMA-P — Preferred Reporting Items for Systematic review and Meta-Analysis Protocols — is the reporting guideline for protocols, published in 2015 as a 17-item checklist available from the PRISMA website. It is the protocol-stage counterpart of PRISMA 2020, which governs the final report. The items fall into three groups: administrative information, introduction, and methods. In practice, these are the decisions PRISMA-P forces you to make in writing:
| PRISMA-P area | What you must pin down |
|---|---|
| Administrative information | A title identifying the document as a systematic review protocol, registration details, authors and their contributions, funding sources and the sponsor's role, and how amendments will be documented |
| Rationale and objectives | Why the review is needed given what is already known, and an explicit question framed with PICO: population, intervention, comparator, outcomes |
| Eligibility criteria | Study designs, populations, interventions, comparators, outcomes, settings, and any language or date limits — with decision rules for edge cases |
| Information sources and search | Every database, trial register, and other source you will search, plus a full draft search strategy for at least one database |
| Study records | How records will be managed and deduplicated, how many reviewers screen at each stage, and how disagreements are resolved |
| Data items and outcomes | Every variable to be extracted, plus primary and secondary outcomes with explicit prioritization |
| Risk of bias | Which appraisal tool applies to which study design, whether assessment is at study or outcome level, and how it feeds into the synthesis |
| Synthesis | Effect measures, meta-analysis model with rationale, heterogeneity handling, and planned sensitivity and subgroup analyses |
| Meta-bias and certainty | How publication and selective-reporting bias will be assessed, and how certainty of evidence will be rated |
A few of these deserve emphasis. The search item is where most first drafts fall short: naming databases is not a search strategy. PRISMA-P expects at least one database's search written out in full — every term, field tag, and limit — ready to run; our guide to building a systematic review search strategy covers how to construct one. The risk-of-bias item requires tools matched to design — RoB 2 for randomized trials, ROBINS-I for non-randomized studies, as explained in our risk of bias guide. And the final item asks how you will rate confidence in the cumulative evidence; for intervention reviews that almost always means GRADE, and the protocol should say so.
For study records, name the actual tooling: a reference manager plus spreadsheets, or a dedicated screening platform such as AutoEvidence. What matters is not the software but that every screening and extraction decision is logged and auditable.
Register on PROSPERO — Early, and for Free
PROSPERO is the international prospective register of systematic reviews, maintained by the Centre for Reviews and Dissemination at the University of York. Registration is free, and the record it creates is public and dated — which is precisely what gives your protocol its evidentiary force.
Timing matters more than most first-timers realize. PROSPERO is a prospective register: it is designed for protocols submitted before the review's results could be known, and registrations are not accepted once data extraction is underway. The workable sequence is: draft the protocol, pilot your criteria on a small sample of records, submit the registration, then begin screening in earnest. This is the same timing logic that PRISMA 2020's registration-and-protocol item asks you to report at the end — see our PRISMA 2020 walkthrough for the reporting side of that item.
PROSPERO's scope centers on reviews with health-related outcomes. If your review falls outside it, the principle survives the register: deposit the protocol somewhere public and timestamped — an institutional repository or a journal that publishes protocols — so that a dated version demonstrably predates your results.
How Specific Is Specific Enough?
Apply one test to every sentence of your eligibility criteria: could an independent reviewer, given only the protocol, make the same include/exclude decision you would? If a criterion requires judgment you have not written down, it is not yet specific enough. Some hypothetical worked examples of the difference (illustrative criteria, not from a real protocol):
| Vague (as first drafted) | Precise (as it should read) |
|---|---|
| Adults with type 2 diabetes | Adults aged 18 or over with type 2 diabetes as defined by study authors; mixed-population studies eligible only if results for the type 2 subgroup are reported separately |
| Exercise interventions | Structured aerobic or resistance training programs of at least eight weeks' duration; studies where exercise is one component of a multicomponent intervention are excluded unless its effect is separable |
| Recent, high-quality studies | Randomized controlled trials published from the year the comparator became standard of care onward; methodological quality is assessed with RoB 2, not used as an eligibility filter |
| Relevant clinical outcomes | Primary: the named outcome, measured at a stated time point with a stated instrument or definition. Secondary: a named, closed list. Trials that did not measure any of these outcomes are excluded at full text; trials that measured but did not report them are retained and the authors contacted |
Notice the pattern: each precise version anticipates an edge case — mixed populations, multicomponent interventions, the temptation to use "quality" as a gate — and writes the decision rule in advance. Note also that methodological quality belongs in risk-of-bias assessment, not in eligibility: excluding "low-quality" studies by feel is criteria drift wearing a lab coat.
Plan the Synthesis Before You See a Single Result
The synthesis section is where protocols most often go thin. "Data will be pooled where appropriate" defers every consequential decision to the moment you are looking at results — exactly the moment your judgment is least trustworthy. A defensible synthesis plan specifies, in advance:
- ·Effect measure. Risk ratio, odds ratio, or hazard ratio for dichotomous and time-to-event outcomes; mean difference or standardized mean difference for continuous ones — with a reason, such as reserving SMD for outcomes measured on different instruments.
- ·Model and rationale. Fixed-effect versus random-effects, justified by the expected clinical and methodological diversity of eligible studies — not chosen after inspecting an observed heterogeneity statistic.
- ·Heterogeneity handling. How heterogeneity will be assessed, and what happens if pooling is not sensible. A structured synthesis without meta-analysis is a plan; "we will discuss the studies" is not.
- ·Sensitivity analyses. Named in advance — for example, excluding studies at high risk of bias, or re-running the analysis under the alternative model.
- ·Subgroup analyses. Few in number, each with a stated clinical or biological justification, and where possible an expected direction of effect.
PRISMA-P also asks for a meta-bias plan: how you will look for publication bias and selective outcome reporting. Concretely, that usually means comparing published outcomes against prospective registry entries on ClinicalTrials.gov, and using funnel-plot-based methods only when enough studies are available for them to be interpretable — the Cochrane Handbook gives detailed guidance on both.
Treat the number of pre-specified subgroup analyses as a credibility signal. A protocol listing a dozen subgroups without justification reads like a plan to find something, somewhere. Three subgroups, each with a mechanism attached, reads like science.
Amendments: Allowed, Dated, Disclosed
A protocol is a commitment device, not a prison. Legitimate reasons to amend arise constantly: a database becomes unavailable, an appraisal tool is updated, screening exposes a genuinely ambiguous criterion nobody foresaw. The discipline is in how you amend: date the change, state the rationale, record it in your PROSPERO entry — which preserves an audit trail of revisions — and disclose it in the final report.
The test readers will apply is simple: was the change made blind to results? Clarifying an eligibility criterion during screening is usually defensible. Changing the primary outcome after extraction is complete demands extraordinary justification. When an amendment could plausibly look results-driven, run the original analysis too and present both.
How Long It Takes, and Mistakes First-Timers Make
Expect a first protocol to take several weeks of part-time work, with the search strategy and eligibility criteria consuming most of it. That time is frontloaded, not wasted: every decision rule written now is a debate the team does not have during screening, and every hour spent piloting criteria saves multiples downstream. The mistakes below account for most of the protocol trouble first-time reviewers run into:
- ·Registering too late. Once data extraction is underway, prospective registration is off the table. Best practice is to register before screening begins.
- ·Databases without a search string. Listing PubMed and Embase is not a search strategy; PRISMA-P wants at least one full draft search.
- ·Criteria without edge-case rules. Mixed populations, multicomponent interventions, and partial outcome reporting will all occur. Decide now.
- ·No outcome prioritization. If every outcome is primary, none is — and readers cannot tell whether the reported emphasis was planned.
- ·Model choice after the fact. Picking fixed- versus random-effects once the data are visible undermines the pooled estimate's credibility.
- ·Open-ended subgroup fishing. "We will explore sources of heterogeneity" pre-specifies nothing.
- ·Skipping the pilot. Criteria and extraction forms that have never touched real records will fail on contact with them.
- ·Treating the protocol as paperwork. It is the review's operating manual; written as a formality, it protects nothing.
A strong protocol converts the rest of the review into execution: the questions are settled, the rules are written, and what remains is the work itself — searching, screening, extracting, and synthesizing exactly as planned, with any departures dated and explained.