All articles
7 min read

How to Build a Systematic Review Search Strategy

How to build a systematic review search strategy: PICO concept blocks, MeSH plus free-text terms, Boolean logic, grey literature, and PRISMA 2020 reporting.

Sensitivity vs. Precision: The Trade-Off That Shapes Everything

Every downstream step of a systematic review — screening, data extraction, risk-of-bias assessment, synthesis — operates only on the records your search retrieved. An eligible trial your strategy missed is not merely a limitation; it is invisible, and no amount of careful screening can recover it. That is why the search deserves more design effort than almost any other part of the review, and why it helps to name the single trade-off behind every decision you will make: sensitivity versus precision.

Sensitivity (also called recall) is the proportion of truly relevant studies your search retrieves. Precision is the proportion of retrieved records that are actually relevant. The two pull against each other: every synonym, truncation, or broader subject heading you add tends to raise sensitivity and lower precision.

Systematic reviews deliberately favor sensitivity. Missing an eligible study biases the review in ways you cannot detect or correct, whereas an irrelevant record merely costs screening time. The Cochrane Handbook frames searching in exactly these terms: accept a large screening burden in exchange for confidence that little was missed. It is normal for the ratio of records screened to studies included to be very large. Screening platforms such as AutoEvidence, with AI-assisted abstract screening, can compress the mechanical workload of a large result set — but the reason to favor sensitivity is methodological: a missed study is an error you cannot see or correct.

Turn Your PICO Into Concept Blocks

A search strategy is not one long string of keywords. It is a set of concept blocks combined with AND, where each block represents one concept from your review question and contains every reasonable way of expressing that concept, combined with OR. Your question — already fixed in your protocol before you run any definitive search — usually follows PICO: Population, Intervention, Comparator, Outcome.

Here is the counterintuitive part: you almost never search on all four elements. Most strategies use only two blocks — population and intervention — sometimes with a third block for study design.

Why outcome terms usually stay out of the search

Outcomes are reported too inconsistently in titles and abstracts to search on. A trial may measure your outcome of interest as a secondary endpoint and never mention it in the abstract; older abstracts may not list outcomes at all; authors describe the same measure in many different ways. An outcome block therefore silently discards eligible studies — the worst kind of error, because you never see what you lost. The same logic applies to comparators: control conditions are described too variably (placebo, usual care, standard therapy, or not at all) to be safe search terms. Handle outcomes and comparators at screening — in practice usually at full-text assessment, where the whole paper is available — not at the retrieval stage.

  • ·Always build blocks for: the population or condition, and the intervention or exposure.
  • ·Sometimes add: a validated study-design filter (for example a randomized-trial filter) when your protocol restricts designs.
  • ·Almost never search on: outcomes, comparators, or settings.

Controlled Vocabulary and Free-Text Terms: Always Pair Both

Major databases index records with a controlled vocabulary — MeSH (Medical Subject Headings) in MEDLINE, Emtree in Embase. Indexers read each article and assign standardized terms, so a controlled-vocabulary search finds papers regardless of the words the authors chose: a search on the heading for myocardial infarction also retrieves papers that only ever say heart attack. Many platforms let you explode a heading, automatically including every narrower term beneath it in the hierarchy.

Controlled vocabulary alone is not enough, for three reasons. Recently added records have not yet been indexed, so the newest studies are invisible to a heading-only search. Indexing is done by people and is imperfect. And new or niche concepts may have no adequate heading at all. Free-text searching of title and abstract words covers those gaps — but it depends entirely on the authors' wording, so it needs an explicit list of synonyms, spelling variants, and abbreviations. The rule that follows is simple: every concept block pairs controlled-vocabulary terms and free-text terms with OR.

Controlled vocabulary (MeSH/Emtree)Free-text title/abstract terms
Finds records regardless of author wordingYesNo — you must list synonyms yourself
Finds records not yet indexedNoYes
Portable across databasesNo — each database has its own thesaurusMostly — terms carry over; field tags change
Vulnerable to indexing errorsYesNo

Boolean Logic, Truncation, and Phrase Searching

Three Boolean operators do all the structural work:

  • ·OR joins synonyms within a concept block — any one of these terms is enough.
  • ·AND joins the blocks to each other — every concept must be present.
  • ·NOT excludes records, and should almost never appear in a systematic review search. A record indexed for both animals and humans, removed by NOT animals, is an eligible study lost.

Truncation and phrases

Truncation (commonly an asterisk) retrieves all endings of a word stem: randomi* covers randomized, randomised, randomization, and randomisation in one term. Choose stems deliberately — truncating too early makes the stem match hundreds of unrelated words and precision collapses.

Phrase searching (commonly quotation marks) requires words to appear together in order. Phrases are precise but brittle: a search for the exact phrase water-based exercise will miss a paper that says exercise in water. Proximity operators — these words within a set distance of each other, in any order — sit usefully between phrases and AND, keeping much of a phrase's precision with far better sensitivity. Think in proximity terms whenever a concept is a multi-word idea whose word order varies.

Search syntax is platform-specific. Truncation symbols, proximity operators, phrase handling, and field tags all differ between platforms — sometimes even for the same database offered through different vendors. Build a master strategy in one database, translate it line by line for each additional database, and save every translated version verbatim.

Worked Example: From PICO to Search String

Sample question: in adults with knee osteoarthritis (P), does aquatic exercise (I), compared with land-based exercise or usual care (C), improve pain and physical function (O)? Following the logic above, the search uses only two blocks — population and intervention. The illustrative lines below use PubMed-style syntax, with MeSH and tiab field tags. They are a worked illustration, not a validated strategy: adapt the terms to your own question, test them against known included studies, and translate them line by line for any other platform.

#ConceptPubMed-style search line
1Population — controlled vocabulary"Osteoarthritis, Knee"[Mesh]
2Population — free text"knee osteoarthritis"[tiab] OR "osteoarthritis of the knee"[tiab] OR gonarthr*[tiab] OR (knee[tiab] AND osteoarthr*[tiab])
3Population block#1 OR #2
4Intervention — controlled vocabulary"Hydrotherapy"[Mesh]
5Intervention — free text"aquatic exercise"[tiab] OR "aquatic therapy"[tiab] OR "water-based exercise"[tiab] OR "water exercise"[tiab] OR "pool exercise"[tiab] OR hydrotherap*[tiab]
6Intervention block#4 OR #5
7Combined search#3 AND #6

Notice what is absent: no pain or physical-function terms anywhere — outcome concepts are handled at screening. Notice also the deliberate redundancy inside each block: gonarthr* catches gonarthrosis and gonarthritis, and the final term on line 2 catches titles that separate the two words entirely. If the protocol restricts inclusion to randomized trials, a study-design filter is added as a third block with AND — use a published, validated trial filter for the database in question rather than improvising one, since filters are themselves search strategies whose sensitivity has to be earned.

Why One Database Is Never Enough

No single database covers the whole biomedical literature. MEDLINE and Embase overlap substantially, but each indexes journals the other does not; CENTRAL aggregates trial records from several sources; regional and subject-specific databases capture literature the large general databases under-index. Indexing lag compounds the coverage problem: a study can be findable in one database weeks before it is findable in another. Searching multiple databases is therefore not a nicety — the Cochrane Handbook treats it as a methodological requirement, and a single-database search is one of the most common reasons a review's search is judged inadequate.

Grey literature and trial registries

Studies with positive, statistically significant results are more likely to be published and easier to find than studies with null results. A search confined to published journal articles therefore inherits publication bias. Grey literature — conference abstracts, dissertations, preprints, and reports — is the partial antidote, and your protocol should state which grey-literature sources you will search.

Trial registries deserve special attention. Searching ClinicalTrials.gov and comparable registries reveals trials that were registered, and often completed, but never published — exactly the studies most likely to carry unwelcome results. Even when a registry entry yields no usable data, knowing that an unpublished trial exists lets you contact investigators for results and lets readers judge the risk of reporting bias in your synthesis.

Finally, supplement database searches with citation chasing: check the reference lists of every included study, and where your tools allow it, the newer papers that cite them. Studies found this way often expose synonyms your strategy missed.

Peer Review, Documentation, and Rerunning the Search

Have the strategy peer reviewed

Search strategies contain bugs the way code does: a missing OR that turns a synonym list into nonsense, a heading left unexploded, a truncation that quietly drops a spelling variant. Before running the definitive search, have a second information specialist or experienced searcher review the strategy line by line — formal peer review of search strategies before the definitive search is a standard step recommended by the Cochrane Handbook. A complementary validation step: assemble a small set of studies you already know must be included, and confirm the strategy retrieves every one of them. Any miss points to a missing term or a flawed block.

Document everything for PRISMA 2020

PRISMA 2020 (Page MJ et al., BMJ 2021;372:n71) requires the full search strategy for every database, register, and website searched — including all filters and limits — usually presented as a supplement. For each source, record the database and platform, the date searched, the complete line-by-line strategy, any limits applied, and the number of records retrieved. Those counts feed directly into the flow diagram; our PRISMA 2020 guide walks through both the checklist and the diagram. If the review is registered on PROSPERO, the strategy you run should match the one in the record, and any deviation should be explained in the manuscript.

Rerun before submission

The literature keeps growing while you screen, extract, and write. Rerun the full strategy in every database shortly before submission, screen the newly retrieved records, and update the flow diagram and, if needed, the synthesis. A search that is visibly out of date at submission invites a revision request, so building the rerun into your timeline is far cheaper than doing it under a revision deadline.

Sign-off checklist: every concept block pairs controlled vocabulary with free-text terms; outcomes and comparators are not in the search; multiple databases plus at least one trial registry are covered; a second searcher has reviewed the strategy; the full strategy for each database is saved verbatim with dates and record counts; and a rerun is scheduled before submission.

Run your next systematic review with AI

AutoEvidence automates search, deduplication, dual-pass screening, extraction, and meta-analysis — while you keep the scientific judgment.

Start your first review