Risk of Bias Assessment: RoB 2 and ROBINS-I Explained
Learn how to assess risk of bias with RoB 2 and ROBINS-I: domains, signaling questions, per-result judgments, and how to use ratings in synthesis and GRADE.
What Risk of Bias Means — and What It Doesn't
Every systematic review reaches the point where searching and screening are done and a harder question takes over: how much should we trust each included study? Risk of bias assessment is the structured answer. Bias, in this context, means systematic deviation from the truth — a flaw in design, conduct, or analysis that pushes a study's results away from the true effect, in a direction and by an amount you usually cannot know. It is worth separating this from two neighboring concepts that often get lumped together as "study quality."
- ·Bias is systematic error. An unconcealed allocation sequence that lets clinicians steer sicker patients away from the experimental arm will distort the estimate no matter how large the trial grows.
- ·Imprecision is random error. A small, impeccably conducted trial gives a noisy estimate with wide confidence intervals — but it is not biased. Imprecision is handled in certainty-of-evidence ratings, not in risk of bias tools.
- ·Reporting quality is how completely a study is described. A badly reported trial may have been well conducted, and vice versa. Reporting is the territory of the guidelines cataloged by the EQUATOR Network — though poor reporting often forces "no information" answers during bias assessment.
This distinction is why modern Cochrane tools abandoned numeric quality scales that summed points across unrelated items. RoB 2 and ROBINS-I instead ask targeted questions about specific mechanisms of bias, one domain at a time. Both tools are covered in depth in the Cochrane Handbook, the authoritative reference for applying them.
Why Risk of Bias Is Not a Checkbox
It is tempting to treat risk of bias assessment as a compliance exercise: fill in the table, color the cells, move on. That wastes the most decision-relevant information a review produces. The assessments should actively shape three things downstream: how strongly you phrase your conclusions, which sensitivity analyses you run, and how far you downgrade the certainty of the evidence. A meta-analysis dominated by high-risk results and one dominated by low-risk results can share the same pooled estimate and still warrant completely different conclusions.
RoB 2: The Tool for Randomized Trials
RoB 2, the revised Cochrane risk-of-bias tool, is the current standard for randomized controlled trials. It organizes the assessment into five domains that follow the life of a trial from allocation through to the reporting of the result. (RoB 2 covers bias within a study; non-publication of whole studies is an across-studies problem, handled in frameworks like GRADE.)
| Domain | Core question |
|---|---|
| 1. Randomization process | Was the allocation sequence truly random and concealed until assignment? Do baseline imbalances suggest a problem? |
| 2. Deviations from intended interventions | Did departures from the assigned interventions — or an inappropriate analysis — distort the effect? |
| 3. Missing outcome data | Were outcome data available for all or nearly all randomized participants, and is the missingness likely related to the true outcome? |
| 4. Measurement of the outcome | Was the measurement method appropriate, and could assessors' knowledge of the intervention received have influenced the measurement? |
| 5. Selection of the reported result | Was the reported result chosen from multiple eligible measurements or analyses on the basis of its findings? |
Effect of assignment vs. effect of adherence
Before touching domain 2, decide which question your review is actually asking. The effect of assignment is the effect of being allocated to an intervention, whether or not it was fully received — the question an intention-to-treat analysis answers, and usually the relevant one for clinical and policy decisions. The effect of adhering is the effect of receiving the intervention as intended, which matters for questions about what a treatment can do under ideal use. The distinction changes what counts as a problem in domain 2: for the effect of assignment, non-adherence that would also occur in routine practice is part of the answer rather than a bias — but deviations that arose because of the trial context (say, unblinded clinicians providing differential co-interventions), and any analysis that fails to keep participants in their randomized groups, are still domain 2 problems. For the effect of adherence, non-adherence itself is a genuine threat, and one that is hard to correct because adherence after randomization is not itself randomized. Most reviews assess the effect of assignment — state your choice explicitly in the protocol.
How RoB 2 Judgments Are Made
The most misunderstood feature of RoB 2 is its unit of assessment. You do not rate a study; you rate a result — one outcome, measured at one time point, analyzed in one way. The same trial can be low risk for all-cause mortality (objective, fully ascertained) and high risk for a self-reported pain score (subjective, measured without blinding). In practice, this means assessing each result that feeds your syntheses, which often produces several RoB 2 assessments per trial.
Within each domain, assessors answer a short series of signaling questions — factual prompts with the response options yes, probably yes, probably no, no, and no information. An algorithm maps the answers onto a domain-level judgment of low risk, some concerns, or high risk; assessors may override the algorithm, but should document why. The domain judgments then roll up to an overall judgment for the result:
- ·Low risk — the result is at low risk of bias in all five domains.
- ·Some concerns — at least one domain raises some concerns, but none is high risk.
- ·High risk — at least one domain is high risk, or several domains have some concerns in a way that substantially lowers confidence in the result.
ROBINS-I: Non-Randomized Studies and the Target Trial
ROBINS-I (Risk Of Bias In Non-randomized Studies of Interventions) extends the same domain-based logic to cohort studies, controlled before-after studies, and other non-randomized designs. Its organizing idea is the target trial: imagine the hypothetical pragmatic randomized trial you would have run to answer the same question — same participants, interventions, and outcomes, and free of features that would themselves introduce bias — regardless of whether such a trial would be feasible or ethical. Each domain then asks how far the observational study's results might deviate from what that target trial would have found.
ROBINS-I has seven domains, grouped by when the bias arises:
| Stage | Domains |
|---|---|
| Pre-intervention | Confounding; selection of participants into the study |
| At intervention | Classification of interventions |
| Post-intervention | Deviations from intended interventions; missing data; measurement of outcomes; selection of the reported result |
Confounding is almost always the dominant concern. Without randomization, people who receive an intervention differ systematically from those who do not — in disease severity, comorbidity, health-seeking behavior — and those differences are often related to prognosis. ROBINS-I therefore expects you to list the important confounders for your question in advance, at the protocol stage, and to judge whether each study's design and analysis dealt with them adequately.
As with RoB 2, the unit of assessment is a result, not a study — and because the target trial is pragmatic, lack of participant or clinician blinding is not in itself a ROBINS-I concern. Judgments in the original version of the tool use a wider scale than RoB 2: low, moderate, serious, or critical risk of bias, plus no information. The anchoring matters. Low risk means the study is comparable to a well-conducted randomized trial for that domain — rare in practice. Moderate means sound by the standards of non-randomized studies, but short of that benchmark. The overall judgment is at least as severe as the worst domain, and results judged critical are generally considered too problematic to include in synthesis at all.
Run the Assessment as a Dual, Documented Process
The convention for both tools is that two reviewers assess each result independently and resolve disagreements by discussion, escalating to a third reviewer where needed — the same rationale that underlies dual-reviewer abstract screening. Independent judgments surface ambiguities that a single reader normalizes away. Before the full assessment, run a calibration round on a handful of studies to align interpretations of the signaling questions; this is where most systematic disagreements between assessors get resolved cheaply.
Document as you go. For every answer, record the supporting evidence: a quotation and its location, or an explicit note that the information was absent. Whether you track this in a spreadsheet or in a platform such as AutoEvidence, the essentials are the same — judgments at the level of the result, an audit trail for each answer, and disagreements resolved visibly rather than silently overwritten.
Putting the Judgments to Work
A finished set of assessments should show up in at least three places in your review.
- 1.Weighting the narrative. Conclusions should track the trustworthy evidence, not the majority vote. If the results favoring an intervention are concentrated among high-risk results while the low-risk results are equivocal, the synthesis should say so plainly.
- 2.Sensitivity analyses. Rerun your meta-analysis restricted to low-risk results, or excluding high-risk ones. If the pooled effect shrinks or disappears, that is a substantive finding about the evidence base, not a statistical footnote. Our guide to reading a forest plot covers how these comparisons are presented.
- 3.GRADE. Risk of bias is one of the reasons to rate down the certainty of evidence for an outcome under the GRADE approach. Domain-level judgments give a certainty rating the specific, citable justification it needs.
Finally, report the assessments transparently: PRISMA 2020 expects you to describe the tool, the process, and the results, and the familiar "traffic light" plots are a compact way to show per-domain judgments for every result. Handled this way, risk of bias assessment stops being a table in the appendix and becomes what it was designed to be — the mechanism by which a systematic review calibrates its confidence to the evidence.