How to Read a Forest Plot: A Practical Guide
Learn how to read a forest plot: squares, whiskers, the pooled diamond, effect measures, weights, heterogeneity, prediction intervals, and common misreadings.
What a Forest Plot Shows
A forest plot is the signature figure of a meta-analysis. One graphic answers four questions at once: what did each study find, how precise was each finding, how much does each study contribute to the synthesis, and what do the studies say when combined — and how consistently? If a systematic review includes a meta-analysis, the forest plot is where the numbers live. (If the distinction between the two is fuzzy, start with our explainer on systematic reviews versus meta-analyses.)
Forest plots look dense, but their grammar is small. Once you can parse a single study row and the summary diamond, you can read any forest plot in any journal. This guide works through the figure row by row, then walks through a complete hypothetical example from top to bottom.
The Anatomy, Row by Row
A typical forest plot has one row per study, a summary row at the bottom, and an effect-size scale along the horizontal axis. Each row combines the same five elements.
- ·Study labels (left column): usually first author and year. Adjacent columns often show the raw data — events and totals, or means and standard deviations per arm — plus the numeric estimate and confidence interval that the graphic repeats visually.
- ·The square: its position on the horizontal axis marks the study's point estimate. Its size is not decoration — the area of the square is typically proportional to the study's weight in the meta-analysis, so bigger squares mean more influence on the pooled result.
- ·The whiskers: the horizontal line through each square is that study's confidence interval, usually 95%. Long whiskers mean an imprecise study; short whiskers mean a precise one.
- ·The vertical line of no effect: drawn at 1 for ratio measures (risk ratio, odds ratio, hazard ratio) and at 0 for difference measures (mean difference, standardized mean difference). A study whose whiskers cross this line has not, on its own, excluded "no effect."
- ·The diamond: the pooled result at the bottom. Its center sits at the pooled point estimate, and its left and right tips sit at the pooled confidence limits — the diamond's width IS the pooled confidence interval. There are no separate whiskers on a standard summary diamond.
The Effect Measures You Will Meet
The meaning of the horizontal axis depends entirely on the effect measure, which is stated in the column header or the axis label. Four measures cover most forest plots you will encounter.
| Measure | Typical setting | No-effect value | Notes |
|---|---|---|---|
| Risk ratio (RR) | Binary outcomes in trials and cohort studies | 1 | Ratio of risks; the most intuitive ratio measure |
| Odds ratio (OR) | Binary outcomes; case-control studies; adjusted analyses | 1 | Approximates the RR only when the outcome is rare |
| Mean difference (MD) | Continuous outcome measured on the same scale in every study | 0 | Stays in the original units (e.g., mmHg, points on a scale) |
| Standardized mean difference (SMD) | Continuous outcome measured with different instruments across studies | 0 | Expressed in standard-deviation units |
You will also meet the hazard ratio for time-to-event outcomes; like other ratios, its null value is 1. Ratio measures are plotted on a logarithmic scale, so that a halving (RR 0.5) and a doubling (RR 2.0) sit at equal distances from 1 and confidence intervals appear symmetric. Expect axis ticks like 0.25, 0.5, 1, 2, 4 rather than evenly spaced integers.
Significance and Weight: Two Different Questions
Reading significance from a confidence interval
If a 95% confidence interval excludes the no-effect value, the result is statistically significant at the conventional 0.05 level; if the whiskers cross the line, it is not. But the interval tells you more than a p-value would: the position of the square tells you the estimated size of the effect, and the width of the whiskers tells you how precisely it was estimated. A precise estimate of a trivial effect and an imprecise estimate of a large effect can both be "significant" — the plot lets you tell them apart.
Why bigger studies dominate
Meta-analysis typically weights studies by the inverse of their variance: studies with large samples and many events produce precise estimates and therefore receive large weights. The square sizes show this at a glance, and a weight column usually states it exactly. A meta-analysis of one mega-trial and four small trials is, to a first approximation, the mega-trial's result restated — worth knowing before you attribute the pooled estimate to "five studies."
Crucially, weight reflects precision, not quality. A large trial at high risk of bias still gets a large weight. Trustworthiness has to be appraised separately — see our guide to risk-of-bias assessment with RoB 2 and ROBINS-I — and many forest plots display those judgments as colored symbols beside each row.
Heterogeneity: When the Studies Disagree
Start with the visual check: do the confidence intervals broadly overlap? If most whiskers share common ground on the axis, the studies are compatible with a similar underlying effect. If estimates are scattered on both sides of the line with intervals that barely touch, the true effects likely differ across studies — and the pooled number needs careful framing.
Below the studies, the plot usually reports I²: the percentage of the observed variability in effect estimates attributable to genuine between-study differences rather than chance. The Cochrane Handbook deliberately offers rough, overlapping bands for interpreting I² (for example, 0–40% "might not be important," 75–100% "considerable") rather than hard cutoffs, and I² is unstable when a meta-analysis contains only a few studies. You may also see tau² (τ²): the estimated variance of the true effects across studies, expressed on the analysis scale (the log scale for ratio measures). Its square root, tau, is an absolute measure of between-study spread in the same units as the effect.
Fixed-effect vs. random-effects labels
The model label near the diamond matters. A fixed-effect (common-effect) analysis assumes every study estimates the same true effect, so all variation between studies is chance. A random-effects analysis assumes the true effects themselves vary and estimates their distribution. The practical consequences are visible on the plot: random-effects models spread weight more evenly (small studies gain relative influence), the pooled point estimate can shift, and the pooled confidence interval — the diamond — is usually wider, because it carries the between-study variance as well as the within-study variance. When heterogeneity is near zero, the two models converge; when it is not, expect two visibly different diamonds if both are shown.
Prediction Intervals: What the Extra Bar Means
The pooled confidence interval answers one question: how precisely do we know the average effect across the included studies? A prediction interval answers a different one: given the between-study spread, in what range is the true effect expected to lie in a single new setting — a future study, or a particular clinic or population? On the plot it usually appears as a bar or extended line through the diamond, wider than the diamond itself.
The prediction interval is wider than the confidence interval whenever tau² is greater than zero, and it can cross the no-effect line even when the pooled confidence interval does not. That combination is a genuinely useful finding: the intervention helps on average, yet a new setting could plausibly see little or no benefit. Prediction intervals only make sense under a random-effects model and are unstable when the meta-analysis contains very few studies, which is why many plots omit them.
Common Misreadings
- ·"The CI crosses 1, so the treatment doesn't work." A wide interval crossing the null is absence of evidence, not evidence of absence. A risk ratio of 0.80 (95% CI 0.55–1.16) is compatible with a meaningful benefit and with no effect at all; the honest reading is "inconclusive," not "negative."
- ·"The diamond doesn't overlap Study X, so something is wrong." The pooled interval describes the precision of the average effect, so it is expected to be much narrower than individual study intervals — it need not overlap every study, and often won't.
- ·Comparing studies to each other. Eyeballing "Study A beat Study B" treats a chance-laden difference between two estimates as real. Read each study against the no-effect line and the pooled estimate, and let the heterogeneity statistics summarize between-study disagreement.
- ·Reading square size as importance or quality. The square encodes statistical weight only. A huge square can belong to a badly biased trial.
- ·"I² = 0%, so the studies agree." With few, imprecise studies, heterogeneity statistics have little power; an I² of zero can simply mean the analysis couldn't detect the spread that exists.
- ·Misjudging distance on a log axis. RR 0.5 and RR 2.0 are effects of the same magnitude in opposite directions — they only look asymmetric if you forget the scale.
A Worked Walk-Through (Hypothetical Example)
The following plot is entirely hypothetical, with invented study names and numbers, constructed only to practice the reading routine. Imagine five randomized trials comparing Drug A with placebo for preventing 90-day hospital readmission, pooled as risk ratios under a random-effects model.
| Study | RR (95% CI) | Weight |
|---|---|---|
| Anders 2019 | 0.60 (0.40–0.90) | 12% |
| Baptiste 2020 | 0.88 (0.71–1.09) | 26% |
| Chen 2021 | 0.79 (0.61–1.02) | 22% |
| Dahl 2022 | 1.20 (0.82–1.76) | 14% |
| Evans 2023 | 0.75 (0.60–0.94) | 25% |
| Pooled (random-effects) | 0.82 (0.69–0.97) | 100% |
Reading top to bottom: four of the five squares sit left of 1, favoring Drug A, and two studies — Anders 2019 and Evans 2023 — exclude the null on their own. Dahl 2022 points the other way, but its wide whiskers overlap the other studies' intervals — the results are visually compatible with a shared direction of effect, though not identical, and the reported I² of 44% confirms moderate heterogeneity. The largest squares belong to Baptiste 2020 and Evans 2023; together they carry about half the weight, so the pooled result is driven mostly by those two trials.
The diamond sits at RR 0.82 with tips at 0.69 and 0.97 — entirely left of the line, so the pooled result is statistically significant: roughly an 18% average relative reduction in readmission. Now suppose the plot also shows a 95% prediction interval of 0.51 to 1.33. That extra bar crosses 1, so the fair summary becomes: “Drug A reduced readmissions on average, with moderate heterogeneity, but the true effect in a single new setting could range from a substantial benefit to none.” That sentence — average effect, precision, consistency, and what to expect next time — is the whole point of the figure.
A 60-second reading routine
- 1.Identify the effect measure and the direction labels under the axis.
- 2.Locate the no-effect line: 1 for ratios, 0 for differences.
- 3.Scan the squares: which side of the line, and how scattered?
- 4.Scan the whiskers: which studies are precise, and which intervals cross the line?
- 5.Note the biggest squares — they will dominate the pooled result.
- 6.Read the diamond: where is its center, and does its width touch the line?
- 7.Check the model label (fixed vs. random) and the I²/tau² read-outs.
- 8.Look for a prediction interval and ask whether it changes the practical message.
Forest plots are generated automatically by meta-analysis software — including AI-accelerated platforms such as AutoEvidence — but no software interprets one for you. That skill transfers to every evidence synthesis you will ever read. For where the forest plot fits in the full review workflow, from question to write-up, see our step-by-step guide to conducting a systematic review.