Why pool trials at all?

One trial, even a good one, is a single snapshot. It might have got lucky or unlucky; it might have been a bit small. A systematic review systematically hunts down every trial that meets pre-set quality criteria, and a meta-analysis then combines their results statistically — weighting bigger, tighter trials more heavily — into one pooled estimate. The payoff is power and precision: ten trials of 100 people each can settle a question that no single one of them could.

The standard way to see all of this at once is the forest plot. It looks intimidating and is actually simple:

A forest plot of five trials and their pooled estimate Trial A · n=60 Trial B · n=120 Trial C · n=45 Trial D · n=200 Trial E · n=90 Pooled (5 trials) no effect ← favours placebo favours supplement →
How to read it in ten seconds. Each row is one trial: the square is its result (bigger square = bigger, more trustworthy trial), the whisker is its confidence interval. A whisker crossing the dashed "no effect" line (Trial C) means that trial alone wasn't conclusive. The diamond at the bottom is the pooled answer — here it sits clear of the line, so combined, the trials point to a real effect.

The catch: garbage in, garbage out

A meta-analysis inherits the quality of the trials it eats. Pool ten tiny, biased, seller-funded trials and you get a very precise estimate of a biased effect — the tight diamond makes it look authoritative, which is exactly the danger. Three things to check before you trust one:

  • Heterogeneity. If the trials wildly disagree (whiskers all over the place), averaging them is like averaging apples and motorcycles. Reviews report this as "I²" — high I² means treat the single pooled number with caution.
  • Publication bias. Trials that find nothing often never get published, so the pool can be skewed toward flukes that happened to look positive. Good reviews test for this (a "funnel plot").
  • Inclusion criteria. A rigorous review only admits decent trials; a weak one sweeps in everything. Look for the phrase GRADE — a formal rating of how much confidence the evidence deserves ("high" down to "very low").
The green-flag phrase

"Systematic review and meta-analysis of randomised controlled trials, GRADE-assessed." That single line tells you someone gathered the strongest study type, combined them properly, and graded their own confidence. It's about as good as evidence gets for a supplement.

Worked example A trustworthy pool vs a precise mirage

The real thing. Creatine's evidence rests on things like a 2024 GRADE-assessed systematic review of 143 RCTs — many trials, the strongest design, formally graded. When a diamond like that clears the no-effect line, it's about as close to "settled" as this field gets. (See our creatine deep-dive.)

The mirage. A "meta-analysis" of eight 20-person open-label trials, all run by ingredient sellers, with sky-high heterogeneity and no publication-bias check, can produce a confident-looking pooled number for almost anything. Same chart, same diamond shape — but built on sand. The label will cite it exactly as if it were the creatine review.

So "a meta-analysis found…" is not an automatic trump card. It's the top floor of the pyramid only when the trials underneath were solid and combined honestly. Read the diamond, but also glance at what it's standing on.

What to remember

Read the diamond — then check what it's standing on.

  • A forest plot at a glance: squares are individual trials (bigger = weightier), whiskers are uncertainty, the diamond is the pooled answer. Crossing the "no effect" line means inconclusive.
  • Garbage in, garbage out. Pooling weak trials yields a precise-looking wrong number. Watch heterogeneity, publication bias, and inclusion standards.
  • Green flag: "systematic review and meta-analysis of RCTs, GRADE-assessed." That's the strongest evidence a supplement realistically offers.
Try it · ~1 min

Weigh two “meta-analysis” claims.

  1. For a supplement you're curious about, search "[ingredient] systematic review meta-analysis".
  2. Check how many trials it pooled and whether they were RCTs. Five is not fifty.
  3. Look for a GRADE rating or a note on heterogeneity/publication bias. "Very low certainty" on the diamond means don't over-trust it.

References

1
Page MJ, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71
2
Guyatt GH, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336:924. doi:10.1136/bmj.39489.470347.AD

Progress saves automatically in this browser · skip ahead anytime · no streak.