Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Evidence

Pooling studies fixes the sample size and not the bias

Combining many studies produces a precise estimate whether or not the underlying literature was any good, and the precision is what makes a compromised pooled result so persuasive.

By Rohan D’Souza3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

What pooling is for

Individual studies are usually too small to resolve the question they ask. Combining them, weighting each by how precisely it estimated the effect, produces a summary with a much narrower range than any single contributor, and it can reveal a consistent signal that was invisible in any one report.

That is a genuine advance and it’s why systematic review sits near the top of most descriptions of the evidence hierarchy. It also creates a specific hazard, because the summary inherits every property of the studies it contains, and inherits them with an authority that none of them individually possessed.

The missing studies problem

If work that found nothing was less likely to be published, then the visible literature is a biased sample of the work actually done, and pooling the visible literature estimates the bias very precisely. No statistical procedure applied afterwards recovers what was never written down.

There are detection methods. The most familiar looks at whether smaller studies report larger effects, which is what you would expect if small studies needed a dramatic result to be publishable while large ones got published regardless. Asymmetry of that kind is a warning sign rather than a proof, since several other things produce it, and adjustments that attempt to correct for it rest on assumptions that can’t be checked. The reliable fix is upstream: registries that record studies when they begin, so the ones that vanish can at least be counted.

Quality does not average out

A common misconception is that pooling many weak studies produces a strong result. It does not. If the individual studies share a flaw — an unblinded outcome measure, a comparison against nothing, a population selected the same way everywhere — then every contributor is wrong in the same direction and the average is wrong with greater confidence.

Errors cancel only when they’re independent, and methodological conventions are the opposite of independent. A field that has taught one flawed design for twenty years will produce a hundred studies sharing its weakness, and their agreement is evidence about the convention rather than about the world.

Reviewers know this, which is why quality assessment is part of the standard procedure, and the assessment itself involves judgement calls that different teams make differently. Two competent reviews of the same question can reach different conclusions by weighting the same studies differently or by drawing the inclusion boundary in a different place. When that happens, the disagreement is worth reading closely, because it usually identifies exactly which studies the answer depends on.

Combining things that are not the same thing

The other central difficulty is whether the studies are estimating a common quantity at all. Pooling requires that the interventions, populations and outcome measures be similar enough that one number describes them, and researchers disagree honestly about where that line falls.

Where the studies genuinely differ, the summary estimate describes an average across conditions that may not exist anywhere. Statistical measures of variation between studies are reported for this reason, and high variation is a signal to explain the differences rather than to average across them. A pooled figure with wide disagreement underneath is a description of a disagreement, and it should be read as one.

How to read a review without checking every study

Look for whether the search strategy is described and whether unpublished work was sought, since a review of only the easily found literature has selected on visibility. Look for whether study quality was assessed and whether the result changes when the weakest studies are excluded — that comparison is far more informative than the headline figure.

Look at the variation between studies, and at whether the authors explain it or bury it. And check whether the review itself was registered in advance, because the same analytic flexibility that affects primary research affects reviews: which studies to include, which outcome to treat as primary, which subgroup to report.

What a good pooled result is worth

A great deal, when the underlying literature contains preregistered studies of reasonable size, when unpublished work has been chased, when the studies agree with each other, and when excluding the weakest ones does not move the answer. Under those conditions the narrow interval means what it appears to mean, and the conclusion is about as solid as observational or experimental evidence gets.

When those conditions don’t hold, the pooled figure is a precise summary of a compromised record, and its precision actively misleads. That is the specific danger worth remembering: the failure mode of meta-analysis isn’t vagueness but false confidence, and false confidence is much harder to argue with.

Common questions

Does a meta-analysis outrank a single large trial?

Not automatically. One large, well-conducted, preregistered trial can be more trustworthy than a pooled summary of many small studies of uncertain quality, because it has one clear design rather than an inheritance of many unclear ones. The comparison depends on the ingredients.

Can publication bias be corrected statistically?

Only partially and under assumptions that cannot be verified. The available methods estimate what missing studies might have shown, which is guesswork constrained by a model. They are worth reporting as a sensitivity check and should not be treated as having recovered the missing evidence.

What does high variation between studies mean?

That the studies are not all estimating the same quantity, usually because populations, interventions or measures differed. It is a prompt to investigate what distinguishes them, and it makes a single pooled number a poor summary of what is actually known.

Evidencemeta-analysispublication biasevidenceresearch methods
Rohan D’Souza
Features writer, Think Twice Today

Rohan writes the explanatory pieces on biases, choices, risk and would rather show the working than assert the conclusion.