Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Evidence

Regression to the mean makes useless interventions look like they worked

Extreme measurements are followed by less extreme ones for purely statistical reasons, and any programme aimed at the worst cases inherits that improvement for free.

By Rohan D’Souza4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A statistical fact, not a force

Any measurement that contains a component of chance will, on average, be closer to the average the next time it is taken, whenever the first measurement was extreme. This is regression to the mean, and it is not a tendency of the world towards mediocrity or a pull of any kind. It is a consequence of how extreme values arise.

An unusually high measurement typically reflects two things at once: the underlying value being genuinely high, and the random component happening to be favourable that day. The underlying value persists into the next measurement. The lucky component does not, because it was random. So the second measurement contains the true part and a fresh, ordinary error, and is therefore usually less extreme. The stronger the role of chance in the measure, the larger the effect.

Why this fools evaluations so reliably

Interventions are almost always aimed at cases that are currently doing badly, which is sensible policy and a statistical trap. The worst-performing schools, the sickest patients, the machines that failed most often, the sales region that had a terrible quarter — these are selected precisely because their most recent measurement was extreme.

That selection guarantees improvement at the next measurement whether or not the intervention does anything at all. A programme applied to the bottom group and then evaluated by whether the bottom group improved will report success on a completely inert treatment, and everyone involved will have acted in good faith. This is one of the most common ways that ineffective interventions acquire evidence in their favour.

The name is a historical accident

The term arrived in the nineteenth century out of work on inherited height, where it was observed that unusually tall parents tended to have children closer to average. The original phrasing spoke of a reversion towards mediocrity, which made it sound like a biological law pushing populations back towards the middle.

That reading is wrong, and the wrong reading has never quite gone away. Nothing pushes. The pattern appears in any measurement with a random component, including ones with no biology in them at all, which is why the same shape turns up in test scores, quarterly figures, injury rates and rainfall. Knowing that it is a property of measurement rather than of the thing measured is most of what you need to avoid being fooled by it.

The mirror image in feedback

The same mechanism distorts beliefs about praise and criticism. If someone is criticised after an unusually poor performance, the next performance is likely to be better; if praised after an exceptional one, the next is likely to be worse. Observed repeatedly, this produces a strong impression that criticism works and praise backfires.

The impression is generated by the statistics regardless of what praise and criticism actually do. This is a widely cited illustration of the effect, and the honest way to state it is as a demonstration that observational impressions cannot distinguish the two explanations, not as a finding about the effects of feedback. What praise and criticism really do is a separate question, studied separately, with results that are more complicated than either popular position.

How to tell the difference

A control group solves it completely. If cases are selected for being extreme and then split at random, both groups regress equally, and any difference between them is attributable to the treatment. This is one of the clearest examples of why comparison groups exist, because no amount of careful measurement of the treated group alone can separate the two explanations.

Where a control is impossible, there are partial defences. Selecting cases on one measurement and evaluating on a different one avoids the worst of it. Taking several measurements before the intervention gives a more stable baseline than a single extreme reading. Looking at whether the improvement exceeds what would be expected from regression alone requires knowing how much of the measure is noise, which is rarely known but is worth asking about.

Where it shows up outside research

Sports commentary is full of it: the second season that disappoints after a spectacular first, the player who slumps after appearing on a magazine cover. Some of that may be real, and pressure and complacency are plausible mechanisms. But an exceptional first season is exactly the kind of extreme measurement that regression predicts will not repeat, so the statistical explanation has to be ruled out before the psychological one is needed.

Business commentary has the same problem. Companies identified as outstanding by any performance screen tend to look less outstanding afterwards, and books explaining what made them great are written at the peak. The subsequent decline is then attributed to losing the discipline the book described. A simpler explanation is available, was there all along, and requires no story about culture.

Common questions

Does regression to the mean mean everything becomes average?

No. The underlying differences between cases persist; only the chance component fails to repeat. Genuinely superior performers remain superior on average, they just tend not to repeat their most extreme results.

How large is the effect?

It depends entirely on how much of the measurement is noise. A measure that is mostly stable and precisely taken regresses very little; a noisy measure of a short period regresses a great deal. That is why the reliability of the measure is worth knowing before interpreting any change in it.

Is this the same as the gambler’s fallacy?

No, and they are almost opposites. The gambler’s fallacy expects independent events to compensate for each other, which they do not. Regression to the mean is about extreme measurements of a stable underlying quantity, where the extremity partly reflects error that will not recur.

Evidenceregression to the meanstatisticsevaluationcausation
Rohan D’Souza
Features writer, Think Twice Today

Rohan writes the explanatory pieces on biases, choices, risk and would rather show the working than assert the conclusion.