Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Evidence

A natural experiment borrows its randomness from circumstance

When nobody can assign people to conditions, the next best thing is a situation where something arbitrary did the assigning, and the whole argument rests on how arbitrary it really was.

By Gautam Pillai4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The problem being solved

Randomised assignment is valuable because it balances the factors nobody measured, which is the only defence against confounders you never thought of. For a large share of the questions people actually care about, randomising is impossible, unethical or absurd — you can’t assign people to a childhood, a policy regime, a recession or a hometown.

A natural experiment is an attempt to find a situation where something outside anybody’s control performed an assignment that happens to be unrelated to the outcome. If circumstance divided people into groups for a reason that has nothing to do with what you’re measuring, then comparing those groups approximates the trial you could not run. The entire credibility of the method sits on that conditional.

The recognisable forms

One family exploits a sharp cutoff. Where eligibility for something depends on a threshold — a date of birth, a score, a boundary line — the people just above and just below are similar in almost every respect except the treatment, so comparing them isolates its effect near that point. The comparison is credible precisely because being a fraction either side of an arbitrary line isn’t something anybody chose.

A second family exploits a change that arrived in one place and not another. Comparing how outcomes moved in the affected area against how they moved in a comparable unaffected one removes anything that was trending everywhere, which is the main threat to a simple before-and-after study. A third family uses genuine lotteries, which occasionally allocate places, permits or obligations at random for administrative reasons and produce something very close to a trial.

Where the argument gets its strength

The persuasive part of any of these is never the statistics. It is the case that the assignment really was arbitrary with respect to the outcome, and that case has to be made in words, with evidence about how the mechanism worked. A cutoff can be gamed by people who know where it sits. A policy adopted in one region and not another was probably adopted for reasons — and those reasons may be exactly the things that drive the outcome.

This is why the best work of this kind spends more space defending the design than reporting the result. It shows that the groups looked alike before the change, that nothing else happened at the same moment, that people could not sort themselves across the boundary, and that the effect appears where the mechanism says it should and not elsewhere.

The characteristic failure modes

The first is a boundary that people can cross deliberately. Where the threshold is known and the benefit worth having, applicants bunch on the favourable side, and the two groups stop being comparable in the way the design assumed. The bunching is usually detectable, which is why careful papers plot the distribution around the cutoff and worry about it in print.

The second is a comparison region that was never really comparable. Two places diverge for many reasons, and the assumption doing the work is that they would have moved in parallel absent the change. That assumption isn’t directly testable, since it concerns a world that did not happen, and the usual evidence for it — that they moved in parallel beforehand — is suggestive rather than conclusive.

The third is narrowness. A cutoff design tells you about the effect near the threshold, on the people at the margin, and may say very little about anybody far from it. That is a real limit rather than a quibble, because the people at the margin are often the least typical.

Why these results can still be worth a great deal

A well-constructed natural experiment often addresses a question no trial could reach, at a scale no trial could afford, in a real setting rather than a controlled one. Several of the most useful causal claims in economics and public health rest on designs of this kind, and their external validity can be better than an equivalent trial precisely because nothing was staged.

The trade is that the assumptions are stronger and less visible. A randomised trial wears its main assumption on the surface, where anybody can inspect it. A natural experiment hides its central assumption inside a story about how the world happened to be arranged, and a reader has to evaluate that story rather than a procedure.

Reading one without specialist training

Ask what did the assigning, and whether the people involved could have influenced it. Ask what else changed at the same time or in the same place. Ask who the comparison group is and why they are a fair stand-in. Ask which population the estimate applies to, which for threshold designs is usually narrower than the summary suggests.

Those four questions get a non-specialist most of the way, because the vulnerabilities of these designs are conceptual rather than technical. When a paper answers all four explicitly and without prompting, that is itself a signal about how seriously the design was taken.

Common questions

Is a natural experiment as good as a randomised trial?

Sometimes close, sometimes much weaker, and it depends entirely on how arbitrary the assignment really was. A genuine lottery approaches a trial; a policy that some places adopted and others did not may be badly confounded by whatever made them adopt it.

What does it mean that the groups moved in parallel beforehand?

It is the main evidence offered for the assumption that they would have continued to move together without the change. It supports the assumption without proving it, since two series can track each other for years and then diverge for reasons unrelated to the intervention.

Why do threshold studies only apply near the threshold?

Because that is where the two groups are genuinely comparable. Someone far above the cutoff differs from someone far below in many ways besides the treatment, so extending the estimate to them requires an argument the design itself does not supply.

Evidencenatural experimentscausationstudy designinference
Gautam Pillai
Reporter, Think Twice Today

Gautam writes about biases, choices, risk, mostly the parts other people skip and reads the small print so you do not have to.