Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Evidence

Every finding comes attached to the people it was found in

A result is about the sample that produced it, and the leap from that sample to anybody else is an argument that has to be made rather than assumed.

By Tara Mukherjee3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two different questions about a study

A study can be judged on whether its conclusion is correct about the people in it, and separately on whether that conclusion travels to anyone else. The first is about internal validity — whether the comparison was fair, the measurement sound, the analysis honest. The second is external validity, and it is the question that gets skipped almost universally in the summary that reaches a reader.

The two are in tension. The design choices that make a study internally strong — tight eligibility criteria, controlled conditions, close supervision — are the same choices that make its participants less like the general population. A perfectly clean result about a narrow group is a common and perfectly respectable outcome, and it’s not a finding about everybody.

Who ends up in a sample, and why it is never random

Participants are people who were reachable, eligible, willing and able to complete the thing. Each of those filters selects. Volunteers differ from non-volunteers in ways that often relate to the outcome being studied — interest, conscientiousness, available time, health. This isn’t a flaw somebody could remove; it is what recruitment is.

There is also a well-documented skew in whole literatures. A large share of published behavioural research has been conducted on university students in a small number of wealthy countries, and comparative work has found that on several measures these populations are unusual rather than typical. That doesn’t invalidate the findings, but it does mean the phrase "people tend to" in a summary often refers to a narrow and specific group.

What has to be true for a result to transfer

Generalisation is an argument with premises. It requires that the mechanism producing the effect operates in the new setting, that the conditions the effect depends on are present, and that nothing in the new context opposes it. Stated that way, it becomes obvious that transfer should be argued rather than assumed, and equally obvious why so few summaries bother.

The most reliable transfers involve mechanisms that are basic and shared: physiology, arithmetic, physical constraint. The least reliable involve behaviour embedded in institutions, incentives and norms, since all three vary enormously and all three can reverse the sign of an effect rather than merely shrinking it.

Scale changes things that ought not to matter

An intervention that works in a small trial frequently disappoints when rolled out, and the reasons are structural rather than mysterious. Small studies are delivered by people who are unusually motivated and closely supervised, often the people who designed the thing. At scale it’s delivered by whoever is available, with less training and competing demands, to a broader population including many who would not have qualified for the trial.

There is also a supply effect: doing something for a few hundred people uses spare capacity, while doing it for a hundred thousand competes for scarce resources and changes prices, staffing and the behaviour of everyone nearby. Neither of these is captured in the original estimate, and both push in the same direction.

Reading for transferability

The questions are simple and rarely asked. Who was in it, and who was excluded — the exclusion list is often more informative than the inclusion criteria. Where and when was it done. Who delivered it, and how much attention did each participant receive. What was the comparison group actually getting.

Then compare that description with the situation you care about, honestly. Most of the time the match is partial, and the correct conclusion is a partial one: this is decent evidence that the mechanism exists, weaker evidence about the size of the effect here, and no evidence at all about a group the study excluded.

The alternative to assuming is testing locally

Where a decision is repeated and the stakes justify it, the strongest answer to a transferability question is a small local test rather than an argument about someone else’s sample. Organisations that try things at small scale before committing are doing something more valuable than reading the literature more carefully, though it works best when the local test is designed with a comparison group rather than run as an enthusiastic pilot with nothing to compare against.

Where a local test is not possible, the honest position is to act on the best available evidence while marking clearly which parts of the decision rest on a transfer that has not been checked. That marking is what makes the reasoning revisable later, and it costs a sentence.

Common questions

Does a large sample fix this?

No. Size improves precision within the population studied and does nothing about whether that population resembles yours. A very large study of an unrepresentative group produces a very precise answer to a question about that group.

Are laboratory findings useless outside the laboratory?

Not useless, but they are usually better evidence that a mechanism exists than evidence about how much it matters in ordinary conditions, where many other forces are acting at once. Field studies trade precision for that relevance.

What is the single most useful thing to check?

The exclusion criteria. They tell you who the finding is explicitly not about, and that list frequently includes exactly the people a reader has in mind.

Evidencegeneralisationsamplesexternal validitystudy design
Tara Mukherjee
Contributing editor, Think Twice Today

Tara writes the explanatory pieces on biases, choices, risk and thinks most subjects are more interesting once you know how they work.