Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Evidence

A relationship measured across groups can reverse inside them

Data about areas, teams or years supports conclusions about areas, teams and years, and the same pattern can point the opposite way once you look at the individuals inside.

By Rohan D’Souza3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two levels and two different questions

A great deal of available data describes groups rather than people: rates by region, averages by school, totals by year. Such data can answer questions about groups perfectly well. The difficulty starts when a correlation between two group-level quantities is read as a statement about the individuals who make up those groups, which is a different claim requiring different evidence.

The classic illustration is a region where a characteristic is common and an outcome is also common. That is compatible with the people who have the characteristic being the ones with the outcome, and equally compatible with those being entirely different people. The area-level number cannot distinguish the two, because it has no information about which individuals are which.

Why the reversal is not exotic

The more striking case is where the direction flips. Combine several groups, and a relationship that holds within every one of them can appear reversed in the pooled data, because the groups differ both in the thing being measured and in their composition. The pooled comparison is contaminated by which groups contribute most of the observations.

This is a matter of arithmetic rather than a paradox, and it has been documented in real datasets often enough that anyone who works with aggregated data learns to check for it. The practical signature is a comparison between two totals that were made up of different mixtures. Whenever the mixture differs, the total is comparing the mixtures as much as the quantity of interest.

Neither level is automatically the right one

The instinct after learning this is to insist on individual-level data always, and that instinct is also mistaken. Some questions are genuinely about groups: whether a policy applied to an area changed outcomes in that area is an area-level question, and answering it with individual data would require modelling the aggregation anyway.

There is a mirror error in which properties of individuals are used to infer properties of populations, ignoring how the individuals interact and combine. A behaviour that benefits one person can be neutral or harmful when everyone adopts it, and no amount of individual-level data reveals that. The correct level is the level at which the decision operates, which is a question you have to answer before choosing the data.

Spotting it in a claim

The diagnostic question is what a single row of the data represents. If each row is a country, a hospital, a school or a month, then every correlation computed is a correlation among countries, hospitals, schools or months. Any sentence in the write-up that describes people rather than units has made a leap, and the leap needs an argument.

The leap is often invisible in the prose because the language shifts quietly. A finding about districts with higher rates becomes a finding about the people in them within two paragraphs, with no acknowledgement that anything changed. Reading for the unit is one of the fastest checks available and it catches a surprising amount.

What can be done about it

Where individual data exists, the analysis can be done at that level and the group pattern predicted from it as a check. Where it does not, the honest response is to state the conclusion at the level the data supports, which usually means talking about places rather than people.

Stratification helps with the reversal case specifically. Reporting the relationship separately within each group, alongside the pooled figure, shows immediately whether the two tell the same story. Where they disagree, the within-group version is generally the one relevant to a decision about an individual, and the pooled one may still be relevant to a decision about the whole.

Why this keeps happening

Aggregated data is what is available. It is cheaper to obtain, easier to publish, less encumbered by privacy constraints, and it arrives already tabulated. Researchers use it because the alternative is often no study at all, and that is a defensible choice as long as the conclusions stay at the level of the data.

What turns an acceptable compromise into an error is the summary sentence, written for an audience, that quietly promotes a claim about regions into a claim about people. That sentence is where almost all the damage happens, and it is usually written last, by whoever is least attached to the caveats.

Common questions

Is group-level data unusable?

Not at all — it answers group-level questions properly, and for many policy questions that is exactly the right level. The error is transferring the conclusion to individuals without evidence about individuals.

How can a relationship reverse when the data are combined?

Because the groups being combined differ in composition as well as in the quantity measured, so the pooled comparison partly reflects which groups contributed most of the observations rather than the underlying relationship.

What is the quickest check?

Ask what one row of the dataset is. Every correlation is a correlation among whatever those rows are, and any claim about a different kind of thing has made a leap that should be stated and defended.

Evidencestatisticsinferenceaggregationcausation
Rohan D’Souza
Features writer, Think Twice Today

Rohan writes the explanatory pieces on biases, choices, risk and would rather show the working than assert the conclusion.