Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Evidence

Analysis has too many degrees of freedom, which is why preregistration exists

The same dataset can be analysed in dozens of defensible ways, and choosing among them after seeing the results turns noise into findings without anyone intending it.

By Tara Mukherjee4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The choices nobody counts

Between collecting data and reporting a result sits a long series of decisions, each of which is individually reasonable. Which participants to exclude for inattention or implausible responses. Which of several related outcome measures to treat as primary. Whether to combine two conditions or analyse them separately. Which control variables to include. Whether to transform a skewed variable. When to stop collecting.

Each choice has a defensible argument behind it. The difficulty is that there are many of them, they multiply, and the resulting number of possible analyses is large. If those choices are made while looking at how they affect the result, the analysis reported is the one that worked, selected from a set nobody counted and nobody sees.

Why this is not the same as cheating

The uncomfortable part is that it does not feel like anything. A researcher tries a reasonable specification, finds a result that is ambiguous, notices that two participants clearly misunderstood the instructions, excludes them for good reasons, and the picture becomes clearer. Nothing here is fabricated. But the decision to look for those participants happened after seeing a disappointing result, and would probably not have happened after a clean one.

This has been called the garden of forking paths, and the name captures the important feature: the researcher need not have tried many analyses. It is enough that the analysis they did try was contingent on the data. The count that matters is the number of paths that could have been taken, not the number that were.

What a p-value actually says, and does not

The standard statistical test answers a narrow question: if there were no real effect, how often would data at least this extreme arise by chance? A small value means the data would be surprising in a world with no effect. That is all it means.

It does not say how likely the hypothesis is to be true, which is a different quantity requiring information the test does not use — including how plausible the hypothesis was to begin with. It does not say the effect is large or important. And crucially, the whole calculation assumes the analysis was specified without reference to the data. Once the specification was chosen after looking, the stated figure no longer describes the procedure that produced it, which is why the flexibility problem and the p-value problem are really the same problem.

What preregistration changes

Preregistration means writing down the hypothesis, the design, the sample size and the analysis plan, and depositing it in a public registry with a timestamp, before the data are collected. It does not prevent anyone from doing anything. It makes the distinction between a prediction and an observation visible afterwards.

That distinction is the entire value. Exploratory analysis is legitimate and productive — most interesting hypotheses come from somewhere, and looking at data is one of the better sources. The damage is done when exploration is presented as confirmation, because the reader then applies the wrong standard of evidence. A preregistered plan lets a paper report both honestly: here is what we predicted and found, here is what we noticed afterwards and have not yet tested.

A stronger version, the registered report, sends the plan to a journal for review before the study runs, with acceptance decided on the design rather than the outcome. This removes the incentive to produce a positive result, and studies published this way return null findings far more often than the conventional literature does. That difference is itself evidence about how much the old incentives were shaping what got published.

The limits of the fix

Preregistration is not a guarantee of quality. A badly designed study registered in advance is still badly designed, and a plan can be vague enough to permit almost anything, which defeats the purpose. Deviations from the plan are sometimes necessary and are not misconduct, provided they are declared. Some fields work with data that already exist, where the concept has to be adapted rather than applied directly.

It also imposes real costs on genuinely exploratory work, and there is a legitimate worry that a culture demanding advance specification could discourage the open-ended looking that produces new ideas. The reasonable position is that preregistration is a tool for confirmation rather than a universal requirement, and that the important reform is labelling which mode a piece of work is in.

What this means for a reader

When assessing a study, the question is not only what it found but how much freedom the analysis had. A single prespecified outcome measure is worth more than a favourable result among many. A subgroup finding that was not predicted is a hypothesis rather than a conclusion, and reading it as a conclusion is one of the most common errors in the interpretation of otherwise good research.

None of this requires statistical training. It requires asking whether the thing reported was the thing the researchers set out to test, which is a question anyone can ask and which the better papers now answer explicitly.

Common questions

Is exploratory research bad?

Not at all — it is where hypotheses come from and it deserves more respect than it sometimes gets. The problem is exclusively about labelling. Exploration presented as confirmation invites the reader to apply a standard of evidence the work cannot support.

Why do subgroup findings need special caution?

Because dividing a sample many ways creates many chances for a difference to appear by chance, and the subgroups are often chosen after the overall result disappointed. A subgroup effect predicted in advance is a genuine finding; one discovered afterwards needs a fresh sample.

Should p-values be abandoned?

Some statisticians argue so and others defend them as useful when correctly interpreted. The debate is live and unresolved. What is broadly agreed is that a single threshold treated as a verdict, without regard to effect size, design or prior plausibility, is a poor way to decide anything.

Evidencepreregistrationp-valuesanalysisresearch methods
Tara Mukherjee
Contributing editor, Think Twice Today

Tara writes the explanatory pieces on biases, choices, risk and thinks most subjects are more interesting once you know how they work.