Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Evidence

Why a shelf of famous psychology findings quietly stopped being true

When large teams began repeating well-known experiments with bigger samples and fixed analysis plans, a substantial share of the results did not come back — and understanding why is more useful than any of the original claims.

By Gautam Pillai4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

What happened, in order

For most of the twentieth century, a psychological finding entered the textbooks on the strength of one or two published studies, usually with modest numbers of participants, and was rarely repeated. Repeating someone else’s experiment was not a route to publication, so almost nobody did it. The literature accumulated results without ever testing whether they were stable.

From around 2011 that changed, for several reasons at once, including a prominent fraud case, a paper demonstrating how ordinary analytic choices could produce evidence for an obviously impossible effect, and a growing unease among researchers who had quietly failed to reproduce well-known results in their own laboratories. Large coordinated projects began running famous experiments again, with sample sizes far exceeding the originals, protocols agreed with the original authors where possible, and analysis plans registered before the data existed.

A substantial proportion of the effects did not appear, or appeared at a fraction of their reported size. This was not confined to psychology, and similar exercises in other fields have produced similarly uncomfortable results.

Some specific casualties

Ego depletion — the idea that self-control draws on a limited resource that is used up by exertion — was one of the most influential ideas in the field, with hundreds of supporting studies. Large multi-laboratory replications did not find the effect, and the theory is now regarded as, at best, unproven. It is a striking case because the supporting literature was so large, which tells you that a large literature is not the safeguard it appears to be.

Social priming produced a series of famous demonstrations in which exposure to words associated with a concept supposedly changed unrelated behaviour, such as walking speed. Direct replications largely failed. Power posing was reported to change both feelings of power and hormone levels; the hormonal claims did not replicate, one of the original authors publicly withdrew confidence in the effect, and what remains under discussion is a much narrower claim about self-reported feelings.

The marshmallow studies are a subtler case, and the honest version is more interesting than either the original story or the debunking. The original work showed that some children could delay a small reward for a larger one later. The famous part — that this predicted life outcomes decades on — came from small follow-up samples. Larger later work with more extensive statistical controls found the association considerably weaker once family circumstances were accounted for. The behaviour is real; the strong causal reading of it was never well supported.

Why so much of it was fragile

Four mechanisms account for most of it, and none requires anyone to have behaved dishonestly. Small samples produce noisy estimates, so a study that finds anything at all with few participants has probably overestimated the effect. Publication favoured positive, surprising results, so the studies that found nothing stayed in the drawer and the published record was a biased sample of the work done.

Analytic flexibility did the rest. A researcher analysing data faces many reasonable choices — which participants to exclude, which measure to use, which covariates to include, when to stop collecting — and making those choices while looking at the results can manufacture apparent findings from noise without any intent to deceive. And a hypothesis formed after seeing the data, then reported as though it had been predicted, converts an exploration into a confirmation on paper only.

What survived

Plenty did, and the field is not a smoking ruin. Effects that involve straightforward perceptual or cognitive processes, with large signals and simple designs, generally reproduced. Anchoring came through the replication projects in reasonable shape. So did much of the work on how the framing of a numerical problem changes the answers people give, and on the difference between recognition and recall in memory.

A pattern is visible in what survived: effects with large signals relative to their noise, in designs that do not depend on subtle manipulations of context, tend to hold. Effects requiring an elaborate causal chain from a small nudge to a distant behaviour tend not to. That is worth carrying as a rough prior when you meet a new claim.

How to read this field now

Ask how many independent groups have found the effect, not how many papers exist, since a single laboratory can produce a great many papers. Ask whether the replications were direct or conceptual, since a conceptual replication that changes the method can succeed for reasons unrelated to the original. Ask whether the analysis was registered in advance. And treat any result you first encountered in a popular book published before roughly 2015 as requiring a check, because that is the literature the corrections landed on.

None of this is a reason for cynicism about research. The crisis was identified by researchers, using research, and the reforms that followed — preregistration, registered reports, data sharing, larger samples — are a field correcting itself in public. That is what the process is supposed to look like. It is just slower and less flattering than the version in which findings arrive true.

Common questions

Does a failed replication prove the original was wrong?

Not on its own. A single failure could reflect a flawed replication, a genuine dependence on conditions, or bad luck. What moves the conclusion is a pattern of well-powered failures, especially when the original authors were consulted on the protocol.

Is this problem unique to psychology?

No. Comparable difficulties have been documented in preclinical biomedical research, in economics, and elsewhere. Psychology gets discussed most because it examined itself most publicly, which is arguably to its credit rather than otherwise.

What should I do with a striking finding I read about?

Check its date and its provenance. Look for whether independent teams have reproduced it and whether the result predates the reforms. A finding that has been repeated by people with no stake in it is a different kind of thing from a finding that has been repeated in books.

Evidencereplicationpublication biaspsychologyevidence
Gautam Pillai
Reporter, Think Twice Today

Gautam writes about biases, choices, risk, mostly the parts other people skip and reads the small print so you do not have to.