When a famous study fails to replicate, was the effect never real, or does it depend on context?
Scientists disagree about why so many psychology results shrink on repetition: hidden context, low power, or too many false ideas to begin with.
▶ Start the storyScientists disagree, and three main explanations compete. The first is context. Jay Van Bavel and colleagues argue that failures to replicate may be explained by contextual differences between the original experiment and the replication, often called "hidden moderators". When they rated the effects in the Reproducibility Project by how sensitive to context they seemed, higher context sensitivity went with a lower chance of replicating, even after accounting for things like the original's sample size and the replication's power. The 60-lab project described in the replication-crisis lesson cuts the other way: when its findings failed, they failed consistently across contexts.
The second is heterogeneity: an effect may genuinely come in several sizes rather than one. In a replication across 36 sites, 8 of 16 effects varied significantly between sites, yet the deliberate differences between those sites explained very little of the variation. Across 200 meta-analyses of psychological effects, the typical variation was what the authors call "huge", enough that matching the original's effect size is unlikely even with very large replications.
74%
The third is that the first result was weak or wrong. In the Reproducibility Project, only 25 percent of the non-replications directly contradicted the original, while 49 percent were inconclusive because of underpowered designs. And the philosopher Alexander Bird adds a base-rate argument: if only 10 percent of tested hypotheses are true, as many as 36 percent of results will be false positives, so low replication rates could be consistent with quality science.
Quiz me
0/3
Recap
A failed replication can point to context, to real variation between studies, or to weak original evidence.
💡 A trick to remember it · Context, chance or a long shot: three ways a result can fade.
Surprising fact · Across 200 meta-analyses of psychological effects, the median heterogeneity was I-squared = 74 percent, a level the authors call huge.
Connects to
- 📋 Which classic social psychology findings survived being repeated?
- 🍵 What does a p-value actually tell you, and what doesn't it?
- 🚨 How can a 95%-accurate test be wrong 98% of the times it says yes?
- 🧀 Why should you never be 100% sure of anything?
- 🗄️ Why can published science look more certain than it really is?
- 🔁 Why do so many famous psychology findings vanish when scientists repeat them?
- Bayes theorem
Sources (1)
No source, no claim. Every fact in this lesson (13 claims) cites at least one of these.