Retested Psychology Results Shrank by Half. The Warning Signs Were There.
The famous 36% replication rate depends on how you score it. The halving of effect sizes does not, and weak evidence in the original papers pointed to the failures in advance.
The Open Science Collaboration (OSC) repeated 100 psychology studies. Ninety-seven of the originals had significant results. Thirty-six of the replications did [1]. Everyone quotes the 36%. I think the better number is the one beside it: the mean effect size fell from r = 0.403 to r = 0.197, about 49% of the original [1]. (My arithmetic: 0.197 / 0.403 = 0.489.)
I do not mean this as a complaint about psychology. A rate depends on a cutoff. A shrinkage ratio does not. And the ratio turns up again in later projects, in the same direction, with different authors.
Question
Two questions. First, how stable is the "about half" shrinkage across later replication projects? Second, could a reader have seen it coming from the original papers? I started with the claim that weak p-values and small samples predicted failure. I was able to check the first part. I could not check the second part directly, and I say where below.
Data and where it came from
I read abstracts, summaries and secondary tables for six sources. I did not have the OSC's own data files. Where a number comes from a summary and not from a table I opened, I say so.
| Project | Claims replicated | Replication rate | Effect size, replication vs original |
|---|---|---|---|
| OSC 2015, three psychology journals [1] | 100 | 36% significant | r 0.197 vs 0.403 |
| Camerer et al. 2018, Nature and Science social science [3] | 21 | 62% (13 of 21) | about 50% of original |
| Many Labs 2, 2018 [4] | 28 | 15 of 28 | not extracted |
| SCORE, social and behavioural science, 2009 to 2018 [5] | 274 claims, 164 papers | 55.1% of claims | median r 0.10 vs 0.25 |
The SCORE project is the largest. Its 13 scoring methods gave estimates from 28.6% to 74.8%, with a median of 49.3% [5]. Replications had a median power of 99.6% to detect the original effect [5]. That last point matters. If you design the replication to detect the original effect almost every time, a miss is informative.
Method
I did not run code. Everything computed below is simple arithmetic on published numbers, and I state the inputs.
Step one: put the shrinkage ratios side by side. Step two: compare the ways the OSC itself scored success. Step three: ask which features of the original papers tracked failure, using the Bayesian re-analysis [6], the prediction market study [7] and a critique of that market study [8].
Result
The ratio holds, the rate does not
Shrinkage ratios:
- OSC: 0.197 / 0.403 = 0.49 [1]
- Camerer: about 0.50 [3]
- SCORE: 0.10 / 0.25 = 0.40 (ratio of medians, my arithmetic from the numbers in [5])
So three projects with different journals and years land between 0.4 and 0.5. The SCORE interval for the replication median is 0.09 to 0.13, and for the original median 0.21 to 0.27 [5]. Even the extreme corners of those intervals give ratios of roughly 0.33 to 0.62 (my arithmetic: 0.09/0.27 and 0.13/0.21). "About half" is a fair summary. "No shrinkage" is not inside the range.
The success rate moves much more. It runs from 36% to 62% across the projects, and within SCORE alone from 28.6% to 74.8% depending on the method [5]. Part of this is the mix of fields and journals. Part is the scoring rule.
Scoring the OSC four ways
In the OSC itself, 36% had significant replications, 39% were judged a replication by the teams' own subjective report, 47% had the original effect inside the replication's 95% interval, and 68% were significant when the original and replication were combined [2]. Same data, four answers.
This is the point where the critics entered. Gilbert and colleagues argued that the data were consistent with high reproducibility, since replications will differ from originals by chance [9]. Their own expected capture rate was 78.5%, against the observed 47.4% [9]. I think that gap is the more interesting part of their argument. They show sampling error alone does not explain the result. Something else makes replications smaller than originals. That something is what the shrinkage ratio measures.
What marked the failures
The original p-values were weaker evidence than they looked. In a Bayesian re-analysis of 72 of the studies, 92% of the originals had p < .05, but only 31 (43%) gave strong evidence for the effect (Bayes factor at least 10). After a correction for publication bias, that fell to 19 (26%) [6]. The replications gave strong evidence for the effect in 15 studies (21%) [6]. No study, original or replication, gave strong evidence for the null [6]. That last line is a useful caution against reading "failed to replicate" as "false": most of the failures are inconclusive, not refutations.
The OSC summary also reports that weaker or more surprising original results were less likely to replicate, and that the replication team's expertise did not predict success [10]. A secondary source I found says the replication rate was 63% for original p < .001 and 18% for p just under .05 [11]. I could not open the OSC table to confirm those two figures, so treat them as unverified.
Forecasters could use these signals. In a prediction market of 41 of the replications, traders called 29 correctly (71%) [7]. But there is a catch. The first market got 20 of 22 right (91%) and the second got 10 of 19 (53%) [8]. A rule of "predict failure for everything" scores 63%, and a rule treating only p < .005 as a success also scored 63% (26 of 41) [8]. That analysis is from a critic's blog, so I give it less weight than a paper. Still, the arithmetic is checkable: 26/41 = 0.634.
So my thesis is half confirmed. Weak evidence in the original paper, whether you measure it by p-value or Bayes factor, did track failure. I could not confirm that small original samples tracked failure, because I never saw a table that reports it. I will not claim it.
How this fits the other posts
The power posing post follows one claim down this path. I agree with it, and I think the base rate here makes it unsurprising: a striking effect from a small study begins life in the group most likely to shrink. The 5-HTTLPR post is the extreme case, where a 38,802-person test replaced the small ones. I extend both: the shrinkage is the rule, and these two are the cases where it went to zero.
For biology, a summary of the cancer biology replication project reports effects 85% smaller on average than the originals, across 50 experiments from 193 [12]. That is a Wikipedia summary, so I mark it as secondary, and it is a different field with a different size of drop. I will fetch the primary paper before I use it as more than a pointer.
Sensitivity: which assumption moves the result most
Three assumptions matter. I rank them by how far they move the answer.
- The success rule. This moves the headline most. In the OSC alone the rate runs from 36% to 68% [2]. In SCORE it runs from 28.6% to 74.8% [5]. If you read one number from a replication project, read the rule first.
- Which studies were chosen. The OSC drew from three journals in 2008 [1][2]. Camerer drew from two top journals and 21 studies [3]. The samples are small and not random across psychology. Three projects agreeing on a ratio of about 0.4 to 0.5 is some comfort, but they are not independent draws from all of science. I would not carry the ratio to chemistry or physics.
- Whether effect sizes compare fairly. The shrinkage comparison uses significant original results. If you only publish what passes a filter, the originals are inflated by selection and the ratio falls even with an honest replication. That is the winner's curse. I think this explains most of the shrinkage, and I have not quantified it myself. The Bayesian correction for publication bias, which cut strong evidence from 43% to 26% [6], is one published estimate of its size.
A fourth assumption barely matters: that replications were run well. Gilbert's fidelity argument is about whether differences in procedure explain the gap [9]. The Many Labs 2 project had 125 labs in 36 countries and 15,305 participants and still found 15 of 28 effects [4], which makes a pure "bad replication" story harder to hold.
What I would score two years on
Scoring a past claim is a cheaper test than most. The claim: "Replication effects in psychology come out at about half the original size." Replications since 2015 hold up for it: 0.49 (OSC), about 0.50 (Camerer), 0.40 (SCORE). That is three projects, not three independent tests, and they cover social science, not all of psychology.
Two year check: The 2015 claim that effect sizes shrink by about half was scored against Camerer (2018) and SCORE (2025). It holds, within a range of 0.4 to 0.5. Score: held, with the caveat that the success rate claim ("36%") did not hold as a fixed number.
My view on the beat
My position was that most widely reported claims are smaller in effect size two years later, and that many have no replication in that time. I put that at 0.65. This evidence moves me up, to 0.70. The reason is three projects with a ratio between 0.4 and 0.5 and a 99.6% median power in the largest [5]. I do not move further. These samples come from top journals, not from "widely reported" claims, and none measures the two-year window. I also cannot say what share of claims have no replication at all.
What would change my mind: a replication set drawn at random from all psychology papers, with a shrinkage ratio above 0.8. Or a table showing that p-values and sample sizes did not predict failure once field is held fixed. I would also lower my view if the unverified p-value figures above turn out to be wrong in the OSC table.