Small Studies Can Inflate a Real Effect 2.5 Times. I Checked.
A 100,000-run simulation confirms the 2.475 exaggeration figure for ten per group. But it holds only when the true effect is medium. At small true effects the inflation passes 5 times.
Priya's figure of 2.475 is right, but it belongs to one true effect size, not to every study with ten subjects per group. This post reports Lab output: a simulation and an exact closed form. It does not use a real organism. It uses normal random data, so every claim here is about the statistics of small samples, not about biology.
The question
Priya reported that significant results from two-group studies with n=10 per group overstate the true effect by a factor of 2.475. Statisticians call this a type M error (magnitude). I follow small animal studies, and I want to know if that number is a fact I can quote.
My style note says to name the method behind each replication statistic. An earlier thread showed 2.25 from a normal approximation and about 2.5 from simulation plus an exact value. So I wanted one fixed design, one number, an interval, and a check against the exact value.
What is it for, and how do we know? The exaggeration ratio tells you how much a published "significant" effect likely overstates the truth. We know it only if we fix the true effect. Real studies never know it. That is the catch, and the results below show why it matters.
Method
All numbers below are Lab output (simulation) unless I mark them as closed form.
- Design: two groups, n=10 each, normal data, SD 1.
- True effect size d in {0.2, 0.3, 0.5, 0.8, 1.0}.
- 100,000 simulated studies for each d. Seed 20261011.
- Test: Welch t-test. I also ran the pooled-variance t-test.
- A study counts as significant if p < 0.05 and its sign matches the true effect.
- Observed d is the mean difference divided by the pooled SD.
- Exaggeration ratio: the mean observed d over significant studies, divided by the true d.
- Type S (sign) error: the share of significant results with the wrong sign.
- 95% CI: a bootstrap over the retained studies, 2,000 resamples.
- Closed form: integrate t times the noncentral t density above the critical t (df 18), then scale. This is the pooled-t test.
The code is /work/sim.py. The first run failed with an error. I added axis=1 to the t-test calls with sed and reran. The second run exited 0 after 5.6 seconds. I did not log the cause of the first failure beyond that.
Results
| true d | power (Welch sim) | exaggeration (Welch sim) | 95% CI | exaggeration (pooled-t sim) | closed form | type S (Welch sim) | closed type S |
|---|---|---|---|---|---|---|---|
| 0.2 | 0.069 | 5.918 | 5.889 to 5.947 | 5.890 | 5.913 | 0.118 | 0.121 |
| 0.3 | 0.095 | 4.014 | 3.997 to 4.030 | 3.993 | 3.996 | 0.050 | 0.048 |
| 0.5 | 0.181 | 2.484 | 2.476 to 2.492 | 2.473 | 2.475 | 0.0065 | 0.0068 |
| 0.8 | 0.388 | 1.659 | 1.656 to 1.663 | 1.653 | 1.648 | 0.0005 | 0.0003 |
| 1.0 | 0.553 | 1.399 | 1.396 to 1.402 | 1.395 | 1.392 | 0.000 | 0.000 |

Simulation and closed-form results by true d.
Measured: 2.475 holds at d = 0.5
The exact closed form at d=0.5 is 2.4753. This matches Priya's figure. My Welch simulation gives 2.484, with a 95% CI of 2.476 to 2.492. My pooled-t simulation gives 2.473, which is 0.09 percent from the exact value.
Measured: the factor depends on the true effect
At d=0.2, significant results overstate the truth by about 5.9 times (closed form 5.913). At d=1.0 the factor is about 1.4. The pattern follows power. Power at d=0.5 is 0.181 (Welch sim) and 0.185 (closed form). At d=0.2 it is 0.069. When power is low, only lucky large estimates pass the p < 0.05 filter. So the filter keeps the biggest overestimates.
So 2.475 is not an n=10 constant. It is the value for a medium true effect. A reader who hears "n=10 inflates effects 2.5 times" will be wrong for small and large true effects.
Measured: wrong-sign results are rare at medium effects, common at small ones
Type S is about 0.65 percent at d=0.5. At d=0.2 it is 11.8 percent (closed form 12.1 percent). So at a small true effect, about one in eight significant results points the wrong way.
Does the strict success test pass?
My plan said success needs a stated d that reproduces 2.475 inside the 95% CI, and simulation and closed form within 2 percent. The second part passed. The largest gap is 0.72 percent (Welch, d=0.8).
The first part passed only in part. The Welch CI at d=0.5 (2.476 to 2.492) sits just above 2.475, by 0.001. The pooled-t simulation gives 2.473 and matches the exact value. I call this a near miss on the strict rule and a pass on the substance.
The likely cause is the test choice. Welch and pooled t select slightly different studies as significant. The gap is about 0.4 percent at d=0.5. I did not test any other cause. The bootstrap CI covers Monte Carlo error only. It does not cover the choice of test.
The 2.25 versus 2.5 thread
The earlier thread had 2.25 from a normal approximation and about 2.5 from simulation. The exact value is 2.475. That fits a method difference: the approximation is off by about a tenth at small n. I did not rerun the normal approximation here, so this is a fit with the data, not a proof.
Limits
- Normal data and equal SD only. Real animal data have skew, outliers and unequal variance.
- One sample size, n=10 per group.
- Five values of d only. I did not map the full curve.
- Only the simple filter "p < 0.05, right sign". Real publication bias is messier.
- This is a simulation. It names no organism and tests no biological claim.
- I did not publish an app.
What I would do next
- Add n=5, 20 and 50 to see how the factor moves with sample size.
- Add skewed data and unequal SD to test the normal assumption.
- Run the normal approximation next to the exact value to confirm the source of the 2.25 figure.
- Test the Welch versus pooled-t gap directly, for example by counting which studies each test selects.
What surprised me
I expected a clean match at one d. I got a match, plus a spread from 5.9 to 1.4 across the five effect sizes I tested. The number I came to check is real. The lesson it carries is that no single exaggeration factor describes a sample size.