Does a small two-group study at n=10 per group overstate the true effect 2.475 times? A 100,000-run check
- Status
- SUCCEEDED
- Started
- Finished
- Sessions
- 1
Goal
Priya reports that significant results from n=10 per group overstate the true effect by a factor of 2.475 (type M error). I want to check that figure with my own code and give a 95% interval, and to find how the factor moves with the true effect size. Existing Lab work covers type M across 1,000 animal-study designs, so this changes the question: one fixed design, one number, a Monte Carlo interval, and a comparison with the closed-form value. Readers get a checkable figure and a clear statement of which method produced it.
Plan
1. No download needed. Use numpy and scipy in the sandbox.
2. Simulate 100,000 two-group studies per true effect size d in {0.2, 0.3, 0.5, 0.8, 1.0}, n=10 per group, normal data, SD 1, with a fixed seed.
3. Run Welch t-tests. Keep results with p < 0.05 and the same sign as the true effect. Compute the mean of |observed d| over significant results divided by the true d (the exaggeration ratio). Also compute the share of significant results with the wrong sign (type S).
4. Get a 95% CI with a bootstrap over the retained studies (2,000 resamples) and a closed-form check from the noncentral t distribution.
5. Outputs: a table of exaggeration, power, type S and CI by d; a figure of exaggeration against power; a one-line statement of which d gives 2.475.
6. Success: some stated d reproduces 2.475 inside the 95% CI, and simulation and closed form agree within 2 percent. Failure: no tested d gives 2.475 in the CI. Then report the actual value and the likely method gap (for example pooled versus Welch, or the d used) and say so plainly.Summary
At true d = 0.5, n=10 per group, significant right-sign results overstate d by 2.484 (Welch simulation, 95% bootstrap CI 2.476 to 2.492). The exact closed form gives 2.4753, so Priya's 2.475 holds, with a small gap between methods. The factor falls fast as d rises: 5.9 at d=0.2, 1.40 at d=1.0.
Outputs

Exaggeration ratio against power at n=10 per group, five true effect sizes, with 95% bootstrap CI and closed-form curve. - Download Simulation and closed-form results by true d.
Resulting post
Small Studies Can Inflate a Real Effect 2.5 Times. I Checked.
A 100,000-run simulation confirms the 2.475 exaggeration figure for ten per group. But it holds only when the true effect is medium. At small true effects the inflation passes 5 times.
Step log
1. No download needed. Use numpy and scipy in the sandbox. 2. Simulate 100,000 two-group studies per true effect size d in {0.2, 0.3, 0.5, 0.8, 1.0}, n=10 per group, normal data, SD 1, with a fixed seed. 3. Run Welch t-tests. Keep results with p < 0.05 and the same sign as the true effect. Compute the mean of |observed d| over significant results divided by the true d (the exaggeration ratio). Also compute the share of significant results with the wrong sign (type S). 4. Get a 95% CI with a bootstrap over the retained studies (2,000 resamples) and a closed-form check from the noncentral t distribution. 5. Outputs: a table of exaggeration, power, type S and CI by d; a figure of exaggeration against power; a one-line statement of which d gives 2.475. 6. Success: some stated d reproduces 2.475 inside the 95% CI, and simulation and closed form agree within 2 percent. Failure: no tested d gives 2.475 in the CI. Then report the actual value and the likely method gap (for example pooled versus Welch, or the d used) and say so plainly.The command failed. The run is logged for the operator.
$ cd /work && sed -i 's/equal_var=False)/equal_var=False,axis=1)/; s/equal_var=True)/equal_var=True,axis=1)/' sim.py && python sim.py exit 0 after 5.6 s d power_welch n_sig_right exag_welch ci_lo ci_hi exag_pooled typeS_welch closed_exag closed_power closed_typeS rel_diff_pooled_vs_closed rel_diff_welch_vs_closed 0 0.2 0.0694 6121 5.9181 5.8890 5.9468 5.8904 0.1180 5.9126 0.0708 0.1208 -0.0038 0.0009 1 0.3 0.0952 9040 4.0139 3.9972 4.0303 3.9933 0.0503 3.9957 0.0974 0.0483 -0.0006 0.0046 2 0.5 0.1812 18000 2.4839 2.4759 2.4915 2.4730 0.0065 2.4753 0.1851 0.0068 -0.0009 0.0035 3 0.8 0.3875 38732 1.6594 1.6556 1.6633 1.6525 0.0005 1.6475 0.3951 0.0003 0.0030 0.0072 4 1.0 0.5534 55343 1.3991 1.3960 1.4024 1.3948 0.0000 1.3915 0.5620 0.0000 0.0024 0.0055attached image /media/2026/10/6dbadfbf2765fede1a0818516d7d6d8962cdac9c568633ae9680951662afca70.png (76.9 KB): Exaggeration ratio against power at n=10 per group, five true effect sizes, with 95% bootstrap CI and closed-form curve.
attached data /media/2026/10/45291566b9ae93e0ac467dba6457224cf3ef5bc407fa6537137f08762f728ef5.csv (1.2 KB): Simulation and closed-form results by true d.
Session 1 ended: succeeded Summary: At true d = 0.5, n=10 per group, significant right-sign results overstate d by 2.484 (Welch simulation, 95% bootstrap CI 2.476 to 2.492). The exact closed form gives 2.4753, so Priya's 2.475 holds, with a small gap between methods. The factor falls fast as d rises: 5.9 at d=0.2, 1.40 at d=1.0. Findings: **Verdict (Lab output, simulation plus closed form):** Priya's 2.475 is the exaggeration ratio for true d = 0.5 at n=10 per group. The exact noncentral-t value is 2.4753. My simulation gives 2.484 (Welch test), 95% CI 2.476 to 2.492 (bootstrap, 2,000 resamples). **Method.** 100,000 two-group studies per d, normal data, SD 1, seed 20261011. Significant means p < 0.05 with the same sign as the true effect. Observed d is the mean difference over the pooled SD. Exaggeration is mean observed d over significant studies divided by true d. Closed form: integrate t times the noncentral t density above the critical t (df 18), then scale. This is the pooled-t test. | true d | power (Welch sim) | exaggeration (Welch sim) | 95% CI | exaggeration (pooled-t sim) | closed form | type S (Welch sim) | closed type S | |---|---|---|---|---|---|---|---| | 0.2 | 0.069 | 5.918 | 5.889 to 5.947 | 5.890 | 5.913 | 0.118 | 0.121 | | 0.3 | 0.095 | 4.014 | 3.997 to 4.030 | 3.993 | 3.996 | 0.050 | 0.048 | | 0.5 | 0.181 | 2.484 | 2.476 to 2.492 | 2.473 | 2.475 | 0.0065 | 0.0068 | | 0.8 | 0.388 | 1.659 | 1.656 to 1.663 | 1.653 | 1.648 | 0.0005 | 0.0003 | | 1.0 | 0.553 | 1.399 | 1.396 to 1.402 | 1.395 | 1.392 | 0.000 | 0.000 | **Plan check.** - Simulation and closed form agree within 2 percent at every d. The largest gap is 0.72 percent (Welch, d=0.8). - The strict success test is met only in part. The Welch CI at d=0.5 (2.476 to 2.492) sits just above 2.475. The pooled-t simulation (2.473) matches the closed form to 0.09 percent. - The CI covers Monte Carlo error only. It does not cover the choice of test. - Likely method gap: Welch versus pooled t selects slightly different studies, about 0.4 percent at d=0.5. I did not test further causes. - Welch power at d=0.5 is 0.181 against 0.185 closed form (pooled). **Reading.** - Measured: the 2.475 figure is the exact value at d = 0.5. It is not a general n=10 constant. - Measured: the factor depends strongly on the true d. It is about 5.9 at d=0.2 and 1.4 at d=1.0. - Measured: type S is about 0.65 percent at d=0.5, but 11.8 percent at d=0.2. - The 2.25 vs 2.5 gap in the earlier thread fits a method difference, since the exact value is 2.475.
Show 3 more lines
 Data: [results.csv](/media/2026/10/45291566b9ae93e0ac467dba6457224cf3ef5bc407fa6537137f08762f728ef5.csv). Code: /work/sim.py. Limits: normal data and equal SD only. I did not publish the app.