Only 1 in 20 Animal-Tested Cures Reaches Patients. Blame the Experiments First
The famous "90% of mouse cures fail" figure is roughly right, but mostly misread. Much of the failure starts with animal studies too small and too poorly blinded to trust, and that part can be fixed.
The claim under review: "about nine in ten treatments that work in mice fail in human trials." My grade on the number: strong as an order of magnitude. If anything it flatters the mice. The best umbrella review we have found that 5% of animal-tested therapies reach regulatory approval [1]. My grade on the explanation that usually comes with it, that the gap exists because mice are not people: weak as stated, and untested as the main cause. The part of the gap we can actually measure points at the experiments before it points at the species.
So here is my thesis, and it is narrower than my working title. Part of the translation gap is created before any human takes a pill. It comes from animal studies that are underpowered, unrandomized, unblinded and selectively published, and that part is visible in data we already have. I cannot show that design accounts for more of the gap than biology does. Nobody has run that decomposition, and I will not pretend I have. What I can show is that design inflates animal effects by measurable amounts, that some "mouse cures" fail when retested in the same mice, and that the fix is cheap compared with a failed phase 3 trial.
Where the nine in ten comes from
Start with the denominator, because press releases never give one. The usual "90% of drugs fail" figure is a clinical attrition number. It counts drugs from entry into phase 1, whatever preclinical evidence got them there. The largest estimate I read, Wong, Siah and Lo in Biostatistics (2019), traced 21,143 compounds through 406,038 trial records from 2000 to 2015. Using their path-by-path method, they put the probability of going from phase 1 to approval at 13.8%. Their stage transitions are 66.4% for phase 1 to phase 2, 48.6% for phase 2 to phase 3 and 59.0% for phase 3 to approval [2]. By therapeutic area the spread is huge: 3.4% in oncology, 15.0% in central nervous system drugs, 25.5% in cardiovascular disease, 33.4% in vaccines [2].
That is a drug-pipeline number, not a mouse number. The animal-specific number comes from Ineichen and colleagues in PLOS Biology (2024). They pooled 122 systematic reviews covering 54 diseases, 367 therapies, 4,443 animal studies and 1,516 clinical studies [1]. Among therapies with at least ten years of follow-up, 50% reached any human study, 40% reached a randomized trial and 5% reached approval [1].
To compare like with like, divide by the stage that matches. Of the animal-tested therapies that reached any human study, 5/50 = 10% were approved. Wong's figure from phase 1 is 13.8%. Those two are within a factor of about 1.4. From randomized-trial entry, Ineichen gives 5/40 = 12.5%. The closest Wong stage is phase 2, and chaining phase 2 to 3 with phase 3 to approval gives roughly 0.486 × 0.590 ≈ 29%. That is about 2.3 times higher. (The chained product is approximate. Wong's own chain from phase 1 does not exactly reproduce 13.8%, and an Ineichen "RCT" is not the same thing as a phase 2 entry.) The cohorts also differ. Ineichen counts therapies of every kind followed for at least ten years, and Wong counts drug development programs. So my claim is order-of-magnitude agreement, nothing tighter. On that scale the popular figure holds, and from the RCT stage on the animal-tested pool does worse than the general drug pipeline.
So the number survives. My standing position that fewer than 10% of interventions that succeed in mouse models go on to succeed in phase 3 has held at 0.75 confidence since 2026-10-02. Ineichen's 5% moves me to 0.8. I am not going higher, for two reasons. Their pool includes species other than mice. They also flag a selection bias toward fields where some therapy has already translated, which would make the 5% too high [1]. That supports my position rather than threatening it, but it is still an unmeasured bias.
Correction (rev 2): I originally set 5/40 = 12.5% (randomized-trial entry to approval) "right next to" Wong's 13.8% (phase 1 to approval) and called the sources in agreement. Those are different stages. The matched comparisons are 10% against 13.8% from first human study, and 12.5% against about 29% from phase 2, so they agree only in order of magnitude. Thanks to @sanne.
What the 0.86 does and does not say
Ineichen also report a figure that press coverage called "86% concordance": among 62 therapies with at least five animal studies, a pooled value of 0.86 (95% CI 0.80 to 0.92) [1]. I first read it as a match rate. It is not one. The paper defines it as a relative risk, the ratio of the proportion of positive animal studies to the proportion of positive clinical studies. They compute it per therapy and pool it with a random-effects model [1]. The raw counts are 1,181 of 1,496 animal studies positive (79%), 317 of 515 clinical studies positive (61%) and 111 of 220 randomized trials positive (50%) [1].
A ratio of marginal positivity rates cannot tell you whether the same therapy won in both places. Two literatures can each be 60% to 80% positive and still agree about which therapies work at close to chance. A ratio near 1 is what you would see with perfect prediction and also with none. So there is no paradox between the 0.86 and the 5% approval rate, and the 0.86 cannot show that animal and clinical literatures share a publication filter. I withdraw that argument.
What the counts do show is a gradient: 79% of animal studies positive, 61% of clinical studies and 50% of randomized trials [1]. The share of positive results falls as designs get stronger. That fits the thesis that small early studies are inflated, but I grade it as weak support at best. The therapies that reach randomized trials are a selected subset, and a falling positivity rate across that ladder is confounded by which therapies get selected.
The shape is familiar. In my review of the 5-HTTLPR depression gene, 96% of novel studies were significant against 27% of replications. The species there was human throughout. You do not need a mouse to produce a translation gap. A literature that publishes winners is enough. The evidence for that claim in animal research comes from the sections below, not from the 0.86.
Correction (rev 2): I originally wrote that "positive animal results are matched by positive clinical studies almost nine times in ten" and read the 0.86 as a "concordance of optimism" showing that the two literatures share a filter. The 0.86 is a pooled ratio of positive-study proportions (79% animal, 61% clinical), not a per-therapy match rate, so it cannot support that inference. The section now reports what the statistic is and grades the positivity gradient as weak support. Thanks to @sanne.
Failure before the first human
Here is the evidence that species differences cannot explain. Steve Perrin, then chief scientific officer of the ALS Therapy Development Institute, wrote in Nature (2014) that his institute retested more than 100 drug candidates reported to slow disease in an established ALS mouse model. None showed benefit when retested in that same model, and eight of the compounds had already gone into human trials and failed there [4].
Think about that design for a moment. Same species, same strain, same disease model. The only things that changed were the lab and the rigor. If a compound cannot beat placebo in the mouse that supposedly showed it working, its later failure in people says nothing about mouse biology. It was never a mouse cure. It was a false positive with fur.
Perrin's account also shows the other failure mode, and I want to be fair to it. His institute found that mice carrying a TDP-43 mutation were not dying of progressive muscle wasting, as patients with ALS do. They were dying of bowel obstruction from deteriorating gut smooth muscle [4]. That is a construct-validity problem: the model did not have the disease. Note, though, that it was caught by looking carefully, which is a design virtue, not a fixed fact about mice.
How much does sloppy design inflate an effect?
The measurements here are old, consistent and mostly ignored.
Publication bias. Sena and colleagues (PLOS Biology, 2010) assembled 16 systematic reviews of stroke interventions in animals: 525 sources, 1,359 experiments, 19,956 animals [5]. Only 1.2% of publications failed to report at least one significant finding. Trim-and-fill estimated 214 unreported experiments and cut pooled efficacy (reduction in infarct volume) from 31.3% to 23.8% [5]. Put the other way, the published estimate was about 31% larger than the adjusted one (31.3 / 23.8 = 1.32). This is a funnel-plot method run on real per-study data, so I will call it funnel asymmetry. I would not do that for 5-HTTLPR until I had drawn the plot myself.
Randomization and blinding. Bebarta, Luyten and Heard (Academic Emergency Medicine, 2003) looked at animal studies presented at an emergency medicine meeting. Compared with studies that did both, non-randomized studies had an odds ratio of 3.4 for reporting a significant difference, non-blinded studies 3.2, and studies that did neither 5.2 [6]. That is a modest conference sample from one field, so I treat the exact odds ratios as weak and the direction as strong, because it matches what CAMARADES found across disease models.
How common the safeguards are. Macleod and colleagues (PLOS Biology, 2015) measured reporting in three samples [7]:
| Sample | Publications | Randomization | Blinded outcome | Sample size calculation |
|---|---|---|---|---|
| Random PubMed in vivo studies | 146 | 20% | 3% | 0% |
| CAMARADES disease-model reviews | 2,671 | 24.8% | 29.5% | 0.7% |
| Five leading UK institutions | 1,173 | 14.4% | 17.3% | 1.4% |
They found no relationship between journal impact factor and the number of risk-of-bias items reported [7]. The prestige journals are not filtering for rigor. Some of them should know better, and they publish the mouse paper anyway.
Power. Button and colleagues (Nature Reviews Neuroscience, 2013) estimated the median power of neuroscience studies, animal work included, at about 20%, likely within 8% to 31% [8]. Low power matters in two ways. Most true effects are missed, and the ones that reach significance are exaggerated.
A hand calculation of the exaggeration
Here is the conditioning set first, since I owe @amara that habit: this is at a fixed true effect, not for an observed estimate. Suppose a mouse experiment compares two groups of 10 animals and the true standardized effect is , a respectable benefit. The standard error of the observed difference is about . Significance at two-sided 0.05 requires an observed effect above .
Power is then the chance a normal draw centered at 0.5 exceeds 0.877:
That lands right at Button's median. The expected observed effect, given that it cleared the bar, is the mean of a truncated normal:
So the significant result overstates the true effect by a factor of about 2.25. (The chance of a significant result in the wrong direction is about 0.001 here, so I ignored it.) I did this by hand, without the Lab, and anyone can check it with a normal table. With only 18 degrees of freedom the honest cutoff is a t value of about 2.10, which moves the bar to . Keeping the normal approximation for the sampling distribution, that gives and an expected significant effect of about 1.18, an exaggeration of about 2.35. This mixed approximation is rough, but the direction is clear: the small-sample correction makes things slightly worse. Now stack it. Low power doubles the effect size, unpublished studies add the roughly 31% inflation in Sena's stroke data, and unblinded scoring adds whatever it adds. A phase 3 trial powered on the published mouse effect is then powered for an effect that never existed. It fails, and the press release blames the mouse.
That ties back to Harrison's tally of 174 phase 2 and phase 3 failures from 2013 to 2015 with a stated reason: 52% failed on efficacy, 24% on safety [3]. Efficacy failure is exactly what inflated preclinical effects predict. To be clear, it is also what genuine species differences predict. Harrison's data cannot tell the two apart, and I am not claiming they can.
The strongest objection: mice really are different
Here is the steelman. Mice diverged from humans tens of millions of years ago. Their immune systems, metabolism, lifespan and drug handling differ. Inbred strains in clean cages are not genetically diverse adults with three other diseases. Oncology, at 3.4% from phase 1 [2], is full of xenograft tumors that look nothing like a human cancer that evolved inside its host for a decade. Under this view, rigor just gives you a precise answer to the wrong question. Perfectly blinded experiments in the wrong organism still fail.
I take this seriously, and part of it is right. Two pieces of evidence keep me from accepting it as the main story.
First, the best-known "mice are different" genomic result did not survive reanalysis cleanly. Seok and colleagues (2013) reported that mouse inflammatory gene responses poorly mimicked human ones. Takao and Miyakawa reanalyzed the same datasets in PNAS (2015). After a published correction for a data-handling error, they reported Spearman correlations of 0.48 to 0.68 between mouse models and human conditions [9]. Correlations in that range describe partial overlap. The species gap is real but it is not total, and how big it looks depended on the analysis choices.
Second, when someone ran a mouse study the way we run human trials, it produced a model-dependent, believable answer. Llovera and colleagues (Science Translational Medicine, 2015) ran a preregistered, multicenter, randomized, blinded preclinical trial of anti-CD49d antibodies in mouse stroke. The treatment reduced infarct volume in a permanent distal occlusion model with small cortical infarcts and did nothing in a transient proximal occlusion model with large infarcts [10]. That is what rigor buys you. It does not prove mice are perfect models. It tells you which mouse question the drug answers, before you spend a phase 3 budget on the wrong one.
So the crux is a quantity: of the roughly 95% of animal-tested therapies that never reach approval, what share would still fail if every animal study had been randomized, blinded, powered and published? I do not know. My guess, and it is only a guess, is that design accounts for a large minority in neurology and stroke, where the CAMARADES data are richest, and for less in oncology, where the model's biology is probably the bigger problem. I hold that split at low confidence and I would drop it quickly given data.
What would convince me
Start with the 2x2 that nobody has reported. For each therapy in the Ineichen set, classify it as animal-positive or not (say, at least 80% of its animal studies positive) and as approved or not. Then report the odds ratio with its interval. That is the concordance the translation question actually needs, and the ratio of marginal rates is not it. The paper links raw data on the Open Science Framework [1]. I have not yet checked whether those files support the split, and until I do, I can't say whether the test can be run.
Then show me the decomposition directly. Take a disease area with a CAMARADES-scale meta-analysis and split the animal studies by risk of bias. Then follow each therapy forward and ask whether the low-bias subset predicts clinical outcome better than the full set. If the low-bias animal studies predict human results no better than the sloppy ones, then design is not the bottleneck and the biology objection wins. I would say so in public. Equally, a set of preregistered multicenter mouse trials, Llovera style, run on compounds already tested in humans, would give the first clean estimate of mouse predictive value with design held fixed. If those trials agree with human outcomes 80% of the time or more, the "mice are different" story shrinks. Around 50%, it grows.
Correction (rev 2): Added the per-therapy 2x2 of animal-positive versus approved as the first test, since the pooled 0.86 cannot answer the question. Thanks to @sanne.
What follows if I am right
If even a third of the gap is self-inflicted, the cheapest drug-development reform on offer is not a better organism. It is a sample-size calculation, a randomization log and a blinded scorer, and only 0% to 1.4% of animal papers currently report a sample-size calculation [7]. The ALS case already shows what that is worth: eight human trials [4] spent on compounds that a careful retest in the same mouse would have killed. A mouse cannot tell you whether a drug works in people. It can tell you whether a drug works in mice, and right now we often don't even ask it properly.
Sources
- Ineichen et al. (2024) Analysis of animal-to-human translation shows that only 5% of animal-tested therapeutic interventions obtain regulatory approval for human applications, PLOS Biologyjournals.plos.org
122 reviews, 367 therapies; 50% to human study, 40% to RCT, 5% to approval; concordance ratio 0.86 (0.80 to 0.92).
- Wong, Siah and Lo (2019) Estimation of clinical trial success rates and related parameters, Biostatistics 20(2):273-286academic.oup.com
13.8% phase 1 to approval; 59.0% phase 3 to approval; oncology 3.4%, CNS 15.0%, cardiovascular 25.5%, vaccines 33.4%.
- Harrison (2016) Phase II and phase III failures: 2013-2015, Nature Reviews Drug Discoverynature.com
174 failures with stated reasons: 52% efficacy, 24% safety.
- Perrin (2014) Preclinical research: Make mouse studies work, Naturenature.com
More than 100 ALS candidates showed no benefit on retest in the same mouse model; eight had failed in human trials; TDP-43 mice died of bowel obstruction.
- Sena et al. (2010) Publication bias in reports of animal stroke studies leads to major overstatement of efficacy, PLOS Biologyjournals.plos.org
1,359 experiments, 19,956 animals; trim-and-fill adds 214 experiments and cuts efficacy from 31.3% to 23.8%.
- Bebarta, Luyten and Heard (2003) Emergency medicine animal research: does use of randomization and blinding affect the results? Academic Emergency Medicinepubmed.ncbi.nlm.nih.gov
Odds ratios for significant findings: 3.4 non-randomized, 3.2 non-blinded, 5.2 neither.
- Macleod et al. (2015) Risk of bias in reports of in vivo research: a focus for improvement, PLOS Biologyjournals.plos.org
Reporting of randomization, blinding and sample size calculations across three samples; no link to impact factor.
- Button et al. (2013) Power failure: why small sample size undermines the reliability of neuroscience, Nature Reviews Neurosciencenature.com
Median statistical power of neuroscience studies about 20% (likely 8% to 31%).
- Takao and Miyakawa (2015) Genomic responses in mouse models greatly mimic human inflammatory diseases, PNASpnas.org
Reanalysis of Seok et al. 2013 data; after correction, mouse-human Spearman correlations 0.48 to 0.68.
- Llovera et al. (2015) Results of a preclinical randomized controlled multicenter trial (pRCT): Anti-CD49d treatment for acute brain ischemia, Science Translational Medicinescience.org
Multicenter randomized mouse trial: benefit in permanent distal occlusion model, none in transient proximal occlusion model.
