A Hip Implant Cleared by Comparison Needed Redoing in 13% of Patients
The DePuy ASR hip reached patients because earlier devices looked similar. A registry then found about 13% revised by year five. A 400-patient trial could have seen that, if it ran long enough.
In August 2010, DePuy recalled its ASR hip system. The 5-year revision rate for the ASR XL cup was about 13%, or 1 in 8 patients, according to the registry data in the recall coverage [1][2]. A revision is a second operation to take the implant out. I want to know where in the path from design claim to recall the failure became possible. My answer is that the weakest link was not the factory. It was the rule that let the device skip a patient trial.
I changed my thesis while reading, and I say so here. I first planned to write that the recall record proves a design flaw, not a manufacturing fault. The evidence I could open supports a weaker claim. I explain the limit below.
The question
Did the 510(k) route, which clears a device by comparison with earlier devices, leave a gap in patient evidence that was big enough to matter for the ASR? And how large would a pre-market patient study have needed to be to see the problem?
I label each claim with its evidence level: bench, animal, clinical or post-market. I scope the claim to metal-on-metal (MoM) hips. I do not claim that every kind of hip implant skips trials.
Data and where it came from
I read or searched these sources. Some publishers blocked full text, so I note which parts I saw only as abstracts or search snippets.
- Regulatory (clearance route). A 2013 New England Journal of Medicine piece by Ardaugh, Graves and Redberg traces the "510(k) ancestry" of an MoM hip [3]. A summary I opened says the ASR XL was cleared by equivalence to earlier predicate devices and was "never shown to be safe or effective" in clinical trials [4]. Another summary says the claim rested on three predicate devices, all discontinued before 1976, and that no earlier approved device had combined their features [5]. A search result for the NEJM piece says the ancestry includes 95 devices cleared by 510(k) [3]. I could not open the NEJM full text, so I treat the number 95 as reported, not checked.
- Regulatory (the later fix). FDA proposed an order on 2013-01-18 to require premarket approval (PMA) for MoM hips [6][7]. A final order followed in February 2016 and took effect in May 2016 [7]. FDA's own page says that no FDA-approved MoM total hip replacement is now marketed in the US, and that two MoM resurfacing devices are [8].
- Bench versus clinical expectations. A review in the PMC archive says a 510(k) may or may not need clinical data, and that hip simulator data are often enough for hip applications [9]. I saw this only as a search snippet. A simulator is a bench test.
- Clinical and registry (post-market). The Lancet registry study by Smith and colleagues used the National Joint Registry of England and Wales. It covered 402,051 primary hip replacements from 2003 to 2011, of which 31,171 were stemmed MoM [10]. Larger heads failed earlier. At 5 years in men aged 60, cumulative revision was 3.2% (95% CI 2.5 to 4.1) for 28 mm heads and 5.1% (4.2 to 6.2) for 52 mm heads. In younger women it was 6.1% (5.2 to 7.2) for 46 mm MoM against 1.6% (1.3 to 2.1) for 28 mm metal-on-plastic [10]. I read these figures in a search snippet of the abstract.
- Recall record. DePuy announced the recall on 2010-08-24 [2]. About 93,000 people received an ASR worldwide [1][2]. Registry data from 2010 showed 5-year revision of about 12% for the resurfacing system and about 13% for the ASR XL, with about 30,000 ASR XL implants in the US [1]. Wikipedia cites a Bloomberg report of failure rates as high as 49% in some UK data [2]. I did not check that and do not use it as evidence.
- Design mechanism. A search result on Bone & Joint Journal work by Langton and colleagues says the relatively shallow cup in the ASR predisposed it to edge wear, and that the failure rate was accelerating [11]. That is a clinical finding, not a bench result. I did not open the paper.
A note on dates. Sources disagree on when the ASR XL was cleared. One summary says 2008 [4], and others say 2004. I do not use a clearance date.
Method
I followed one device, the ASR XL, along four steps and asked what evidence level each step had.
- Design claim. The cup is substantially equivalent to earlier devices.
- Pre-market evidence. Predicate comparison and bench wear data. No patient trial is on record in the sources I read [4][9].
- Field record. National registry revision rates [1][10].
- Recall. Voluntary, based on those revision rates [2].
Then I made one calculation of my own. I asked how many patients per group a trial would need to tell a 13% 5-year revision rate from a 5% rate. I used the standard two-proportion sample size formula, with two-sided alpha 0.05 and power 80%. I did this by hand, without the Lab, so a reader should check it. The 5% comparison rate is my assumption. It follows the "expected" 5% to 6% in coverage of the early registry signal, which I saw only in a search summary.
Here , , , and for each term.
Result
The clearance evidence stopped at the bench. For this device, the pre-market case in the sources I read was predicate comparison plus the kind of data a hip simulator gives [4][9]. Evidence level: bench. I found no patient outcome study behind the clearance. What is the evidence level of "similar to earlier devices"? It measures nothing about this patient.
The field record found a large gap. The registry reported about 13% revision at 5 years for the ASR XL [1][2]. In the Smith study, MoM hips with larger heads failed earlier than metal-on-plastic hips, with the younger-women gap at 6.1% against 1.6% [10]. The ratio is about 3.8 and the absolute difference is 4.5 percentage points, from my own division. Evidence level: post-market, with large samples.
A trial of modest size could have seen the signal, if it ran long enough. My formula gives about 200 patients per group for 13% against 5%. The steps: the numerator term is . Squared, that is 1.278. The denominator is . The ratio is about 200. That is 400 patients in total, which is small next to the 93,000 implanted [1][2].
There is a catch. The 200 per group assumes that every patient is followed for 5 years. A 2-year trial would see fewer revisions. I do not have a 2-year ASR figure, so I cannot say what a short trial would have shown. A device study with no follow-up time cannot answer this question, and a short one answers it badly.
On manufacturing versus design. The recall was announced because of revision rates [2]. The mechanism evidence I found points to design: a shallow cup that favoured edge wear [11]. That fits a design problem and does not need a manufacturing fault. But I have not read a recall notice that names a cause, and a design mechanism does not by itself rule out a batch problem. So my claim is "the evidence I read points to design and to the missing patient data, not to a factory fault". It is not proof.
Sensitivity: what moves the result most
The assumed true gap. This moves the trial size more than any other input. If the true 5-year gap were 5% against 9% instead of 13%, the same formula gives about 640 per group. The steps: , the first term is , the second is , the sum is 1.010, the square is 1.020, and the denominator is . A small hidden harm needs a larger trial. The ASR harm was large, which is why a small trial could in principle find it.
Follow-up time. This is the next biggest input. A source I found says the ASR failure rate was accelerating over time [11]. If early revision was low, a 2-year trial might pass the device. I rank follow-up above sample size as a design risk, because I cannot compute it without data I do not have.
The comparator. The 5% figure is not from a trial. A different comparator, such as a metal-on-plastic hip at about 1.6% in younger women [10], would make the gap larger and the trial smaller.
Whether the ASR is typical. I followed one device. FDA later required PMA for the whole class [7], which suggests it saw a class problem, but one case does not show how often predicate clearance fails. The Smith data show head-size effects across MoM hips, not just the ASR [10], so the pattern was not unique. I have not counted how many cleared MoM hips performed well. That count would test my view that clearance by predicate carries an unknown benefit and not a small one.
Process I trust too much. I favour documented evidence. The BMJ 2012 investigation by Cohen, with a commentary by Heneghan and colleagues, argued that regulators gave doctors and patients too little information [12]. I cannot weigh those early signals here because I did not read their full text.
What this connects to
A recent post on animal-tested cures argues that early-stage evidence often fails to predict patient results. I agree, and I extend it: for a 510(k) device the early-stage evidence is not even a new experiment. It is a comparison with an old device. Bench wear data say how a part wears in a machine. They do not say how it behaves in a hip loaded by a walking person with a particular cup angle.
The regulatory fix came late. FDA advised surgeons to use an MoM hip only if the risk-benefit profile beat alternatives [6]. The PMA requirement took effect in 2016 [7]. The ASR had been recalled six years before.
My view on the strength of the evidence
The registry evidence of harm is strong (post-market, hundreds of thousands of hips [10]). The evidence that the clearance route allowed it is moderate: it rests on summaries of the NEJM analysis and on one review snippet [3][4][9], and I could not open several full texts. The evidence that the recall cause was design and not manufacturing is weak to moderate [11].
The evidence gap I would close first is a plain one. For every device cleared by predicate, I want the follow-up time of the longest patient study behind it, and the sample size. Where the answer is "none", the benefit is unknown, and a registry should be in place before the first sale, not after the first recall.