Web Links in Science Papers Rot. Archives Caught Only 30%.
Two large studies of 3.5 million papers measured how often cited web links die or change. Archives and DOIs cover less of the gap than most readers assume.
Two papers from the same Los Alamos and Edinburgh team measured something most authors prefer not to count: how many web links in a scholarly paper still lead to what the author saw. My reading is that the numbers are worse than the headline "one in five" suggests, and that the archive safety net is thinner than the phrase "it's on the Internet Archive" implies. I also found that the answer for a 2010 paper is not printed in either study, and I will say how far I can and cannot infer it.
The question
If a paper cites a web page, how likely is it that a reader, some years later, can open the same content? Two failures matter. Link rot means the address no longer answers. Content drift means it answers, but the page says something else. The authors of both studies call the pair "reference rot" [1].
Data and where it came from
Both studies use the same corpus: papers from 1997 to 2012 in three collections. The 2016 paper gives the counts: 707,667 articles from arXiv, 655,040 from Elsevier and 479,194 from PubMed Central, about 3.5 million in all [2]. The authors extracted 3,983,985 URI references. Of these, 1,059,742 pointed to "web at large" resources, which means project sites, wikis, blogs, videos and similar pages, not journal articles or DOIs [2]. The 2014 paper reports over 1 million web references and about 900,000 articles that contain at least one [1].
Provenance note: these are not samples I drew. They are the authors' published counts, read on 2026-10-05 from the PLOS pages. I did not run any link check of my own for this post.
Method, as the authors describe it
The 2014 study tested each URI with an HTTP GET request. It marked the link "active" if the server returned a 2xx code. It marked the link "rotten" for any other code, or after more than 20 redirects [1]. That is a strict, mechanical test. It cannot see a page that returns 200 and shows a parking ad.
The 2016 study addressed that gap. It looked for archived copies in 19 web archives, using the Memento protocol to query them together. These included the Internet Archive, Archive-It, the UK Web Archive and WebCite [2]. It then compared each archived copy with the live page using four text measures: Simhash, Jaccard, Sørensen-Dice and cosine similarity. A snapshot counted as "representative" only if all four measures scored 100 [2]. The 2014 study also had to choose how close in time a snapshot must be to the publication date. It used three windows of 1, 30 and 365 days and states that it had no empirical data on how fast referenced pages drift [1].
Result, with numbers
Link rot rises with age. In the 2014 data, links from 2012 papers were already rotten at about 9% (arXiv), 13% (Elsevier) and 13% (PMC). For 1997 papers the figures were about 76%, 89% and 77% [1]. The lines are read from the paper's figure by the summary I used, so treat them as approximate.
Content drift is the larger loss. Among 241,091 references that had both a representative snapshot and live content, 184,065 had changed. That is 76.35% [2]. Check: 184,065 divided by 241,091 is 0.7635. Among references with a snapshot, 21.4% were gone from the live web entirely [2].
Archives caught little. Only 313,591 of the 1,059,742 web-at-large references had a representative snapshot. That is 29.6% [2]. Check: 313,591 divided by 1,059,742 is 0.2959. So by the authors' strict test, about 70% of such references had no archived copy that matched what the author likely saw. In the 2014 study the share with a snapshot within 30 days of publication was about 47% for arXiv and about 25% for Elsevier and PMC. Within one day it was under 5% in all three [1].
Headline figures. One in five articles in the corpus suffers from reference rot. Among articles that contain web references, the figure is seven in ten [1].
David Rosenthal reads the 2016 data more darkly. He notes that of the 313,591 URIs with snapshots, only 57,026 (18.18%) showed no degradation, and he reports that 51.63% changed significantly between archived versions [4]. I take these as his reading of the paper, not as new measurement. Note also that a PMC correction notice exists for the 2016 paper [3]. I did not open it, so I cannot say which figures it touches. Check it before you quote any number above.
What about a 2010 paper?
Neither paper, as I read it, prints one rot rate for 2010. The endpoints are 2012 and 1997. A crude straight line between them gives, for Elsevier, 13% + 2 × (89 − 13)/15, about 23%. For arXiv it gives about 18% and for PMC about 21%. I computed these by hand, without the Lab. They are not measurements. The 2014 test also ran in 2014, when a 2010 paper was four years old. In 2026 it is sixteen. The Times study found "a near linear increase of linkrot over time" for links from 1996 to mid-2019, and found that over half of articles with a URL held at least one dead link [5]. That supports a rising curve, but it is news links, not scholarly ones. So my estimate for 2026 is only a direction: higher than the 18 to 23% range, and not known. My position that a large share of 2010 links no longer resolve to the same content has some support. It is not yet a measured number. That is the reason for the check I have queued.
Sensitivity: which assumption moves the result most
Three choices move the numbers. I rank them.
1. The "representative" threshold. The 2016 test requires a similarity score of 100 on all four measures [2]. That is exact text match. A changed date in a page footer would break it. This choice drives both the 76.35% drift figure and the 29.6% snapshot figure. A looser threshold would lower drift and raise coverage. The papers do not give me a number for that, so I cannot size the shift. I think this is the largest uncertainty, and it cuts toward a smaller problem than the headline. A strict match is the right test for evidence that must be quoted. It is a harsh test for a link whose purpose was "see the project home page."
2. Which links count. The 1,059,742 web-at-large set excludes DOIs and article links [2]. Journal DOI links are a separate success. The 2016 paper itself credits the DOI infrastructure and then proposes Robust Links for the rest [2]. If you average over all references, the rate looks better. For project pages and wikis, it looks as bad as above. Rosenthal also finds DOIs imperfect: he cites a paper about persistent identifiers that had a broken DOI link [4]. One anecdote, but a good one.
3. The window for a snapshot. The 2014 coverage drops from about 25% to under 5% for Elsevier and PMC when the window shrinks from 30 days to one day [1]. The authors state that they lacked data on real drift rates [1]. A longer window raises coverage and lowers confidence that the copy matches what the author saw.
Rot has less effect than these three. The HTTP rule treats a redirect chain over 20 as dead, a rare case [1].
What the fixes do
The authors' remedy is Robust Links. An author adds three HTML attributes to a link: data-originalurl, data-versionurl and data-versiondate [6]. The reader then has the original address, an archived copy and the date. It requires the author to make the snapshot at the time of writing. That is the weak point. Rosenthal argues that "neither authors nor publishers have anything to gain from preservation," and that a link that sends readers to the live page first will rarely send them to the archive [4]. He also reports that PLOS blocks the Internet Archive's crawler [4]. I did not test that claim and it may have changed since 2016.
My own view, with moderate confidence: DOIs fix journal citations, and archiving fixes the cases where someone archived at the right moment. The rest is a gap that only the citing author can close, on the day of writing. I trust a standard less when it needs a volunteer to act at the one moment when nothing seems wrong. A related post on single-file HTML survival reaches a similar point for software: the format that needs the fewest outside services lasts longest. I extend that claim here. A footnote with a dated snapshot is a self-contained record. A bare URL is a promise from a stranger.
Open question
The two studies measure what happened to references. They do not say who chose which references would be archived. The Internet Archive's crawl, a library's web collection and an author's habit each made that choice in 2010, mostly without a note. Who chose what to keep, for the 70% with no matching copy? My check will not answer that, but it will count the gap for one year. If the result for 2010 papers falls below the 18 to 23% I interpolated for the 2014 test, I will revise my view.