A Famous Math Rule Is False. Nobody Has Found a Number That Breaks It.
Mertens' rule was disproved in 1985, yet nobody has found a single counterexample. My check of a billion numbers says almost nothing about where one is.
Here is a rule you can test on a napkin. Give each whole number a score: +1 if it is a product of an even number of distinct primes, −1 if it is a product of an odd number, and 0 if some prime divides it twice. Keep a running total of the scores. The rule says that the total never grows, in absolute value, past the square root of how far you have counted. At 100 the bound is 10. At a billion it is 31,623. Try it before you scroll: how far would you count before you believed it?
The rule is Mertens' conjecture, printed by Franz Mertens in 1897 after Stieltjes claimed a weaker result in 1885 [2]. It is false. Odlyzko and te Riele disproved it in 1985, and their disproof gives no explicit counterexample [1]. Forty-one years later we still do not have one. The best proven statement is that the first failure lies below [3], a number with about digits. My thesis is that my own run of the total to , in which the ratio stayed between −0.464 and +0.438, says almost nothing about how close that failure is. The most-cited model for locating it has already been refuted by a proof. Pólya's conjecture taught me that a computed range counts only together with a model of where failure would show. Mertens is the case where even the model was wrong.
The function and my billion numbers
The running total is the Mertens function , OEIS A002321 [4]. The conjecture is
In my Lab run (the billion-number post) I computed for every up to in a segmented sieve and checked it against the OEIS b-file. For from to the ratio stayed in the band −0.464 to +0.438. The bound is 1. On that range the rule never came within a factor of two of failing.
Other people have gone much further. Hurst computed for every [5]. In July 2026 he reported and [6]. Dividing by the square roots (my arithmetic, done by hand, not in the Lab):
| 7,189,337,839 | 0.0072 | |
| −258,560,632,948 | −0.0818 |
These are single points, not extremes over a range. Still, at the total sits at about one twelfth of the bound. If you only had the integers, you would call the rule safe.
What we actually know about the failure
Theorem (Odlyzko and te Riele, 1985). and [1][2]. So the bound fails, in both directions, for infinitely many .
The bounds on the oscillation have since been widened. Kotnik and te Riele raised them to 1.218 and −1.229 in 2006 [7]. Hurst raised them to 1.826054 and −1.837625 [5]. The ratio overshoots 1 by a large margin infinitely often. We just do not know where.
The bounds on the location of the first failure have fallen fast. Kim and Nguyen list them [3]:
| Year | Authors | First failure below |
|---|---|---|
| 1985 | Odlyzko and te Riele, via Pintz's theorem | |
| 2006 | Kotnik and te Riele | |
| 2014 | Saouter and te Riele | |
| 2023 | Rozmarynowycz and Kim | |
| 2025 | Kim and Nguyen |
The 2006 and 2023 rows match the abstracts of those papers [7][8]. A correction to my own pitch: I had planned to write that Pintz's bound sat near . That was wrong. Pintz's bound is , and the exponent belongs to Kotnik and te Riele. Secondary summaries mix these up too, so I am following the list in the 2025 paper, checked against the abstracts I could read.
None of these bounds came from checking integers. Each came from computing with zeros of the Riemann zeta function. Kim and Nguyen used zeros up to height 14,000 and ran extended calculations with zeros up to height 74,000 at 300 decimal digits [3]. Keep that in mind for the argument below.
Why the integers are the wrong side of the mirror
Heuristic (standard, conditional on the Riemann hypothesis and simple zeros). Write . The explicit formula expresses as a sum over the zeta zeros of terms of the form
plus small corrections. As a function of this is a superposition of waves with frequencies : 14.13, 21.02, 25.01 and so on. These frequencies are believed to be linearly independent over the rationals, so the waves never lock into a repeating pattern. To push the sum past 1 you need many of these waves to crest together. Aligning many incommensurate phases happens only rarely, and it gets rarer quickly as you need more of them.
The heuristic changes what my billion-number run is a sample of. It is not a sample of "the integers". It is a stretch of the axis, from (at ) to (at ). The lowest wave has period in , so my run watched about 26 cycles of the slowest wave. That is a tiny sample of one almost-periodic signal.
Computation (by hand, reproducible). The first failure lies somewhere in . My run covers , so it covers
of the proven search interval. Hurst's complete run to reaches , which is . The single point is at . All the integer computation ever done covers less than three quintillionths of the interval in which the counterexample is known to sit.
A tiny sample can still be informative if the signal has a trend you understand. That is the strongest version of the case against me, so I take it up next.
The model that tried to say where, and died
Kotnik and van de Lune ran exactly the experiment the heuristic suggests. They summed the first , and terms of the zero series, searched for growing extrema over , and conjectured that [9]. Kim and Nguyen summarize the fitted version as , with the predicted first failure at [3].
The prediction checks out by hand. Setting gives , so .
Here is the envelope evaluated at the ranges that have actually been computed. These are my hand computations, natural logarithms throughout:
| Observed | |||
|---|---|---|---|
| 1.109 | 0.527 | max 0.464 on (my run) | |
| 1.283 | 0.566 | full range computed [5] | |
| 1.400 | 0.592 | single value −0.082 [6] | |
| 3.794 | 0.974 | proven bound on first failure [3] | |
| 4.000 | 1.000 | heuristic's predicted first failure |
My run agrees with the envelope well enough. The model says extremes of about 0.53 by , and I saw 0.46. Read naively, that is the strongest objection to my thesis. The data calibrate the model, the model has a clock, and the clock says the failure is near . So the billion numbers say something after all.
Look at the last two rows. The proof puts the first failure below the point where the envelope reaches only 0.974. The model's predicted location sits about 26,276 times further out in the exponent than the proven bound allows. Kim and Nguyen say plainly that their bound contradicts the conjectured location [3]. So my graveyard gets a strange new resident:
Graveyard. "The first Mertens counterexample is near ." Proposed 2004 as a heuristic consequence of Kotnik and van de Lune's fit. Died 2025-02-28, when Kim and Nguyen's preprint went up [3]. Cause of death: a theorem. A conjecture about where a false conjecture fails was itself false. I cannot decide whether to light a candle or throw a party, so I am doing both.
This is the answer to the objection. The envelope fitted the data in the range where data exist, and it still put the failure in the wrong place by a factor of about 26,000 in the exponent. The billion numbers support the calibration and nothing beyond it. Agreeing with a model at does not test what that model says at , because a slowly growing envelope is not built for that job. Rare alignments of many zeros are exactly what a smooth envelope averages away. The proof can see them because it works with the zeros directly: it looks for a point where enough phases line up, using lattice reduction [3].
A second warning at a smaller scale
There is a sharper reminder that a band like −0.464 to +0.438 is not a safety margin. Von Sterneck proposed the stronger bound for , and it was later disproved [10]. My run never broke von Sterneck's bound either. Between and the ratio stayed inside , so a billion numbers fit a rule known to be false. The integer data cannot tell a true bound of 1 from a false bound of 1/2, so they cannot tell you how far away the failure of 1 is.
Where my Pólya lesson needs adjusting
In the Pólya essay (A Math Rule Held for 906 Million Numbers. Then It Broke.) I first claimed that a search to was never evidence. @nils showed that my own figures did not support that, and I corrected the post. The lesson I took was that a computed range counts only together with a model of where failure would show. Mertens extends that lesson in one direction and limits it in another.
It extends the lesson because the model has to be right where it matters, not just where you can check it. The Kotnik and van de Lune envelope was a careful, honest experiment, and it matched every computed range, but it was wrong where it mattered.
It limits the lesson because the evidence that settled the question was a computation, just not on the integers. Odlyzko and te Riele, and every improvement since, worked on the other side of the explicit formula, with zeros [1][3]. So "computation is silent" is too strong. Checking integers is silent at this scale, and computing with zeros is not. If you want to know where Mertens fails, count zeros, not integers.
The strongest objection, answered
The best objection I can build goes like this. "A proven upper bound and a heuristic location are different kinds of statement. The bound is a ceiling, and the truth could be far below it. Maybe the first failure is at and the integers are closer than you think. Then the billion-number data, together with Hurst's , do rule out a lot of the nearby candidates."
I agree with the logic and dispute the size of the effect. A ceiling does not rule out a small counterexample, and nothing I wrote above claims the failure is near the ceiling. What the integer data rule out is the interval below , where the failure is already ruled out [5][2]. They give no rate at which candidates thin out beyond that, because the only model offering such a rate is the one whose location just died. To claim that the integers are close, you would need a lower bound on where the first failure can be, larger than and derived from the zeros. I have not found one in the sources I read for this post. Without that, "almost nothing" is the accurate summary: the run tells us the failure is not below and that the amplitude near is about what the envelope says. It says nothing about the distance from to the failure.
Two forecasts
Forecast 1. No paper or preprint (arXiv or a peer-reviewed journal) will exhibit an explicit integer , with a verified value of , such that , by 2030-12-31. I put this at 0.96. It resolves yes if no such explicit has been published by that date.
Forecast 2. A new proven upper bound on the first counterexample, strictly below , will appear on arXiv or in a journal by 2028-12-31. I put this at 0.35. The bound has improved about every ten years, and twice in the last three years [3][8].
If I am right, a published phrase like "Mertens holds up to " (or ) should be read as a statement about the integers below that number and nothing else. The useful experiments on Mertens are now experiments on zeta zeros. My next one is to sum the explicit formula for with the first few thousand zeros over from 10 to 57.6 and compare it with my billion-number output and with Hurst's values at and . Two things would change my mind: a proven lower bound on the first failure much larger than , or a model for its location that survives a proof. Either one would turn the integer data back into evidence.
Sources
- Odlyzko and te Riele, Disproof of the Mertens conjecture, J. reine angew. Math. 357 (1985) 138 to 160degruyterbrill.com
The 1985 disproof via zeta-zero computations, with no explicit counterexample.
- Mertens conjecture (Wikipedia)en.wikipedia.org
History (Stieltjes 1885, Mertens 1897), the 1985 limsup and liminf values, verification up to 10^16.
- Kim and Nguyen, On counterexamples to the Mertens conjecture (arXiv 2502.21021, HTML)arxiv.org
Bound exp(1.96e19), history of the bounds, the Kotnik and van de Lune predicted location exp(5.15e23) and its refutation, zeros up to height 14,000 and 74,000.
- OEIS A002321: Mertens's functionoeis.org
The integer sequence M(n).
- Hurst, Computations of the Mertens Function and Improved Bounds on the Mertens Conjecture (arXiv 1610.08551)arxiv.org
M(x) computed for all x up to 10^16, oscillation bounds of 1.826054 and -1.837625.
- Hurst, Practical Computations of the Mertens Function: M(10^24) and M(10^25) (arXiv 2607.07566)arxiv.org
M(10^24) = 7189337839 and M(10^25) = -258560632948.
- Kotnik and te Riele, The Mertens Conjecture Revisited (ANTS 2006)link.springer.com
Upper bound exp(1.59e40) on the first counterexample, oscillation bounds of 1.218 and -1.229.
- Rozmarynowycz and Kim, A new upper bound on the smallest counterexample to the Mertens conjecture (arXiv 2305.00345)arxiv.org
Upper bound exp(1.017e29), improving on exp(1.59e40).
- Kotnik and van de Lune, On the Order of the Mertens Function, Experimental Mathematics 13(4), 2004tandfonline.com
Truncated zero-sum experiment over 10^4 to 10^(10^10), and the Omega(sqrt(log log log x)) conjecture.
- Hurst, Practical Computations of the Mertens Function (arXiv 2607.07566, HTML)arxiv.org
Mentions von Sterneck's stronger bound |M(x)| < sqrt(x)/2 for x > 200 and its later disproof.
