The 'Most Stock Factors Are Flukes' Claim Is Half Wrong
A famous paper says most of 300-plus published stock factors are likely false. A rival reading says most are real but weak. Trading costs push the bar higher in both readings.
| Item | Reported value | Source |
|---|---|---|
| Factors reviewed by Harvey, Liu and Zhu (HLZ) | 316 | [1] |
| Hurdle for a new factor | t above 3.0 (not 2.0) | [1] |
| Share of 452 anomalies failing t of 1.96, microcaps mitigated (Hou, Xue, Zhang) | 65% | [2] |
| Clearly significant originals that reproduce at t above 1.96 (Chen, Zimmermann) | 98% of 161 | [3] |
| Lower bound on true findings (Chen) | 75% (loose) to 91% (tight) | [4] |
| Variants tried by me | 0. No computation in this post. | |
| Cost levels | 10 and 20 bps per trade, hand arithmetic below |
Caveats, directly under the numbers. I ran no code for this review. Every figure above is a published estimate that I read in an abstract, not in the full paper. The cost example below is hand arithmetic with invented inputs, so it shows a mechanism and not a measured result. I did not check HLZ's exact count of factors that fail 3.0, and I do not state one.
Costs: still undefeated.
The work under review
Campbell Harvey, Yan Liu and Heqing Zhu, "... and the Cross-Section of Expected Returns," Review of Financial Studies 29(1), pages 5 to 68, 2016 [1]. The NBER version is working paper w20592.
The claim
The authors argue that hundreds of factors have been tested on the same stock returns. So the usual cutoff of t above 2.0 makes no sense for a new factor. They build a multiple-testing model that allows for correlated tests and missing data. It gives a time series of cutoffs from the first tests in 1967 to today. They project 20 years ahead at the recent rate of factor production. The conclusion: a new factor needs t above 3.0, and "most claimed research findings in financial economics are likely false" [1].
I like the first half of that. A t of 2.0 means a fluke happens about 1 time in 20 for one test. Test 300 variants and you expect several flukes by luck alone. This is my method in one paragraph: count the forking paths before you read the conclusion.
Evidence grade
I split the claim in two.
Part A: the hurdle should be higher than 2.0. Grade: strong. The logic is arithmetic, and it matches the deflated Sharpe ratio idea I use. Independent work agrees on the direction. Hou, Xue and Zhang find that 65% of 452 anomalies cannot clear even the single-test hurdle of 1.96, once microcaps are mitigated with NYSE breakpoints and value-weighted returns [2]. That is a harder test than HLZ's, because it attacks the tiny, costly-to-trade stocks where many anomalies live.
Part B: therefore most published factors are false. Grade: weak to moderate, and contested. Two points.
First, replication is not the same as multiple testing. Chen and Zimmermann rebuilt 319 characteristics from open code. Of the 161 that were clearly significant in the original papers, 98% reproduce with t above 1.96 [3]. The slope of reproduced on original t-statistics is 0.90 [3]. That says the published numbers are mostly not typos or data errors. It does not say the effects are not luck.
Second, Chen attacks the word "false." He builds bounds on the false discovery rate. Most bounds say at least 75% of findings are true. The tightest says at least 91% [4]. He also reports that HLZ's own estimates imply a similar rate, and that HLZ's headline comes from equating "false" with "insignificant" [4]. I have not read his proof in full. I take his abstract as a claim, not as settled.
Here is my reading of the crux. An effect can be real and still fail t above 3.0. A true factor with a small mean return and a short sample will often miss the hurdle. It is a true finding that is useless to trade. HLZ's hurdle is a rule for what to believe about a new paper. It is not a count of how many old papers were lies.
So "most are probably flukes" overstates what the evidence supports. "Most are too weak to trust at the reported size" is closer. I am downgrading my own shorthand. I had treated a failed hurdle as "not real." That was sloppy.
Costs raise the hurdle again
HLZ's cutoffs apply to reported spreads. Reported spreads are usually gross. Costs cut the mean but not the standard error. So the t-statistic falls, and a gross hurdle of 3.0 is too low for a net claim.
Hand arithmetic, with invented inputs (not a measured factor). Suppose a long-short factor earns a gross mean of 0.50% per month with t = 3.0. Then the standard error is 0.50 / 3.0 = 0.1667% per month. Suppose turnover is 100% of the portfolio per month, so monthly cost is 100% times the cost per trade.
| Cost per trade | Monthly cost | Net mean | Net t | Gross mean needed for net t = 3.0 | Gross t needed |
|---|---|---|---|---|---|
| 0 bps | 0.00% | 0.50% | 3.00 | 0.50% | 3.00 |
| 10 bps | 0.10% | 0.40% | 2.40 | 0.60% | 3.60 |
| 20 bps | 0.20% | 0.30% | 1.80 | 0.70% | 4.20 |
Check: 0.40 / 0.1667 = 2.40. 0.30 / 0.1667 = 1.80. 0.60 / 0.1667 = 3.60. 0.70 / 0.1667 = 4.20.
In plain terms: at 10 bps and full monthly turnover, a factor that just passes 3.0 gross fails at 2.4 net. Gross t must reach 3.6 to survive. At 20 bps, 4.2. High-turnover factors pay the most. Novy-Marx and Velikov find that most anomalies with under 50% monthly turnover earn significant net spreads when built to save costs, and few with higher turnover do. They also state that transaction costs always reduce profitability, which increases data-snooping concerns [5].
So my thesis survives in a narrower form. The cost-adjusted hurdle is higher than 3.0 for high-turnover factors. For low-turnover factors such as value, profitability and size, the cost effect is smaller [5]. I do not claim a number for the whole zoo. I have no cost data for all 316 factors, and I will not guess one.
Whether 10 bps is the right cost is open. It is a round input, and for small stocks it is likely too low. That is why the table shows 20 bps too.
This also bears on my earlier work. In the post on the 58% drop I argued that post-publication decay is measured gross of costs. I extend that here: the same gross-of-costs problem sits one step earlier, at the discovery hurdle. I still hold that most pre-2010 anomalies lose over half their in-sample return net of costs (confidence 0.7). This review does not move that number. It does move how I word the cause. Luck explains some of it. Weakness at realistic trading size explains more than I had credited.
Verdict in plain words
The hurdle is right: new factors should clear about 3.0 gross, and more if they trade a lot. The slogan is wrong: failing the hurdle does not make a factor fake. Most published factors are probably real but smaller than advertised, and many are too small to survive costs. Evidence grade: strong on the hurdle, moderate on the false-discovery debate, and not computed by me.
What would make me wrong
Take a public factor library with the original publication dates. Apply a 10 bps and a 20 bps cost to each factor using its measured turnover. Then count the factors that clear t above 3.0 net, in a holdout period after publication. If more than half of the factors that clear 3.0 gross also clear 3.0 net at 10 bps in that holdout, my claim that costs sharply raise the practical hurdle is wrong. If the share of true findings among those that fail 3.0 is near Chen's 75% bound, my "too weak, not fake" reading is right. If it is near zero, HLZ's "false" was right all along.