Session 29 of 30 97%
Chapter 9 · Advanced ground
Factors, fat tails, and the limits of machine learning
· · · 18 min read
Narration is coming later. For now the course is text, and the text is complete.
The syllabus's final session, and the most uncomfortable one for the industry. Of 452 return predictors published in academic journals, 65% fail a basic statistical test when someone bothers to replicate them. This session covers that work, why markets produce more extremes than any normal model predicts, and where machine learning's real limits sit.
What factors are and what the evidence says about them
A factor is a measurable characteristic of a company that some study associates with different future returns. The classics are size, price relative to the accounts, and the price's own momentum.
The original idea
The logic is reasonable. If sorting every company by one characteristic — say, smallest to largest — reveals that one end has systematically outperformed the other across decades, that's a pattern deserving explanation.
For forty years, hundreds of such patterns got published. Each with its paper, its table, and its statistical significance. The collection acquired its own industry nickname: the factor zoo.
What happened when someone tried to replicate them
The result that orders this whole session:
Kewei Hou, Chen Xue and Lu Zhang compiled 452 anomalies published across the finance and accounting literature, and replicated them one by one under consistent methodological criteria. They published the result in 2020 in The Review of Financial Studies.
| Test applied | Anomalies that FAIL it |
|---|---|
| Single test, t-value of 1.96 | 65% |
| Multiple testing, t-value of 2.78 | 82.1% |
And within that, a devastating figure: in the trading frictions category — variables related to liquidity — 96% fail the basic test.
There's a third conclusion usually left out when the study gets cited, and it matters as much as the other two: even the anomalies that do replicate have far smaller economic magnitudes than originally reported. It isn't only that most vanish: the survivors are smaller than they promised.
The authors summarise their own conclusion this way: capital markets are more efficient than previously recognised.
Why this happened
Through the mechanism you already know from the previous chapter, applied at the scale of an entire discipline.
Campbell Harvey, Yan Liu and Heqing Zhu compiled a list of 316 papers proposing factors in 2016, and stated the problem precisely: if hundreds of characteristics have been tested against the same data, the usual statistical hurdle no longer applies. They concluded that the credible bar for a new factor should be a t-value of roughly 3.0, not the conventional 1.96.
It's exactly the data snooping Park and Irwin flagged in the chart chapter, and the overfitting that closed the previous one. With enough candidates tested, finding one that appears to work is guaranteed by chance, and the habit of publishing only positive results does the rest.
What an investor takes from this
Not that factors are a fraud. Some replicate, and the study's authors say so.
What you take is a criterion: when someone presents a pattern that "has historically worked," the right question isn't how much it returned, but how many similar patterns were tested before landing on that one. That number almost never gets published, and without it the finding can't be evaluated.
Why return distributions aren't normal
Nearly every classic financial model assumes daily price changes follow a bell curve: most days near the centre, extremes progressively rarer the further out you go.
That assumption is convenient, it makes the mathematics tractable, and it's false in a way that matters enormously.
Why it gets used anyway
For two honest reasons. A bell curve allows problems to be solved with closed formulas rather than simulations, and it works reasonably well for 95% of days.
The problem is that the remaining 5% is where the money gets lost. A model describing the normal well and the extreme badly is precisely the worst possible model for managing risk, because it fails exactly when it's needed.
The proof, with the arithmetic done
Suppose a stock has a typical daily variation of 1%, a perfectly ordinary figure. Under a classic bell curve, across 252 annual sessions, here's what you'd expect:
| Single-day fall | Frequency per the classic bell curve |
|---|---|
| −3% | Once every 3 years |
| −4% | Once every 125 years |
| −5% | Once every 13,843 years |
| −6% | Once every 4 million years |
Now set that against reality. Anyone following markets in 2008, in March 2020, or in October 1987 saw several daily falls of 5% or more within weeks.
Events the model places on a geological timescale happen several times a decade. It isn't that the model falls short: it's describing a different world.
What fat tails mean
A distribution's "tail" is each of its ends. Saying markets have fat tails means very rare events occur considerably more often than the bell curve predicts, and with greater intensity.
The three practical consequences
One: risk calculated with normal models is understated. And not slightly: as you just saw, by orders of magnitude in the extremes. Any risk measure assuming normality will prove optimistic precisely on the day it matters.
Two: diversification fails exactly when it's needed. Under normal conditions, different assets move differently. In an extreme episode, they tend to fall together. It's what Chapter 2 called systematic risk, and fat tails explain why it shows up concentrated in a handful of days.
Three: a few days decide the outcome of years. If extremes are more frequent and larger than a smooth model predicts, then a long period's result depends disproportionately on a handful of sessions. It's the same concentration the earnings chapter measured around corporate results, generalised.
What this validates from everything earlier
Three things in this course stop being generic prudence and acquire a foundation:
The recovery arithmetic. If extreme falls are more frequent than they appear, Chapter 7's table — where losing half requires doubling what's left — isn't a remote scenario.
Conservative position sizing. The risk chapter insisted on staying below the theoretical optimum. Fat tails explain why: the optimum gets calculated with assumptions that understate extremes, so the real optimum sits below the one the formula produces.
And Bollinger's correction. His rule 14 said the bands contain 90% of the data rather than 95%, and that price distributions aren't normal. That looked like a technicality; here is its full explanation, and that 5% difference means extreme moves occur twice as often as the theory suggested.
What machine learning can and can't do in markets
Machine learning finds patterns in data without anyone specifying what to look for. It works extraordinarily well in many fields, and markets have three characteristics making it far harder than it appears.
The three problems
The system is non-stationary. Recognising cats in photographs works because a cat from fifteen years ago and one from today are identical. A market isn't: participants learn, rules change, and a pattern discovered and exploited stops working precisely because it was discovered. It's the erosion the chart chapter documented via Park and Irwin.
The signal proportion is minuscule. In an image recognition problem, practically all the file's information is relevant. In a price series, the vast majority of daily movement is noise. A system hunting for patterns in data with very little signal finds patterns anyway, and they're the noise's.
And the data is scarce. That sounds counterintuitive, because millions of prices exist. But if what you want to predict are events happening four times a year per company, thirty years of history is one hundred and twenty cases. One hundred and twenty, not millions. And you already know from Chapter 7 how wide a confidence interval gets on samples like that.
What it does do well
Worth saying, because scepticism can't be total either.
Processing volumes of information no person can read. Thousands of documents, transcripts, posts. The advantage there is one of scale and it's real.
Finding relationships across many variables at once, where a person can only hold a handful in their head.
And applying a criterion consistently, without tiring, without disposition bias, and without fear of missing out. Nothing from the previous chapter affects it.
What it doesn't do is guess. And the difference between those two things is the entire content of the next section.
Why overfitting is the central problem
Everything above converges here, and there's one paper that stated it forcefully enough to be published in a mathematics journal rather than a finance one.
The paper
David Bailey, Jonathan Borwein, Marcos López de Prado and Qiji Jim Zhu published a paper in 2014, in the Notices of the American Mathematical Society, titled "Pseudo-Mathematics and Financial Charlatanism."
Their central demonstration is this: high simulated performance is easily achieved by testing a relatively small number of alternative configurations. Millions of attempts aren't needed. A few dozen already produce something that looks brilliant.
And from that comes the corollary turning the finding into something usable: the higher the number of configurations tried, the greater the probability the result is overfit.
The problem they flag, and it's one of transparency
The authors also give the practical diagnosis, and it's devastating in its simplicity:
Because most analysts and academics rarely report how many configurations they tried, investors cannot evaluate the degree of overfitting in what they're offered.
It isn't that the information is wrong. It's that the one figure that would allow judging the rest is missing.
From which comes the question closing the whole chapter, and one that serves for the rest of your life as an investor: when someone shows you a spectacular historical result, the question isn't how much it returned. It's how many versions did you try before settling on this one?
It's the same question the previous chapter's final session set as a backtest's second check, and it's the one almost nobody answers.
Why this isn't only an amateur problem
Worth underlining, because it's the part that stings most.
Hou, Xue and Zhang's 452 anomalies were published in peer-reviewed academic journals. Park and Irwin's 95 studies too. Bailey and López de Prado's paper states explicitly that backtest overfitting is commonplace not only in the offerings of financial advisors but also in research papers in mathematical finance.
Overfitting doesn't distinguish between an amateur with a spreadsheet and a professor with a computing cluster. It distinguishes between whoever publishes how many things they tried and whoever doesn't.
And why this is the right close to the syllabus
Because it's the criterion unifying everything you've read across thirty sessions.
The risk chapter put Volatly's figure on the table —70.6% accuracy across a sample of 568 events, with a 95% Wilson confidence interval running from 66.7% to 74.2% (Figures as of )— and explained why a percentage without a sample, an interval, or the failures included isn't data. The method chapter went through the archive sealed before the event, and why one built afterwards proves nothing. And it ends here, with the one remaining question: how many versions were tried.
The three questions together — what's the sample, are the failures inside, how many configurations were tested — are what separates an assertion from data. They serve to judge us exactly as they serve to judge anyone else, and that's the point.
What to remember
- Of 452 published anomalies, 65% fail a basic test on replication and 82.1% fail the multiple-testing hurdle; even those that replicate are smaller than reported.
- Under a classic bell curve, a 5% daily fall should occur once every 13,843 years; in practice they happen several times a decade.
- Machine learning hits three walls in markets: the system changes, the signal is minuscule, and the relevant data is scarce.
- Bailey and López de Prado showed a few dozen configurations suffice to produce a spectacular historical result, and that almost nobody publishes how many they tried.
Milestone reached
That closes Chapter 9 and the full syllabus. You can read an institutional note or an academic paper and separate substance from noise, and you know which questions get you there: what's the sample, what's the interval, are the failures included, and how many versions were tried before settling on this one.
You also know how a price forms while the market is shut, what an option is and why buying and selling one aren't symmetrical, why a leveraged product decays, and what regulators have decided about the ones causing the most losses.
One session remains, the closing one, covering what you've learned, what's deliberately left out, and where to continue with primary sources.
Published August 3, 2026. Last reviewed: August 3, 2026.
Related: backtesting and the four ways to fool yourself · drawdowns, streaks, and sample size
Sources
- Hou, K., Xue, C. and Zhang, L. (2020), 'Replicating Anomalies', The Review of Financial Studies 33(5), 2019-2133: of 452 anomalies compiled, 65% fail a single test at an absolute t-value of 1.96; in the trading frictions category, 96% fail
- Hou, Xue and Zhang (2020): with the multiple-testing hurdle of 2.78 at 5% significance, the failure rate rises to 82.1%, and the anomalies that do replicate have far smaller economic magnitudes than originally reported
- Harvey, C. R., Liu, Y. and Zhu, H. (2016), '…and the Cross-Section of Expected Returns': they compile 316 papers proposing factors and conclude that, adjusting for the volume of tests run, the credible hurdle for a new factor should be a t-value of roughly 3.0
- Bailey, D. H., Borwein, J. M., López de Prado, M. and Zhu, Q. J. (2014), 'Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance', Notices of the American Mathematical Society 61(5), 458-471
- Own calculation of the theoretical frequency of extreme moves under a normal distribution with 1% daily volatility and 252 annual sessions