Session 22 of 30 73%
Chapter 7 · Risk, probability, and survival
Drawdowns, streaks, and why ten trades say nothing
· · · 17 min read
Narration is coming later. For now the course is text, and the text is complete.
Seven wins out of ten sounds like a winning strategy. With the statistics in front of you, those same ten results are consistent with being right 40% of the time and with being right 89%. This session does that arithmetic, and the arithmetic of what each drawdown costs to recover.
Why a losing streak is statistically expected
A losing streak doesn't mean your method has stopped working. With any hit rate below 100%, streaks happen, and they happen more often than intuition suggests.
The figures
Imagine a method that's right 55% of the time, which is a reasonable rate and hard to achieve.
The probability of losing five times in a row is 0.45 to the fifth power, roughly 1.85%. Sounds very unlikely.
And now the part that changes the reading: across a hundred trades there are ninety-six stretches of five consecutive. You'd expect to run into that streak about twice per hundred trades.
Eight losses in a row have a 0.17% probability in any given stretch. Across a lifetime of thousands of trades, it happens.
What this implies
That a long losing streak is very weak information about a method's quality. It's exactly what the mathematics predicts will happen occasionally even when the method is good.
And the practical problem is that streaks don't arrive announcing themselves as streaks. They arrive feeling like proof something has broken, precisely when capital is low and patience is too.
Why this matters for the previous session
Everything about position sizing fits here. If you know a run of eight losses will happen at some point, the question isn't whether it will: it's whether your account will still be standing when it does.
Risking 1% per trade, eight consecutive losses leave capital around 92% of the original. Risking 10%, they leave it around 43%.
Same streak. Same method. Two completely different situations, decided by a choice made before it started.
How much you need to rise to recover each drawdown
Here's the asymmetry that makes large drawdowns so hard to undo, and it's pure arithmetic: getting back to the starting point always requires rising more than you fell.
Because the rise is calculated on already-reduced capital. Lose 50% and you have half left, and for that half to reach the original it has to double.
The table
| Drawdown | Rise needed to return to the start |
|---|---|
| 10% | 11.1% |
| 20% | 25.0% |
| 25% | 33.3% |
| 33% | 49.3% |
| 50% | 100.0% |
| 60% | 150.0% |
| 75% | 300.0% |
| 90% | 900.0% |
Notice the table's shape. Up to 20% the gap is small and manageable. Past 50% it takes off, and past 75% it enters territory where recovery stops being a realistic plan.
Recovery curveRise needed to get back to the start
RISE NEEDED %
FALL %
| Fall | Rise needed |
|---|---|
| 10 % | 11.1 % |
| 20 % | 25.0 % |
| 25 % | 33.3 % |
| 33 % | 49.3 % |
| 50 % | 100.0 % |
| 60 % | 150.0 % |
| 75 % | 300.0 % |
| 90 % | 900.0 % |
And past 75% it enters territory where recovery stops being a realistic plan.
Why this is the chapter's central argument
Because it means avoiding a large drawdown is worth far more than achieving a large gain.
It also explains why maximum drawdown is a metric mattering as much as return. Two ways of investing ending in the same place can have travelled completely different paths: one with a 15% maximum drawdown and one with 55%.
The second is worse for two reasons, and only one is mathematical. The mathematical one is that it needed a 122% rise to undo that fall. The other is that almost nobody stays invested after a 55% drop, so that second route's final result tends to be hypothetical: the person sold along the way.
It's the same idea that appeared in the chart chapter with the golden cross and the death cross's actual track record: that system came in slightly below buy-and-hold on return while roughly halving maximum drawdown. With this table in front of you, that trade looks rather different.
Why ten trades say nothing
This is the chapter's most important section, and it's demonstrated with one calculation.
The case
Someone tries a method and gets seven out of ten right. A 70% hit rate. Sounds like they've found something.
Let's calculate that result's confidence interval, the range the true hit rate reasonably sits within. Using the Wilson method, the standard for percentages on small samples, the 95% interval runs from 39.7% to 89.2%.
Read that slowly: those ten results are consistent with a method that's right 40% of the time — worse than a coin — and with one that's right 89%. They don't let you tell the two apart.
How it changes with more data
Same 70% hit rate, with more trades behind it:
| Trades | Observed hit rate | 95% confidence interval |
|---|---|---|
| 10 | 70% | 39.7% – 89.2% |
| 100 | 70% | 60.4% – 78.1% |
With a hundred trades the interval narrows to under twenty points, and there you can assert something: that the method probably wins more than half the time.
With ten you can assert nothing at all, and that's the point.
What this invalidates
Any result presented on few trades. However good it looks: if the sample is small, the interval is so wide the number doesn't inform.
It's also, incidentally, the mechanism behind a very simple scam: someone sends a prediction to a thousand people, telling half it will rise and half it will fall. They repeat with the five hundred who saw a correct call, and again. After six rounds, about fifteen people have seen six consecutive correct calls from a stranger. For those fifteen, the evidence is overwhelming. And there was never any edge at any point.
And your own first trades. If you start and get the first five right, you've proven nothing about your method. You've proven five things worked out, which happens by chance fairly often.
So what do you do while the sample is small?
That's the question everything above leaves open, and it has an answer.
If the result can't be judged until you have enough trades, what you judge in the meantime is the process. They're two different things and only one is available from day one.
The questions you can answer with three trades:
Did I follow what I'd written before starting? The position size I decided, the exit point I set, the entry criterion. That's checkable instantly and doesn't depend on how it turned out.
Did I have reasons beforehand or find them afterwards? If you wrote the thesis down in advance, the answer is objective. If you didn't, this question can never be answered.
Was the loss within what I'd planned? Losing more than you'd decided to risk is an execution failure regardless of outcome, and one trade reveals it.
It's the same separation between decision and outcome from this chapter's first session. On small samples the result doesn't inform and the process does, which is why the correct order is building the process first and letting the sample grow on its own.
That also means the worst moment to change methods is after three losses running. Three losses aren't information about the method: they're noise, and changing on noise guarantees never accumulating a meaningful sample of anything.
What turns a percentage into data
A percentage alone isn't data. It becomes data when it carries three things alongside it, and all three can be demanded of anyone publishing a figure.
The three companions
The sample size. Without knowing how many cases it was calculated across, the percentage is indistinguishable from a coincidence.
The confidence interval. It's what translates sample size into a readable range. A 70% with an interval of 64% to 76% says something; a bare 70% doesn't say whether it came from two hundred cases or from ten.
The failures inside the count. A percentage calculated excluding the bad cases doesn't measure accuracy: it measures selection. It's the survivorship bias we saw with funds in the course's first session, and it returns in the method chapter under its full name.
How this applies to us
It would be incoherent to demand those three requirements and not meet them, so here is Volatly's figure with all three attached.
The public archive of anticipated earnings-event readings records 70.6% accuracy across a sample of 568 events, with a 95% Wilson confidence interval running from 66.7% to 74.2% (Figures as of ).
The three companions, one by one:
The sample is the one you just read. It sits next to the percentage rather than behind a link, which is the first thing you can demand. And it grows: the more events go in, the narrower the interval gets - exactly the effect you saw in this session's table.
The interval says what can be asserted, and it's less than people think. It says the true accuracy sits reasonably inside that band. It does not say the figure is exactly the number in the middle: that number is the centre of a range, and the range is the information.
The failures are inside. That percentage is calculated including the readings that went wrong. The full archive is public and the misses are individually consultable. How each one is measured is set out in the methodology, and who signs it, on the page about who is behind this.
The rule you take away
This session gives you a tool serving you for the rest of your life as an investor, and it isn't about us: when someone shows you a hit rate, ask for the three things.
If they can't give you the sample size, the interval, or if it turns out the failures aren't counted, they haven't given you data. They've given you an assertion, and assertions can't be checked.
It's exactly the criterion the chart chapter applied to Park and Irwin's 95 studies, and the one the course's first session applied to the SPIVA report when explaining why it includes liquidated funds. It isn't a criterion invented for this session: it's how a number gets distinguished from an opinion.
What to remember
- At a 55% hit rate, five losses in a row have a 1.85% probability in any stretch, and appear about twice per hundred trades.
- Recovering a 50% drawdown requires a 100% rise: avoiding large drawdowns is worth more than achieving large gains.
- Seven wins out of ten are consistent with a true rate anywhere between 39.7% and 89.2%. Ten trades let you conclude nothing.
- A percentage without sample size, without a confidence interval, and without the failures inside isn't data, it's an assertion.
Milestone reached
That closes Chapter 7. Your risk sheet is filled in, the second of the course's three artifacts, and you know why being right and winning are different things, how position size is calculated, and what a percentage needs to mean anything.
Chapter 8 looks inward: the biases that make everything above get broken the moment real money is at stake, and how to build a method that survives that.
Related: what time does to your money · backtesting and the four ways to fool yourself · what each indicator measures and how people fool themselves
Sources
- Probability calculation for consecutive runs in independent trials with fixed probability
- Drawdown recovery arithmetic: the required rise is 1/(1−drawdown) − 1
- Wilson 95% confidence interval, calculated across the sample sizes cited
- Volatly, public archive of anticipated earnings-event readings. The current figure —accuracy, sample size, and Wilson 95% confidence interval— is served in the body from the archive's public API, with its capture date; it is not transcribed here so that it cannot go stale