Part 10
Risk, Statistics, and Returns
On this page
10.1 How much money you need to trade
Before the machinery of this part, measuring an edge honestly, sizing it so a bad run cannot end you, combining streams into a book that survives, there is a plainer question that comes first and almost nobody asks. Is trading worth your time at all, and if it is, how much capital does it actually take to matter? The honest answers are less flattering than the industry that sells courses and signals would like, but they are not discouraging ones. Making money trading is entirely possible, and so, for a few, is turning a small account into a large one. The goal here is not to talk you out of any of it. It is to set expectations that match reality, so that the version of trading you actually run is one you can stick with instead of one that quietly costs you an account before you learn better.
10.1.1 The benchmark you actually have to beat
Trading is not free even when you win. Every hour you spend on it is an hour not spent earning or living, and every dollar you put at risk could have sat in a broad index instead. So the real benchmark is not zero, and it is not "did I make money." It is whether you beat what you would have earned doing nothing, after the time it cost you. For a stock trader that benchmark is the S&P 500. If you cannot beat SPY over a full cycle, with dividends, you have spent your evenings and your risk tolerance to underperform a fund that charges almost nothing and takes no effort. For someone trading higher-beta or crypto markets the bar is higher, not lower, because the passive alternative is higher too: the honest yardstick there is something like the Nasdaq or Bitcoin, whichever matches the risk you are actually taking. Measuring yourself against cash, or against your entry price, or against nothing, is how people convince themselves a losing use of their time is a business.
And the time genuinely counts. A strategy that beats the index by a few points but eats three hours a day is a worse deal than it looks, because those hours have a value and the comparison should include it. Before anything else, decide what benchmark your risk deserves and how many hours a day you are willing to feed the machine, because those two numbers decide whether any of the rest is worth doing.
10.1.2 What good actually looks like
Here is the number the marketing will not give you. A trader who compounds 20 to 40 percent a year, with tight risk rules, sized so no single loss threatens the account, and without gambling, is doing extremely well. Not "getting started" well. Genuinely, durably, top-tier well. The people promising more, consistently, are either gambling with position sizes that will eventually take them to zero, or selling you something. The reason is the through-line of this whole part: returns are compensation for risk, and you cannot manufacture a high compounding rate out of a modest edge without either taking more risk or levering up, both of which raise the odds that a normal bad stretch ends you. A 30 percent year at a Sharpe near one is a real achievement. A 200 percent year is almost always a large bet that happened to win, and the same process produces the minus 90 percent year that never makes it into the screenshot.
Which is why the risk-adjusted view this part keeps returning to matters more than the headline. Two traders both up 25 percent are not equal if one did it smoothly and the other rode a 60 percent drawdown to get there. The smooth one has something repeatable. The other one has a story and a coin that landed heads.
10.1.3 Small accounts can grow, and here is what it takes
None of this means a small account cannot become a large one. It can, and there are real examples. The most famous is Larry Williams, who in the 1987 World Cup Championship of Futures Trading turned ten thousand dollars into over 1.1 million in a single year, a return north of 11,000 percent, trading real money. It happened. The point is not that it is impossible. The point is what it actually took. Williams had a genuine, tested edge, and he sized it with an aggressive version of the Kelly criterion this part comes to later, betting a large fraction of the account on each position. That sizing is what turns a good edge into a parabolic return, and it is also what makes the ride nearly unsurvivable: the same run carried gut-wrenching drawdowns that would have shaken out or wiped out almost anyone, and Williams himself has said the risk he took is not something to emulate or something he could reproduce on command. A parabolic account needs all of it at once, a robust edge that genuinely works, the discipline to hold the line through the drawdowns, sizing aggressive enough to compound fast, and a real dose of luck in the markets it happens to run into. Remove any one and the same approach blows the account up, which is exactly what it does to most of the people who copy the sizing without the edge. So keep the possibility honest in both directions. Growing a small account into a serious one is a documented outcome, not a fantasy. It is also rare, risky, and part luck, and building your plan around being the next Larry Williams means building it around a lottery you still have to be skilled to enter.
10.1.4 Side income versus full time
The realistic path for most people is the one nobody sells, because it is unglamorous. If you run simple systematic strategies, trend following, selling volatility on liquid ETFs, harvesting earnings volatility, the kinds of edges the strategies part laid out, the ones that take about an hour a day to check and execute, then a lower five-figure account is entirely defensible as a side income. It will not replace your salary next year. But compounded over many years at a sane rate, alongside a job that keeps you from ever being a forced seller, a small account run this way has real potential, and the low time cost means you are not sacrificing your main income to run it. This is the version of trading that actually works for most of the people who make it work at all: modest, systematic, part-time, patient.
Going full time is a different proposition, and the arithmetic is unforgiving. It is silly to think anyone replaces a real income with a fifty thousand dollar account. Thirty percent on fifty thousand is fifteen thousand dollars in a very good year, before taxes, before the losing years, and you cannot spend the account you are supposed to be compounding. The capital required to live on trading is a function of your expenses, and there is no way to shortcut it: you need an account large enough that a sane, survivable return covers your life with room for the bad years, which for most people is a multiple of what they imagine. I cannot put one number on it, because a twenty-two year old living cheaply and a forty year old with a mortgage and children are not in the same situation. But the exercise everyone skips is the one that matters: write down what you need to live on, divide by a return you could actually earn without gambling, and look honestly at the account size that falls out. It is almost always far larger than the hopeful version in your head.
10.1.5 Day trading and scalping
The shorter your horizon, the harsher this gets, and day trading and scalping are the harshest of all. They demand hours in front of the screen while the market is open, every day, which is precisely the arrangement that makes building any other source of income difficult. You are trading your entire working day for it. The gains can be larger in percentage terms, that part is true, but so can the losses, and the specific fantasy that carries people into it, turning ten thousand dollars into millions by scalping, is far rarer than the stories make it look. It happens to a small number of people, and those are the only stories that circulate, while the many who tried the same thing and did not make it are invisible. It is not impossible, but the base rate is brutal, and walking in expecting to be the exception is usually how the account gets spent. If you are going to commit your whole day to markets, go in with clear eyes about the odds and a plan that survives being wrong, rather than with a screenshot as a business model.
10.1.6 Market selection sets the ticket size
Which market you trade is not a detail; it often decides how much capital you need before your skill even enters the picture. The clearest case is futures. A single futures contract controls a large notional position, and when the exchange offers no smaller version, you cannot trade a fraction of it. Many contracts have micro versions now, which helps, but plenty do not, some ICE-listed contracts among them, and there the minimum position is simply large. A strategy that is perfectly sound can be untradeable in an account that cannot hold even one contract at a sane fraction of its equity, and forcing it anyway means running at a size where a normal move is an account-threatening loss. Position sizing, the whole subject of this part, assumes you can size to your risk. When the smallest tradeable unit is larger than your risk budget, the math breaks before you start.
The platform's Core Allocation sleeve is the cleanest illustration, and it needs no detail about how it works to make the point. It is an extremely simple, diversified allocation. Over roughly fifteen years it earned a Sharpe of about 1.13, a genuinely strong risk-adjusted result, with a maximum drawdown near 8 percent, and its compound return was only about 6.4 percent a year. That low absolute number is not a defect in the strategy. It is a direct consequence of what it trades: diversified ETFs, which are low-volatility instruments, so even a high-quality edge on them produces a smooth, shallow-drawdown, low-return stream. The upside is that it can be run in a small account, because ETFs are divisible down to a single share. A trader who wants more return from the same idea can express it in futures instead, and the return and the volatility both rise together, exactly as vol targeting later in this part explains. But the moment you do that, the contract sizes mean you need something on the order of a couple hundred thousand dollars to run it efficiently. Same idea, same edge, two completely different capital requirements, decided entirely by the instrument. The closing lesson of this part walks that futures version end to end as a full worked example.
Core Allocation Backtest
That tradeoff, smooth and small-account-friendly and low-return in ETFs, or higher-return and higher-volatility and capital-hungry in futures, is the market-selection decision in miniature, and you make some version of it every time you choose what to trade.
10.1.7 Trading other people's money
There is one way around the capital problem, and it deserves a clear-eyed mention rather than a recommendation. You can trade external capital: proprietary trading firms that fund you after an evaluation, or arrangements where outside investors back your track record. Done well, this decouples your returns from your own savings, and a genuinely skilled trader with a small account can access size they could never fund themselves. Done carelessly, it is a trap. Be very careful choosing who you trade for. Many funding programs are built so that the difficult part is not trading well but keeping the capital, with rules on drawdown, on holding periods, on how and when you can scale, that are designed to be tripped, so that the fees from failed evaluations, and not the profits of funded traders, are the actual business. Read the fine print as if it were written by a counterparty whose interests are opposed to yours, because often it is. External capital can be a legitimate route for a proven trader. It is not a shortcut around being a proven trader, and the programs that market it as one are the ones to walk away from.
10.1.8 The honest close
None of this is meant to talk you out of trading. It is meant to make you decide, with the numbers in front of you, whether the version of it you can actually run is worth the time and money it will cost. For a lot of people the answer is a modest, systematic, part-time book that compounds quietly for years, and that is a genuinely good answer. For a few it is a full-time career that took a large account and a long apprenticeship to reach. For some the honest answer is a smaller, slower, part-time version than they first pictured, and rightsizing the ambition early, before the tuition of a blown account, is itself a win. Everything else in this part is about doing it well. This lesson is about being honest that "well" has to clear a bar the index sets for free.
With expectations set honestly, the rest of this part builds the machinery that turns a modest, survivable edge into a real one: the statistics, the performance math, the sizing, the risk control. It starts where every number on the platform starts, with the handful of statistical ideas the rest of the part is built from.
10.2 The statistics you actually need
The strategy lessons kept handing you statistical numbers. The VRP screener flags a z-score of +2.4. The skew screener filters at plus or minus 2. Funding is three standard deviations above its average. The COT index sits at 87 on a 0 to 100 range. You can trade off these numbers by pattern matching, and plenty of people do: big z-score means stretched, stretched means fade or follow depending on the strategy. But pattern matching without understanding is how you end up shorting a "three sigma" funding rate that goes to six sigma, or trusting a correlation that was never real, or treating a once-a-month event like a once-a-decade one.
The rest of this part is the quantitative backbone, and this lesson is its foundation: the handful of statistical ideas everything else in Part 10 is built from. The list is short. You need the mean and its blind spots, variance and standard deviation, what a distribution is and why the normal one is both everywhere and wrong, fat tails, skewness, correlation, and z-scores. That is the full kit. No calculus, no proofs. You need each of these at the level of understanding what the number claims and where the claim breaks, because every dashboard on this platform, and every risk decision in the lessons ahead, uses this language.
Statistics in trading has one job: compressing a pile of past observations into a few numbers you can act on. Every compression throws information away. The skill is knowing what got thrown away, because that is where the surprises live.
10.2.1 The mean, and what it hides
The arithmetic mean is the sum of the observations divided by their count. Five daily returns of +4, +1, -2, +6, and -4 percent average out to +1 percent per day. Nothing hard about the computation. The traps are all in the interpretation.
The most obvious trap is sensitivity to outliers. The mean gives every observation equal weight, so one extreme value drags it hard. A coin that returned +2 percent on 19 days and +150 percent on one day has a mean daily return above 9 percent, and that number describes exactly none of the 20 days. The median, the middle value when you sort the observations, would say +2 percent, which describes 19 of them. When a distribution has big outliers, and trading data always does, mean and median split apart, and the gap between them is itself information: it tells you the average is being carried by a few extreme observations rather than by typical behavior. Crypto alt returns are the canonical case. Portfolios of small coins often show a decent mean return and a negative median return, which decodes to "most of these bled, a few went vertical, and the verticals carried the average." Whether that profile suits you depends on whether you can hold the bleeders long enough to catch a vertical, which is a sizing and psychology question, not a statistics one. But the statistics tell you the question exists.
A bigger trap: the arithmetic mean of returns isn't what your account compounds at. Two years, +50 percent then -50 percent. Arithmetic mean: zero. Your account: 100 goes to 150 goes to 75, down 25 percent. The number that describes what actually happened to your money is the geometric mean, the per-period growth rate that compounds to the same endpoint. Here it's sqrt(1.5 * 0.5) - 1, about -13.4 percent per year. In plain terms: multiply the growth factors together, take the appropriate root, subtract one. The geometric mean is always at or below the arithmetic mean, and the gap widens with volatility. That gap, volatility drag, is why two strategies with identical average returns can leave you with very different account balances, and it gets its full treatment in the Kelly lesson later in this part. For now the rule is simple: when someone quotes an average return, ask whether it's the average of the periods or the rate the money actually grew at, because sufficiently volatile strategies can have a positive arithmetic mean and still grind an account to nothing.
And a mean without a spread is close to meaningless. "This strategy averages +0.3 percent per trade" tells you nothing until you know whether the trades cluster near +0.3 or range from -15 to +20. That is the next tool.
10.2.2 Variance and standard deviation
Variance measures how spread out observations are around their mean. The recipe: compute each observation's deviation from the mean, square the deviations, average the squares. Standard deviation is the square root of variance, which brings the number back into the same units as the data.
Run it once by hand with the five returns from before: +4, +1, -2, +6, -4 percent, mean +1. Deviations from the mean: +3, 0, -3, +5, -5. Squared: 9, 0, 9, 25, 25. Sum: 68. Divide by five: 13.6. Square root: about 3.7 percent. So this series has a mean of +1 percent and a standard deviation around 3.7 percent, and those two numbers together sketch the behavior: typically up a little, routinely swinging several percent either way.
Why square the deviations instead of just averaging their absolute size? Partly mathematical convenience: variances of independent things add, which makes portfolio math and time scaling work (the sqrt(252) annualization from the realized volatility lesson exists because variances add across days, so standard deviations scale with the square root of time). The reason that matters more: squaring makes large deviations count for far more than their share. A single 10 percent day contributes as much variance as one hundred 1 percent days. Standard deviation is therefore not a measure of the typical day; it weights the biggest moves in the sample most heavily. You saw this in the realized vol lesson, where one crash day carried a 20-day vol reading from 8 to 29. Same arithmetic, and it applies to every standard deviation on this site.
One technical footnote you'll meet in every spreadsheet: dividing the squared deviations by n versus n minus 1. Dividing by n minus 1 (the sample variance) corrects for the fact that you estimated the mean from the same data, which makes the deviations slightly too small on average. With our five observations, 68 divided by 4 gives 17, standard deviation about 4.1 instead of 3.7. With five data points the difference is visible; with 250 it's negligible. Know it exists so a mismatch between your number and someone else's doesn't send you hunting for a bug that isn't there, and then stop worrying about it. The estimation error from having a finite sample dwarfs the n versus n minus 1 choice at any size where the choice is interesting.
Standard deviation gives you a ruler. Raw moves are meaningless without context: a 2 percent day is a nothing day for a small-cap biotech and a five-alarm event for a currency future. Dividing moves by the instrument's own standard deviation converts everything into the same unit, "how unusual is this for this thing," and that conversion is the basis of nearly every normalized number on this platform. The z-score section below makes it explicit. The next few lessons build on the same ruler: risk measured in standard deviations rather than dollars, positions sized so that each contributes equal standard deviation, portfolios scaled to a target standard deviation. One concept, reused throughout.
Same Average, Different Spread
10.2.3 Distributions: the full picture the summary numbers compress
Mean and standard deviation are a two-number summary of something richer: the distribution, the full account of which outcomes occur and how often. The full picture of a return series is its histogram: chop the range of daily returns into bins, count the days in each bin, draw the bars. Do this for a few years of any liquid instrument and you get a recognizable shape: a tall pile near zero, shoulders falling away on both sides, and a scatter of lonely bars far out in either direction.
Daily Returns vs the Normal
The normal distribution, the bell curve, is the default model for that shape, and there is a reason. When an outcome is the sum of many small, independent influences, the distribution of that sum tends toward the bell curve regardless of what the individual influences look like. That is the central limit theorem, and it is why the normal shows up in heights, measurement errors, and lots of natural data. A day's return looks like it qualifies: thousands of trades, many participants, no single dominant influence. So the normal became the working assumption of most of finance, and it sits inside the Black-Scholes machinery you met in the options lessons, inside standard risk models, and implicitly inside every z-score.
The normal distribution is fully described by exactly two numbers, its mean and its standard deviation. That's its appeal: if returns were normal, the two-number compression would lose nothing. Everything about frequency of every size of move would follow from the pair. The following table is what the normal distribution promises, expressed in trading time, assuming roughly 252 trading days per year.
| Move size | Share of days beyond it (either direction) | Expected frequency |
|---|---|---|
| 1 standard deviation | about 32 percent | roughly one day in three |
| 2 standard deviations | about 4.6 percent | roughly one day in 22, about monthly |
| 3 standard deviations | about 0.27 percent | roughly one day in 370, every year and a half |
| 4 standard deviations | about 1 in 15,800 | about once in 63 years |
| 5 standard deviations | about 1 in 1.7 million | about once in 7,000 years |
| 6 standard deviations | about 1 in 500 million | about once in 2 million years |
Both halves of the table matter. The top half is useful calibration: if a series were even approximately normal, two-sigma events are monthly business, not emergencies. Anyone treating a two-sigma reading as a rare crisis hasn't internalized how common two-sigma is. The bottom half flags the trap. Under normality, a five-sigma day shouldn't have happened since the Bronze Age. Markets deliver them every few years. You can verify that against any long return series, and it is the central fact of this lesson.
10.2.4 Fat tails: where the normal model dies
Real return distributions differ from the normal in a specific, consistent way: more days near zero, fewer middling days, and far more extreme days. The shape is called fat-tailed or heavy-tailed. The statistical name for the property is excess kurtosis, the fourth moment of the distribution. It weights deviations by their fourth power, so it is almost entirely a measure of how wild the wildest observations are. You rarely need to compute kurtosis. You need to know what it flags: two return series can have identical means and identical standard deviations while one delivers its variance steadily and the other delivers it as long calm punctuated by explosions.
The classic demonstration is the crash of October 19, 1987. The US equity index fell more than 20 percent in a single day. Against the daily standard deviation of the preceding years, that was a move in the neighborhood of 20 sigmas. Under the normal distribution, a 20-sigma event is so improbable that you wouldn't expect one in a universe of trading days running from the Big Bang to now. It happened anyway, on a Monday. Every serious market since has supplied its own smaller versions: index moves of 5 to 10 sigmas arrive every few years, single stocks gap 4 sigmas on earnings routinely, and crypto runs the whole demonstration on fast-forward, because leverage plus thin liquidity produces tails that make equities look tame. Bitcoin lost close to 40 percent in a single day on March 12, 2020. Individual alts regularly print daily moves that would be multi-decade events under a normal model calibrated to their trailing vol.
Tail Days: Normal Model vs Reality
Why are the tails fat? You've already met both mechanisms in this course. One is volatility clustering, from the realized vol lesson: markets switch between calm and turbulent regimes, and a mixture of quiet periods and violent periods, each individually well-behaved, produces a combined distribution with a sharp peak and heavy tails. Much of measured kurtosis is regime-switching in disguise. The other is feedback. The central limit theorem needs independence, and market participants aren't independent: stops trigger stops, liquidations trigger liquidations, dealers hedging short gamma sell into declines, vol targeting funds cut exposure simultaneously when vol rises. The microstructure and liquidation lessons showed you the machinery. When the actors respond to each other, moves compound instead of cancel, and the bell curve's core assumption is exactly what breaks.
What fat tails mean for you, concretely. Standard deviation understates tail risk by construction, so any risk estimate of the form "a two-sigma loss is the bad case" is optimistic, and the further out you go the more optimistic it gets. Sizing rules in the coming lessons handle this by targeting volatility conservatively and capping losses structurally, not by trusting the sigma count. When someone explains a blowup as "a ten-sigma event nobody could have foreseen," the correct translation is "our model assumed thin tails and the market has fat ones." The event wasn't impossibly unlucky; the model was wrong, and knowably wrong, because every return series ever examined has fat tails. And the sigma-based readings on this platform, all the z-scores, remain useful as rulers for what is unusual, but the probability column of the normal table above should never be applied to their extremes. A three-sigma reading is rare relative to that instrument's history. It's not a 1-in-370 event, and treating it as one is how you underprice the chance of a four.
10.2.5 Skewness: which tail is fat
Kurtosis says the tails are heavy. Skewness says whether the weight sits more in one tail than the other. A distribution with a longer, heavier left tail is negatively skewed: lots of small gains, occasional large losses. A longer right tail is positive skew: lots of small losses or small gains, occasional large wins.
Equity indices are negatively skewed at the daily horizon: the worst days are substantially bigger than the best days, and crashes are faster than rallies. You already know the reasons from earlier lessons: leverage forces selling but rarely forces buying, protection gets panic-bid on the way down, and the dealer hedging flows from the options lessons amplify declines when the street is short downside gamma. This asymmetry is also the deep reason index put skew exists, connecting back to the skew lesson: the options market charges more for the tail that's actually fatter.
Strategies have skew too, and this is where the concept earns money or costs it. Selling volatility, the concave strategies from Part 9, produces return streams with strong negative skew: months of steady small gains, then an occasional loss that takes back many months at once. Trend following and long options produce the mirror image: frequent small losses, occasional large wins. Neither shape is inherently better. But the shapes fail differently, they feel different to trade, and negative skew has a specific danger: it flatters every backward-looking summary right up until the bad tail shows up. A short-vol track record with no crisis in the sample looks smooth, high-mean, low-standard-deviation, and the statistics aren't lying about the past; they're just silent about the tail that hasn't been sampled yet. The performance measurement lesson two lessons from now deals with this at length. The reading habit to carry: whenever you see a suspiciously smooth return stream, ask what its skew is. Mean and standard deviation together cannot distinguish a genuinely safe strategy from a negatively skewed one that has not paid its bill yet.
A quick connection back to means: under skew, mean and median separate in a predictable direction. Negative skew pushes the mean below the median (the typical period is better than the average, the bad tail drags the average down). Positive skew does the opposite, which is the alt-coin portfolio from earlier: median bleed, mean carried by the moonshots. When you know the skew, you know which summary number to distrust.
10.2.6 Correlation and where it goes wrong
Correlation measures how two series move together. The Pearson correlation coefficient, the standard one, is the covariance of the two series divided by the product of their standard deviations:
Covariance is the average of the products of the two series' deviations from their means. When x is above its mean while y is above its mean, the product is positive; when they deviate in opposite directions, negative. Dividing by the two standard deviations rescales the result to sit between -1 and +1 no matter what units the inputs use. In plain terms: +1 means they moved in lockstep, -1 means they mirrored, 0 means no linear relationship in the sample. A correlation of 0.5 means the co-movement is real but loose; knowing one gives you a meaningful but far from complete read on the other.
Correlation matters to a trader for two reasons that will each get a full lesson later. Diversification: the risk of a portfolio depends on the correlations between its pieces, and combining low-correlated return streams is the closest thing to free money in this business, which is where the portfolio construction lesson ends the part. And relative value: the pairs framework from the cross-asset lessons runs on relationships between instruments, and correlation is where measuring a relationship starts, even though (as that lesson explained) cointegration is what a spread trade actually needs. This lesson's job is narrower: making sure the correlation numbers you compute or read aren't garbage. There are four standard ways they go wrong.
The most common: correlating prices instead of returns. Two price series that both trend upward over the sample will show a correlation near +1 even if their day-to-day moves are unrelated, because both series spend the early sample below their mean and the late sample above it. The number is real arithmetic and meaningless information. Bitcoin's price and the number of streaming subscriptions correlate beautifully over the 2010s; both went up. Always correlate returns (or changes), never levels, unless you specifically know why levels are the right question. This single mistake accounts for a large share of the spurious relationships that circulate as trade ideas.
Another: treating a sample correlation as a stable property. Correlation is an estimate from a window, with all the sampling noise that implies, and the underlying relationship itself drifts. You saw the live example in the crypto cycles lesson: BTC's correlation with equity indices runs high in some regimes, near zero in others, and the shift between those regimes is itself the tradeable information. A rolling correlation chart shows the estimate wandering across values that would each imply an entirely different hedging or sizing decision.
Rolling BTC vs SPX Correlation
Small samples are a trap of their own. A correlation computed from 20 observations is close to noise: with unrelated series, samples that size routinely produce correlations of plus or minus 0.4 by luck alone. The relative value screener's statistics are computed over long windows for exactly this reason. When you eyeball a relationship on a chart covering a few weeks, you're estimating a correlation from a sample too small to mean anything, with the pattern-hungry visual cortex the psychology lessons warned you about doing the estimating.
The most dangerous: correlation is a single number describing the average co-movement, and average co-movement isn't what kills you. Many pairs of assets are loosely related in calm markets and tightly related in stressed ones, because the selling in a stress event is indiscriminate: everything liquid gets sold to fund losses elsewhere. A correlation of 0.3 measured over a calm sample can hide a correlation near 1 conditional on a crisis, which means the diversification you measured is exactly the diversification you won't have when you need it. This is tail dependence, the correlation cousin of fat tails, and it's the mechanism behind several of the blowups in the case-study lesson at the end of this part. The drawdown lesson covers what to do about it. For now: read every correlation as "in the sampled conditions," and assume the stressed number is worse.
10.2.7 Z-scores: the platform's native language
The z-score is the most used statistic on this site, and it is nothing more than the tools already covered, assembled together.
Take today's value of something, subtract the average of its own recent history, divide by the standard deviation of that history. The result is how many standard deviations today sits from normal-for-this-thing. In plain terms: it's the "how unusual is this" ruler from the standard deviation section, applied and made comparable.
Comparable is the point. Funding rates are quoted in percent per interval, open interest in dollars, skew in vol points, the VRP in vol points, dark pool short-volume in ratios. Raw, these numbers can't be compared or ranked across instruments. As z-scores, they all speak one language: zero means typical, +2 means unusually high for this instrument's own history, -2 unusually low. That's what lets one screener rank hundreds of symbols across completely different metrics, and it's why the crypto dashboard, the equity screeners, and the futures pages all normalize this way. A worked example with funding: suppose a coin's funding rate has averaged 0.01 percent per interval over the lookback with a standard deviation of 0.008. Today it prints 0.03. z = (0.03 - 0.01) / 0.008 = +2.5. Longs are paying two and a half standard deviations more than usual to hold this position, which is the crowding signal the crypto positioning lessons built a strategy on.
What a Z-Score Measures
As a live instance of that same z-score pushed to an extreme: on 2026-06-29 the crypto dashboard's most stretched funding reading was FARTCOINUSDT, whose funding rate of about 0.063 percent per 8-hour interval (roughly 69 percent annualized) sat about 6.4 standard deviations above its own recent average. That is the worked example carried past where a normal distribution says it can go: a +6.4 z-score is a once-in-the-history-of-the-universe event under the bell curve, and it printed on a Monday. Longs were paying extraordinarily to stay long, exactly the crowding a positioning fade looks for, and exactly the kind of reading the next paragraph warns can stretch further before it snaps.
Here is the literal reading of the platform's thresholds. The site highlights readings beyond plus or minus 2 in amber and beyond plus or minus 3 in red, and the skew and dark pool screeners require a z-score of at least plus or minus 2 to surface a name. If the underlying series were normal, a value beyond +2 would occur on about 2.3 percent of observations, roughly one in 44, so the plus or minus 2 cutoff catches about the most extreme 2 to 3 percent on each side. Beyond 3 would be about one observation in 740 on each side. So the amber and red bands are calibrated to mean "top few percent of this instrument's history" and "genuinely rare for this instrument," which is how to read them.
The fat-tail lesson applies here: for series with heavy tails, the normal probabilities overstate how rare the extremes are. Funding rates, liquidation totals, and single-stock skew all have fatter tails than the bell curve, so their three-sigma readings arrive more often than one in 740, and their extremes go further than a normal-world intuition expects. A funding z-score of +3 is stretched; it can go to +5. This is why the strategy lessons paired positioning extremes with confirmation, momentum turning or the level breaking, instead of fading a big z-score on sight. The z-score tells you the rubber band is stretched. It doesn't tell you the band can't stretch further, and it puts no timestamp on the snap.
The lookback is a choice, and it's doing silent work. A z-score compares today against a window of history, and the window defines "normal." A short window adapts quickly and forgets quickly: after a month of elevated funding, a short-window z-score reads elevated funding as the new normal and stops flagging it. A long window remembers more and adapts slower. Neither is right; they answer different questions ("unusual versus recently" versus "unusual versus the broader regime"). When a z-score reading surprises you, the first diagnostic is always to look at the raw series and ask what window would produce that number. The platform's z-scores use consistent conventions per dashboard, so readings are comparable across symbols within a page, which is the property that matters for screening.
Z-scores of trending series pin at extremes. If a metric moves to a new level and stays there, its z-score spikes and then decays back toward zero as the window fills with the new level, even though nothing reverted. Conversely, a series in a steady trend keeps printing high z-scores indefinitely. A persistent +2 isn't a stronger signal than a fresh +2; it may just be a series that trends. This is the stationarity problem in practical form: mean and standard deviation only summarize a process whose behavior is stable over the window, and markets change regime. It is also the statistical reason Part 6 put regime first: the same z-score means different things depending on whether the underlying process is the one the window sampled.
Percentiles are the assumption-free alternative. A percentile rank says "today's value is higher than X percent of the lookback observations," using ranks instead of means and standard deviations. It makes no distribution assumption at all, which makes it immune to the fat-tail distortion: the 98th percentile is the 98th percentile whatever the shape. The cost is that percentiles saturate: once today's value exceeds everything in the window, it reads 100 whether it beat the record narrowly or by a factor of five, while a z-score keeps measuring how far beyond. That's why IV gets quoted both ways, as you saw in the implied volatility lesson: IV percentile for "where are we in the range," a z-score or rank-plus-distance view for "how violently outside it." The two together read a distribution better than either alone.
10.2.8 Putting the kit together on one screener row
Walk through the volatility screener with everything above loaded. Each row is a symbol whose VRP, implied vol minus realized vol as defined back in the volatility lessons, sits next to where that premium ranks in the name's own history. The rows lit up at the top are the ones whose premium is unusually rich for themselves, not the ones with the biggest raw number.
Take the sell-volatility list from a recent session. PBR shows a VRP of about 8.7 vol points sitting at the 96th percentile of its own history: a middling raw number that is near a record for this name. A few rows up, NN shows a larger VRP of about 13.1 vol points sitting at only the 31st percentile: a bigger raw premium that is actually below normal for NN. The screener is not ranking the size of the premium. It is ranking how unusual today's premium is for each name against its own past, which is the z-score idea from the previous section applied down a whole column.
This is the trap worth naming, because it catches people on every normalized column and hardest on VRP and skew. A high or low z-score does not tell you the sign of the underlying quantity. The z-score measures distance from the name's own average, and that average carries its own sign. VRP usually runs positive, but for a name that habitually trades with realized above implied it runs negative, and such a name can print a high positive z-score, a premium unusually rich for itself, while the VRP level is still negative. Skew makes it starker: index and single-stock put skew is persistently negative, so a skew z-score of +2 does not mean skew turned positive. It means the skew is two standard deviations less negative than normal for that name, which is still a negative number. Read the rank and the level as two separate facts: one says how unusual, the other says which direction.
Decoded, then, a top-of-screen VRP row says this name's options are pricing future movement at a premium over recent movement that is unusually rich for this name, measured over the site's lookback. What it does not say: it does not say the premium must revert this week (trends pin z-scores), it does not say the options are mispriced (a nearby earnings date would justify a fat premium, which is why the screener shows days to earnings), it does not tell you the raw VRP is even positive without checking the level, and it does not put a ceiling on the reading. The number is a ruler reading, not a verdict. The verdict comes from the strategy rules built around it in Part 9, and the sizing that makes being wrong survivable comes in the lessons immediately ahead.
That decoding loop is the practical skill this lesson is after: from the highlighted number back to what was measured, over what window, with what assumptions, in which direction, and what got compressed away. Every number on the platform survives that loop; most numbers on social media don't.
One habit to carry from all this: every number you trust is a summary, and a summary is only as good as your memory of what it discarded. The next lesson takes this kit into the question every trader cares about: what makes a strategy profitable in expectation, why a high win rate proves nothing by itself, and how long pure luck can impersonate skill before the distribution shows its hand.
10.3 Thinking in distributions
The last lesson gave you the vocabulary: means, standard deviations, fat tails, z-scores. This lesson is about the mental shift that vocabulary enables: stop thinking about trades and start thinking about the distributions that trades are drawn from.
Here is the shift in one example. You short a funding extreme in a crypto perp, the market squeezes 12 percent against you, and you stop out for a full loss. Was it a bad trade? The question is malformed. One outcome can't tell you whether a decision was good, any more than one card tells you whether a poker hand was played well. The trade was a single draw from a distribution of possible outcomes, and the only meaningful question is whether that distribution had a positive mean and a shape you could survive. If it did, you made a good trade that lost money. Those two things coexist constantly, and traders who can't hold both in their head at once end up abandoning good strategies after normal losses and doubling down on bad strategies after lucky wins.
Everything in this lesson follows from taking that seriously. We will rebuild expectancy properly, demolish the win rate as a standalone number, work out how much luck sits inside any track record, and put actual numbers on the two questions that decide whether you stick with a strategy: how many trades before you know anything, and how long a stretch of bad results can last while nothing is actually wrong.
10.3.1 Every trade is a draw
Picture a strategy as a bag of tickets. Each ticket has a number on it: +0.4R, -1R, +2.7R, -0.9R, where 1R is the amount you risk per trade, the same unit from the strategy framework lesson. Every time you take a trade, you pull one ticket. You don't get to choose which one. The rules of the strategy determine what tickets are in the bag and in what proportion; the market determines which one you pull today.
This picture sounds trivial and almost nobody trades as if it were true. The natural human move is to treat each outcome as a verdict on the decision that produced it. Win, and the analysis was right. Lose, and something must be fixed. Poker players have a word for this failure, resulting, and it's the correct word for most of what passes for trade review. If the process was sound, a loss teaches you nothing except that losses exist, which you already knew. If the process was unsound, a win teaches you something actively harmful, because it pays you to repeat a mistake.
The discipline that follows from the ticket picture is to evaluate decisions against the set of things that could have happened, not the one thing that did. A trader who sells a 5-delta call for a tiny credit and watches it expire worthless made money in the history that occurred. In a good fraction of the alternative histories, the stock gapped through the strike and the loss was twenty times the credit. He didn't experience those histories, but he was exposed to them, and the exposure is what should be judged. Risk that didn't materialize is still risk that was taken. You'll meet this idea again in the blow-up case studies later in this part: nearly every famous disaster was preceded by years of results that looked like skill and were actually unpriced exposure.
The practical version of all this: your job is not to win trades. Your job is to keep pulling tickets from bags with positive means and survivable shapes, and to be roughly indifferent to any individual pull.
10.3.2 Expectancy is a mean, and means hide things
Back in the strategy framework lesson you met expectancy:
where p is the win probability, W the average winner in R, and L the average loser in R as a positive number. In plain terms, it's the mean of the ticket bag: what you make per trade on average, counting the losers. Positive means the strategy earns money over enough trades; negative means nothing downstream of it can be fixed.
What that lesson used it for was comparing strategies of different shapes. What this lesson adds is a warning: expectancy is only the mean of the distribution, and the mean is one number describing an object that needs several. Two strategies with identical +0.2R expectancy can differ wildly in variance (how far typical outcomes sit from the mean), in skew (which tail is fat), and in how reliably you can estimate any of it from a sample. The mean tells you whether the bag is worth pulling from at all. The rest of the distribution tells you what pulling from it will feel like, how big you can size it, and how easily you can be fooled about it. Most of the damage traders do to themselves comes from knowing the first thing and ignoring the second.
One more property of the mean matters here, because it shapes the next three lessons. Expectancy is an average across many independent pulls, the way a casino experiences its tables. Your account doesn't experience the bag that way. It experiences one specific sequence of pulls, compounding into each other, and a positive-mean bag can still destroy a specific account if the pulls are sized so large that a normal bad sequence digs a hole too deep to climb out of. The full treatment of that problem belongs to the Kelly and drawdown lessons. The distinction to keep: the bag has a mean, but your account lives one path through it.
10.3.3 Why a 78 percent win rate says nothing
Somebody shows you a track record: 78 percent winners over the last six months. Most people hear that number and are done evaluating. It sounds like an answer. It's not even half of one, because expectancy has three inputs and win rate is one of them.
Run the numbers. Suppose those winners average +0.25R because the strategy takes profits quickly, and the 22 percent of losers average -1R because that's where the stop sits.
expectancy = 0.78 * 0.25 - 0.22 * 1.0 = 0.195 - 0.22 = -0.025R
Negative. This trader loses money at a 78 percent win rate, slowly and with great confidence, and the win rate itself is the anesthetic that keeps him from noticing. Winning weeks pile up, the account bleeds a little at a time, and every individual loss looks like an exception rather than the load-bearing part of the arithmetic.
Now flip it. A trend-following approach wins 32 percent of the time, average winner +2.8R, average loser -1R:
expectancy = 0.32 * 2.8 - 0.68 * 1.0 = 0.896 - 0.68 = +0.216R
Positive, and comfortably so. This trader is wrong twice as often as he is right and makes good money, because the size of the payoffs carries the arithmetic that frequency can't.
The general relationship is worth having in closed form. If your average winner is W and average loser is L, the win rate you need just to break even is:
In plain terms: the smaller your winners are relative to your losers, the more often you have to be right, and the relationship is unforgiving. Here's the breakeven line across payoff ratios:
| Average win / average loss | Breakeven win rate |
|---|---|
| 0.25 | 80.0% |
| 0.5 | 66.7% |
| 1.0 | 50.0% |
| 1.5 | 40.0% |
| 2.0 | 33.3% |
| 3.0 | 25.0% |
| 5.0 | 16.7% |
Our 78 percent trader with his 0.25 payoff ratio needed 80 percent just to break even. He was two points short and couldn't see it, because nobody who wins 78 percent of the time feels like they're two points short of anything.
Breakeven Win Rate
So when someone quotes a win rate, whether it's a signal service, a backtest, or your own journal, the immediate follow-ups are: average winner in R, average loser in R, and over how many trades. Without the first two, the number is decoration. Without the third, even the full triple might be noise, which is where this lesson is headed.
10.3.4 You can buy any win rate you want
The deeper reason win rate carries so little information is that it's a choice, not a discovery. You can set your win rate almost anywhere you like by moving your exits, and the market will charge you for it elsewhere in the distribution.
Take one entry signal, any signal, and attach different exits. Version one: take profit at +0.3R, stop at -2R. The profit target is close and gets hit constantly; the stop is far and gets hit rarely. Win rate somewhere in the 80s. Version two, same entries: take profit at +3R, stop at -0.5R. Now the tight stop gets clipped by ordinary noise most of the time and the distant target is reached occasionally. Win rate somewhere in the 20s or 30s. Same signal, same information content, radically different win rates. All you did was slide probability mass between the tails, trading frequency of winning against size of winning at roughly actuarial rates.
Options make the same purchase even more nakedly. Sell a 5-delta option and you've bought yourself roughly a 95 percent win rate before hedging, by construction, and the price is a left tail that can hand back many months of premium in one move. The seller didn't find a 95 percent edge. He selected a 95 percent shape. Whether there's any edge in it at all depends on whether the option was priced above its actuarial value, which is the volatility risk premium question from Part 3 and has nothing to do with the win rate itself.
This is why the strategy framework lesson insisted that high win rate doesn't mean good strategy; it means concave strategy. Add the converse: a low win rate doesn't mean bad strategy; it means convex strategy. Win rate tells you which shape someone chose. Expectancy, net of costs, tells you whether the choice is getting paid. Keep the two questions separate and a large fraction of marketing in this industry stops working on you.
There's one honest use of win rate, and it's psychological rather than statistical. The shape you choose determines the experience of trading the strategy: how often you get the small dopamine hit of a win, how long the losing streaks run, how it feels at dinner parties. Those things matter for whether you can actually execute the system, and the streak arithmetic near the end of this lesson is how you price them. Just never confuse the comfort of winning often with the presence of an edge.
10.3.5 Luck and skill
This lesson moves from one strategy to the population of people trading. Trading results are a mix of skill and luck, and over short horizons the mix is far more luck-heavy than most people accept about themselves.
Here is the standard thought experiment; its arithmetic is the point. Put 10,000 people in a room and have each one manage money by flipping a coin: heads, up year; tails, down year. Nobody has any skill by construction. After one year, about 5,000 have a winning year. After two, about 2,500 have two in a row. After five years, roughly 310 of them have five consecutive winning years. After ten, around 10 people are sitting on a decade of unbroken success, and every one of them is a coin. Give those ten a marketing budget and a confident origin story and from the outside they are indistinguishable from the genuinely skilled.
The mechanism at work is selection. You never observe the full room, only the survivors, because the losers stopped posting, closed the fund, or quietly went back to their day job. Every place you encounter track records (fund league tables, social media, your own circle of trading friends) is a survivor pool, and a survivor pool systematically overstates how much skill is out there and how achievable the visible results are. When you see an impressive short track record, the correct prior is not "this person has an edge." It's "given how many people are trading, records like this must exist even if nobody has an edge, and I can't tell from the record alone which case I'm looking at."
The same reasoning applies inward, and the inward version is the expensive one. Your own results over 20 or 50 trades sit inside the same fog. A strong first year proves much less than it feels like it proves, and the feeling is dangerous precisely because it arrives with money attached: the natural response to a lucky streak is to conclude you're skilled and size up, which maximizes your exposure at the exact moment your self-assessment is most inflated. Regression toward the mean isn't a moral judgment, it's what happens mechanically when a result contained a large luck component and the luck washes out on the next sample. The trader who returns 60 percent in year one and 5 percent in year two usually didn't lose his touch. He revealed his mean.
None of this says skill doesn't exist. It says skill is slow to prove. In an activity where outcomes are mostly determined by ability, like chess, a handful of games separates the strong player from the weak one, because variance is low relative to the skill gap. Trading sits near the other end: per-trade outcomes are dominated by noise, edges are thin, and the skill gap between a decent trader and a mediocre one might be a few hundredths of an R per trade. Thin signal under heavy noise takes a long sample to detect. How long has a numerical answer, given two sections from now.
10.3.6 The noise in your own P&L
One consequence of noise-dominance costs traders money every day. The shorter the window you evaluate over, the less information it contains, and at the windows most people actually watch, the information content is close to zero.
Take a strategy any professional would take: 15 percent expected annual return with 10 percent annualized volatility. Assume returns are roughly normal for this illustration and scale by the square root of time, as covered in the statistics lesson. The probability that any given observation window shows a profit works out approximately as follows: a year is profitable about 93 percent of the time, a quarter about 77 percent, a month about 67 percent, and a day about 54 percent. At shorter intervals the number keeps sliding toward 50.
For an excellent strategy, a single day is 54/46. Checking your P&L intraday means consuming a stream that is nearly a coin flip, and your nervous system doesn't price it that way: losses hurt roughly twice as much as equivalent gains feel good, so a 54/46 stream experienced tick by tick nets out emotionally negative even while the account grows. The trader who refreshes his P&L forty times a day is strapping himself into a machine that delivers mostly noise and mostly pain, then making discretionary decisions in that state.
The point isn't to stop monitoring risk; positions need watching. The point is to match your evaluation horizon to where the signal lives. Risk gets monitored continuously. Performance gets judged on samples large enough to mean something. Those are different activities, and collapsing them into one anxious habit is how good strategies get abandoned in week three.
10.3.7 How many trades before you know anything
Suppose you've been trading a system and want to know what your sample actually tells you.
Start with the win rate, the easiest quantity to pin down. From the statistics lesson, the standard error of an estimated proportion is:
In plain terms, this is the typical distance between the win rate you measured and the true one, and it shrinks only with the square root of the number of trades. A useful rule of thumb is that the true value sits within about two standard errors of your estimate, 95 times out of 100.
With 30 trades and a measured 60 percent win rate: SE = sqrt(0.6 * 0.4 / 30), which is about 0.09. Two standard errors is 18 points, so the true win rate is somewhere between roughly 42 and 78 percent. That interval contains a losing strategy and a spectacular one. Thirty trades, the sample size at which most people have already formed a permanent opinion of a system, distinguishes almost nothing.
With 100 trades the interval tightens to about plus or minus 10 points. Still wide. To pin a win rate down to within 5 points either way, you need on the order of 400 trades. Square-root shrinkage is brutal like that: each halving of the uncertainty costs four times the data.
Win rate is the easy case. What you actually care about is expectancy, and expectancy is a mean, so its uncertainty is:
where s is the standard deviation of your per-trade outcomes in R. To be reasonably confident an edge is real, you want the measured edge to be at least about twice its standard error, which rearranges to a required sample size:
Plug in realistic numbers. A decent swing strategy might earn +0.2R per trade with a per-trade standard deviation of 1.5R. Then n = (2 * 1.5 / 0.2)^2 = 225 trades before the data alone separates your edge from zero. At 40 trades a year, that's more than five years. Now try a thinner but still worthwhile edge of +0.05R with the same 1.5R spread: n = (2 * 1.5 / 0.05)^2 = 3,600 trades. For most discretionary traders that's several lifetimes. The uncomfortable conclusion is that plenty of real, paying edges are statistically unverifiable from any live sample their trader will ever collect. You'll act under uncertainty forever; the point of the math is to know how much.
And it gets worse when the distribution is skewed, which after the framework lesson you know describes most strategies worth running. The formulas above treat every trade as equally informative, but in a skewed strategy the expectancy calculation is dominated by rare tickets. A concave strategy with a 95 percent win rate loses big about once in twenty trades; after 100 trades you've observed the left tail perhaps five times, and your estimate of the average tail loss, the term that decides whether the whole thing is profitable, rests on five data points drawn from the fattest-tailed part of the distribution. A convex strategy has the mirror problem: its expectancy hangs on a handful of large winners, and a sample window that happens to miss one big trend will report a healthy strategy as a losing one. In both cases the effective sample size for the number that matters is a small fraction of the trade count. This is also the statistical root of why backtests mislead, which the backtesting lesson takes up properly: a backtest is just a sample too, with all of these problems plus some self-inflicted ones.
Two working rules fall out of this section. Put confidence intervals on everything you compute from your journal, even rough ones; a win rate without an interval is a feeling, not a measurement. And since live samples will rarely settle the question, the burden shifts to the quality of the reasoning behind the strategy: who pays this edge and why, the question the framework lesson trained you to ask. Structural logic plus a consistent sample beats an impressive sample with no logic, because the impressive sample is exactly what luck produces in a big enough population.
10.3.8 How long bad variance can last
The last piece saves strategies from being abandoned at the worst possible moment: knowing, in advance and in numbers, what normal bad luck looks like for your specific distribution.
Start with losing streaks, because they're what actually breaks people. For a strategy with loss probability q per trade, a good approximation for the longest losing streak you should expect over N trades is:
In plain terms, streaks grow with the logarithm of how long you trade, so more trading guarantees longer worst streaks, slowly but relentlessly. Over 100 trades: a 60 percent win rate strategy (q = 0.4) should expect a worst streak of about 5 consecutive losses. A 40 percent win rate strategy (q = 0.6) should expect about 9 in a row. A 30 percent winner, the profile of many trend systems, should expect a worst run of around 13, and over a 400-trade career closer to 17. None of these streaks would indicate that anything is wrong. They're what the arithmetic promises when nothing is wrong.
The middle number is worth examining: nine consecutive losses, on a strategy that's working exactly as designed. If you haven't computed this before you start trading, trade seven of that streak is where you conclude the edge is gone, and trade nine is where you stop, historically just in time to miss the winner. If you've computed it, the streak is an expected weather event: unpleasant, survivable, and pre-priced. The single highest-value output of the streak formula is a number written down before you begin: "this system's worst expected run over the next 200 trades is X, and I don't get to reevaluate the system on streak evidence until well past X."
Variance operates over whole calendar periods too. Take a solid strategy: +0.15R expectancy, per-trade standard deviation of 1.5R, 100 trades a year. The year's expected total is +15R. But the standard deviation of the annual total is 1.5 * sqrt(100) = 15R, exactly as large as the mean. Under a normal approximation, the chance that a full year finishes negative is the chance of a one-standard-deviation shortfall, roughly 16 percent. One year in six, this good strategy loses money over a full year while remaining exactly as good as it ever was. Thin the edge to +0.05R and the losing-year probability climbs to about 37 percent: more than one year in three. Nobody feels these numbers intuitively. Everybody assumes a positive-expectancy strategy should produce positive years the way a fair salary produces positive months, and the assumption quietly wrecks more systematic traders than any modeling error.
Fifty Years That Never Happened
A related exercise is worth doing on your own trade history once you have any: take your actual logged trades, shuffle their order at random a few thousand times, and look at the spread of equity paths the shuffles produce. The trades are identical; only their order changes. The spread is usually a shock the first time. Paths built from identical trades differ enormously in maximum drawdown and in how long they spend below their prior high, purely from sequencing. The exercise shows that your realized equity curve is one draw from a family of curves you could just as easily have lived, and judging yourself on its specific wiggles is resulting at the portfolio level. The drawdown lesson later in this part builds directly on this picture.
All of this raises the question the math can't fully answer: if bad stretches this long are normal, how do you ever detect that a strategy has actually died? Not from the streak or the losing quarter alone; you now know those prove nothing by themselves. The evidence that means something is a change in the mechanism. The edge you were harvesting had an identified payer, and structural evidence that the payer left is worth more than any run of outcomes: the positioning extreme that stopped mean-reverting because the participant mix changed, the premium that compressed because too much capital crowded in, the regime shift that the regime lessons taught you to read. Outcome data gets a vote only at full sample sizes. Mechanism evidence gets a vote immediately. Traders who monitor the mechanism can hold through noise with justified confidence and still exit dead strategies years before the statistics would have convicted them.
10.3.9 Living with your distribution
The working rules this lesson has generated together amount to an operating manual for the statistical fog.
Judge decisions by process and exposure, not by single outcomes, and audit your winners as ruthlessly as your losers; a paid-off bad process is the most expensive thing you can learn from. Never evaluate a win rate without its payoff sizes, or an expectancy without its sample size, or a sample without its interval. Before trading any system, compute its expected worst losing streak and its probability of a losing quarter and year, write the numbers down, and pre-commit to what evidence would actually justify shutting it down, so the decision is made by a calmer version of you than the one who'll be nine trades into a drawdown. Match your evaluation horizon to where the signal lives, and treat intraday P&L as risk telemetry, not as feedback about whether you're good. And hold every impressive track record, especially your own, against the base rate of what luck alone produces in a population this size.
None of this makes variance hurt less. It makes variance expected, and expected pain is the kind that doesn't force errors.
The natural next question is how to compress a whole return distribution into summary numbers you can compare across strategies, which is what ratios like Sharpe attempt. The next lesson works through those measures and, more importantly, through what each one hides, because a summary statistic that ignores skew will happily award its best scores to the strategies with the worst tails.
10.4 Measuring performance
The last lesson ended on an uncomfortable note: a win rate says nothing by itself, expectancy needs a large sample before you can trust it, and variance can impersonate skill for years. So suppose you have the sample. Two years of your own trades, or a fund's five-year track record, or a backtest of one of the Part 9 strategies. How do you compress that pile of returns into a verdict? This lesson covers the standard performance measures: what each one computes, what each one hides, and the specific way negatively skewed strategies fool all of them for a while. By the end you should be able to look at any track record, yours included, and know which questions the headline numbers have quietly skipped.
The theme from the statistics lesson carries straight through. Every performance measure is a compression of a return series into one number, every compression throws information away, and the discarded information is where the surprises live. The measures in this lesson discard different things, which is why you use several of them together and never just one.
10.4.1 The problem with raw return
Consider the number everyone quotes first: "the strategy made 40 percent last year." Standing alone, that number measures almost nothing, for two reasons.
One is risk. A 40 percent year achieved with 12 percent volatility and a worst drawdown of 8 percent is a different object from a 40 percent year achieved with 60 percent volatility and a 45 percent drawdown, even though the endpoints match. The second version was a coin flip that landed well. Run it again and the same process produces a wipeout as easily as a repeat. Return without a risk denominator has no context.
The other is leverage. Take any strategy with a positive expected return and borrow money to double the position. Return doubles (minus funding costs). Borrow more, it triples. Raw return is a dial you can turn, not a property of the strategy. If someone can manufacture the number by changing position size, the number can't be measuring skill. What leverage can't change is the ratio of return to risk: double the position and the volatility doubles right alongside the return, leaving the ratio where it was. That invariance is the entire reason risk-adjusted measures exist, and it's why every serious comparison of strategies happens in risk-adjusted units.
So the real question is never "how much did it make" but "how much did it make per unit of risk taken." The disagreements between the measures below are disagreements about how to define the unit of risk.
10.4.2 The Sharpe ratio
The Sharpe ratio is the default answer, the one number every allocator, fund, and backtest report leads with. The definition:
where R is the strategy's average return over some period, Rf is the risk-free rate over the same period, and sigma is the standard deviation of the strategy's returns. In plain terms: excess return per unit of volatility. How much you got paid, above what a money market fund would have paid you for zero effort, for each unit of variability you endured.
Both pieces of the definition are doing real work. Subtracting the risk-free rate matters because return you could have earned in T-bills isn't a reward for anything. When cash yields 5 percent, a strategy returning 8 percent with 10 percent volatility has a Sharpe of 0.3, not 0.8, and that difference isn't pedantry: the strategy is delivering 3 points of actual compensation for 10 points of risk. In a zero-rate world the adjustment vanishes and people forget it exists, then rates rise and half the "absolute return" industry turns out to have been repackaging the cash yield. Dividing by the standard deviation matters because of the leverage argument above: it makes the ratio approximately invariant to position size, so it measures the quality of the return stream rather than its volume.
Sharpe scales with the square root of time, for the same reason volatility does (variances add across independent periods, so standard deviations grow with sqrt of time while means grow linearly). To annualize, multiply a daily Sharpe by sqrt(252) and a monthly Sharpe by sqrt(12). A strategy earning 0.05 percent per day over 1 percent daily vol has a daily Sharpe of 0.05 and an annualized Sharpe of about 0.79. Whenever you see a Sharpe quoted, it's annualized by convention, and whenever you compute one, annualize it, or you'll conclude your daily strategy is garbage when it's fine.
Some calibration for the numbers. Buy-and-hold equity indices have delivered a long-run Sharpe somewhere around 0.3 to 0.4: positive, real, and unimpressive per unit of risk, which is exactly what you'd expect for a premium anyone can collect by opening a brokerage account. A strategy that sustains a Sharpe near 1 over years, net of costs, is genuinely good. Sustained Sharpe near 2 is excellent and rare outside of high-frequency niches. Claimed Sharpes of 3 and above at daily-or-slower horizons deserve immediate suspicion: either the sample is short, the risk is hiding somewhere the standard deviation can't see it (the short-vol section below), or the backtest is lying (next lesson's subject).
There is also a statistical honesty question the Sharpe ratio makes tractable: how long a track record do you need before the number distinguishes itself from zero? A useful approximation is that the t-statistic of a Sharpe estimate is roughly the annualized Sharpe times the square root of the track length in years. To clear the conventional significance bar of about 2, a true Sharpe of 1 needs roughly 4 years of data. A Sharpe of 0.5 needs roughly 16 years. The equity premium itself, at 0.3-0.4, needs most of a century, which is why people were still arguing about its existence after decades of data. This connects straight back to the luck-versus-skill discussion in the previous lesson: a two-year track record with a Sharpe of 0.8 is a coin that came up heads a few times. Promising, worth continuing, and proof of nothing.
10.4.3 What Sharpe hides
Sharpe compresses a return distribution into its first two moments, mean and standard deviation. The statistics lesson spent two full sections on what those two numbers miss: skewness and fat tails. Everything Sharpe hides follows from that compression, plus one thing it hides about ordering.
It penalizes upside volatility as if it were risk. Standard deviation is symmetric: a month of +15 percent widens it exactly as much as a month of -15 percent. A trend-following strategy that grinds small losses and occasionally banks a huge winning month gets its Sharpe dragged down by its best months. The measure is treating the thing you want, big wins, as a defect. For roughly symmetric return streams this doesn't matter. For strongly skewed ones it distorts comparisons in a predictable direction: positive skew gets punished, negative skew gets flattered.
It's blind to the order of returns. Shuffle a return series into any sequence and the mean and standard deviation don't move, so the Sharpe doesn't move. But you don't experience returns as an unordered set. A strategy that lost 35 percent in its first year and spent four years climbing out has the same Sharpe as one that delivered the identical returns as a steady grind. Same number, completely different lived experience, completely different odds that you (or your investors, or your own nerve) survive to see the recovery. Path matters, Sharpe can't see path, and that blind spot is the opening for the drawdown-based measures below.
It trusts the volatility estimate, and the volatility estimate can be gamed by smoothness. The sqrt-of-time annualization assumes returns are independent across periods. When returns are positively autocorrelated (this month up makes next month up more likely), true annual volatility is higher than the scaled monthly number, and the annualized Sharpe overstates reality. Where does autocorrelation come from? Sometimes from genuine trending, but more often from smoothed marks: illiquid holdings priced by appraisal or by stale quotes, monthly reporting that averages over intramonth chaos, or any process where losses show up in the marks gradually instead of at once. A return stream that is smooth because someone smoothed it will post a beautiful Sharpe over risk that never made it into the data. When a track record looks too smooth for its asset class, the volatility in the denominator is the first thing to distrust.
And it says nothing about the tails. Two strategies, identical Sharpe of 1.2. One delivers its risk as a steady hum of moderate ups and downs. The other delivers years of serenity and then a cliff. Mean and standard deviation can't tell these apart until the cliff is actually in the sample. Which brings us to the most expensive blind spot in performance measurement.
10.4.4 The short-vol illusion
This pattern shows up over and over: in other people's track records, in products you're pitched, and in your own backtests of the Part 9 concave strategies. A strategy that sells insurance (short options, short VIX futures, harvesting crypto funding, any premium collection with capped upside and open-ended downside) produces a return stream with strong negative skew. Many small wins, rare large losses. Between the large losses, every backward-looking statistic glows.
Run the arithmetic on a concrete case. A strategy earns 1.0 percent per month with a monthly standard deviation of 1.0 percent, for 59 straight months. Monthly Sharpe of roughly 1, annualized to about 3.5. Nearly five years of that track record: smooth equity curve, tiny drawdowns, a Sharpe that beats almost everything on earth. Then month 60 arrives with its crisis and the strategy loses 25 percent.
Recompute over the full 60 months. The mean drops to about 0.57 percent per month. The standard deviation, dominated by that single observation exactly as the statistics lesson warned, jumps to about 3.5 percent. The annualized Sharpe lands near 0.57. One month took the strategy from world-class to mediocre, and the drawdown-based numbers are uglier still: five years of compounding at 1 percent per month builds the account to about +80 percent, and the single bad month hands a quarter of the account back at once.
| Measurement window | Annualized Sharpe | Max drawdown | Verdict a naive reader gives |
|---|---|---|---|
| Months 1-59 | about 3.5 | a few percent | genius |
| Months 1-60 | about 0.57 | 25 percent | ordinary, with a scary tail |
The example isn't saying the strategy was bad. Depending on the size of the premium collected, selling insurance can be a perfectly sound business, and lessons later in this part make the case for harvesting these premia deliberately. But for 59 months the measured Sharpe wasn't information about the strategy. It was information about which part of the cycle the sample happened to cover. The statistics were accurate about the past and silent about the tail that hadn't been sampled yet, and nothing in the Sharpe computation flags the difference. A reader who knew to ask "what is this return stream's skew, and what is it short of?" would have priced the cliff in advance. A reader comparing Sharpe ratios across strategies as if they were comparable would have allocated everything to the steamroller's path right before it arrived.
This isn't a hypothetical failure mode. Inverse volatility products in the mid-2010s compiled multi-year track records spectacular enough to attract billions, then lost most of their value in a single session in early 2018. Sellers of far out-of-the-money puts post the same shape on a slower clock. Crypto funding harvesters print smooth returns until a violent squeeze or a venue failure resets the account. The blowup lesson at the end of this part walks through the biggest cases mechanically; here the useful rule is narrower and more practical. When a Sharpe ratio looks too good, the first hypothesis is not skill. The first hypothesis is negative skew plus a sample that hasn't paid its bill yet. Ask what the strategy is short of, find the last time that exposure got hit, and check whether the sample includes it. If the answer is no, mentally reprice the track record as incomplete, because it is.
The mirror image also holds. Positively skewed strategies, long options and trend following, systematically look worse on Sharpe than they are: frequent small losses drag the mean, occasional huge wins inflate the standard deviation, and the resulting ratio underprices a return stream whose worst case is knowable and whose best case is open-ended. Two strategies with the same Sharpe but opposite skew are not equivalent, and if you must err, paying up in Sharpe terms for positive skew is the defensible direction to err in.
10.4.5 Sortino: charging only for downside
The Sortino ratio is the standard repair for Sharpe's symmetric penalty. Same numerator, different denominator:
where downside deviation is computed like a standard deviation but using only the returns below the target (usually zero or the risk-free rate): square the shortfalls below the target, average them over all periods, take the root. In plain terms: return per unit of bad volatility, with the good kind not held against you. The trend follower's monster winning months no longer inflate its risk number, and the comparison between a positively skewed and a symmetric strategy becomes fairer.
Two warnings before you lean on it.
Sortino numbers aren't comparable to Sharpe numbers, and people compare them constantly. For a roughly symmetric return stream with a mean near zero, the downside deviation is about the full standard deviation divided by sqrt(2), so the Sortino comes out around 1.4 times the Sharpe with no change in the underlying strategy. A fund quoting "Sortino of 1.8" is describing roughly the same return stream as one quoting "Sharpe of 1.3." Compare Sortino to Sortino, Sharpe to Sharpe, and treat anyone who switches metrics mid-pitch as making a sales decision, not a measurement decision. Conventions also vary (which target rate, whether the averaging divides by all periods or only the down periods), so two people can compute honestly different Sortinos from the same data. Ask how it was computed before trusting a cross-source comparison.
The warning that matters more: Sortino doesn't fix the short-vol illusion. It makes it worse. The denominator now depends entirely on the bad periods in the sample, and the whole problem with negatively skewed strategies is that the sample contains almost no bad periods until it suddenly does. Our 59-month insurance seller has nearly zero downside deviation, so its Sortino is even more absurd than its Sharpe. Downside-only measures are the right correction for strategies whose upside volatility was being unfairly punished. They're the wrong tool, actively misleading, for strategies whose downside simply hasn't shown up yet. No ratio computed from a sample can price a tail the sample doesn't contain. The fix for unsampled tails is not a better ratio; it's structural knowledge of what the strategy is short of, which is why Part 7's map of where returns come from matters more than any formula here.
10.4.6 Calmar and MAR: return against the worst stretch
The third family swaps volatility out of the denominator entirely and replaces it with the thing that actually ends trading careers: the maximum drawdown, the largest peak-to-trough decline the equity curve suffered.
A strategy compounding at 15 percent a year with a worst drawdown of 20 percent has a Calmar around 0.75. The MAR ratio is the same construction; by convention Calmar is often computed over the trailing three years while MAR uses the full history, but in practice the names get used interchangeably and the only safe move is to check the window. In plain terms, both answer: for every unit of worst-case pain endured, how much annual growth did you get?
What this fixes is exactly what Sharpe couldn't see. Drawdown is path-dependent: shuffle the same returns into a different order and the max drawdown changes even though Sharpe doesn't. A strategy that delivers its losses in one concentrated stretch gets a worse Calmar than one that spreads the same losses thinly, which matches how survival actually works. Drawdown also respects compounding, since it is computed off the equity curve rather than the return list. And it partially catches the short-vol illusion after the fact: the moment the cliff enters the sample, the max drawdown records it permanently, whereas rolling volatility forgets a bad month once it scrolls out of the window. A ten-year track record's Calmar carries its 2020 scar forever. Its trailing three-year Sharpe doesn't.
What it breaks is statistical reliability. Maximum drawdown is a single observation, the most extreme point of the most extreme episode in one specific sample, which makes it the noisiest statistic you can compute from a return series. Rerun the same strategy over a parallel history and the max drawdown might easily be half or double, because it hinges on whether a few bad weeks happened to overlap. Worse, expected max drawdown grows with track length: a 15-year record has had more chances to print a deep trough than a 3-year record, so its Calmar is mechanically lower for the same underlying quality. Comparing Calmar across track records of different lengths is comparing apples to a longer exposure to apples. And before the first real crisis is in the sample, drawdown-based measures are just as blind as Sharpe: our insurance seller's 59-month Calmar was magnificent too.
So none of the three families is the answer alone. Sharpe is statistically the best behaved and the most comparable across strategies, and it lies about skew and path. Sortino repairs the skew penalty and doubles down on the unsampled-tail problem. Calmar respects path and survival and is too noisy to rank anything precisely. Used together, disagreements between them are the diagnostic: a high Sharpe with a mediocre Calmar says the losses came concentrated; a Sortino far above its expected 1.4x multiple of Sharpe says positive skew (good); a suspiciously high everything on a short sample says the bill has not arrived. One number is a verdict. Three numbers are a description.
10.4.7 Return streams and equity curves
Underneath every metric in this lesson sit two different pictures of the same trading, and knowing which one you're looking at prevents a whole category of confusion.
A return stream is the sequence of period returns: +1.2 percent, -0.4, +0.8, and so on. It's the analyst's object, (approximately) independent of account size and leverage. It's what Sharpe and Sortino are computed from, what you scale when you apply the vol-targeting ideas coming two lessons from now, and what you correlate against other strategies when you build a book at the end of this part. Two traders running the same strategy at different sizes have the same return stream.
An equity curve is what you get by compounding the stream: actual account value over time. It's the trader's object, the thing you live inside. It's path-dependent, it embeds the volatility drag from the statistics lesson (the gap between average return and compound growth that widens with volatility), and it's where drawdowns exist. The same return stream run at 2x leverage doesn't produce 2x the equity curve; it produces a curve with more than double the drawdowns and less than double the long-run growth, because drag scales with the square of volatility. The full treatment of that arithmetic belongs to the drawdown lesson; the wiring you need now is just that metrics computed on the stream (Sharpe, Sortino) measure the strategy, while metrics computed on the curve (max drawdown, Calmar, compound growth) measure the strategy at a specific size along a specific path. When sizing changes, the second family changes and the first mostly doesn't. A strategy isn't "a Calmar of 1.1." It's a return stream that produced a Calmar of 1.1 at the leverage it happened to run.
Reading an equity curve by eye is a skill worth building deliberately, because the eye catches things the summary numbers smear away. Plot it on a log scale, always: on a linear scale, healthy compounding looks like a recent explosion and early history looks flat, and every judgment you make from the picture inherits that distortion. On a log scale, constant percentage growth is a straight line and changes in slope mean something. Then interrogate the shape. Did the growth come as a steady grind, or did two spikes deliver most of it (in which case the strategy is a tail-catcher and the flat stretches are its normal state)? Is the curve suspiciously smooth for what the strategy trades (see the short-vol section)? Did the character change partway through, smooth then choppy, which often marks a regime the strategy stopped fitting or a size it outgrew?
Then plot the same data as an underwater curve: percent below the running high-water mark at every point in time. This is the most honest chart in performance analysis, because it shows what the summary statistics compress hardest: how deep the drawdowns went and, the part that actually breaks people, how long they lasted. Time underwater is the statistic nobody quotes and everybody quits over. A strategy can have a modest 18 percent max drawdown that took three years to recover, and "three years below high water" appears in no ratio while determining, more than any ratio, whether a human being actually holds the strategy to its long-run numbers.
10.4.8 Reading a track record in practice
This is the sequence to run when a return series lands in front of you, whether a fund pitch, a strategy from Part 9 you backtested, or your own last two years of trading.
Start with the sample itself before computing anything. How long is it, and how long is the strategy's natural cycle? Five years means something for a strategy that trades daily and completes its full cycle of conditions many times over; it means little for a strategy short of a premium that detonates once a decade. Ask what regimes the window covers: does the sample include a vol spike, a rate shock, a real bear market, or only the friendly stretch? Confirm the returns are net of costs, funding, and slippage, because gross numbers on a high-turnover strategy are fiction. And recall from the previous lesson how slowly evidence accumulates: the t-statistic approximation above tells you whether this sample could even in principle distinguish the claimed edge from zero.
Then compute the three families and read the disagreements, not just the levels. Sharpe for comparability, Sortino for the skew correction, Calmar for path and survival, expecting Sortino near 1.4x Sharpe as the symmetric baseline and treating deviations from that multiple as skew information. Look at the monthly return histogram directly and check which tail is longer. Find the worst month and the worst quarter, and ask whether losses of that size make sense given what the strategy claims to do; a worst month suspiciously close to zero on an insurance-selling strategy isn't safety, it's an unsampled tail. Plot the log equity curve and the underwater curve and let your eye check what the ratios summarized.
Finally, and this is the question the numbers can't answer for you: identify what the return stream is short of. Every strategy that earns above cash is being paid for bearing something. If you can name the something (a volatility premium, an event premium, crypto funding, trend risk, liquidity provision), you can reason about when the payment stops and how bad the stopping gets, which is worth more than any ratio computed from the sample. Part 7 built that map in full. A track record whose returns you can't attribute to a nameable premium or a nameable edge is a track record you should assume you don't understand, however good its Sharpe is.
| Ratio | Formula | Denominator measures | What it hides | Roughly trustworthy after |
|---|---|---|---|---|
| Sharpe | (R - Rf) / sigma | total volatility, upside and downside alike | skew, the order of returns, and volatility that was smoothed away | years of data: a true Sharpe of 1 needs about 4, a 0.5 about 16 |
| Sortino | (R - target) / downside deviation | downside volatility only | the unsampled tail, which it flatters worse than Sharpe; not comparable to Sharpe (runs about 1.4x it) | at least as long as Sharpe, and it says nothing before the first real loss |
| Calmar / MAR | annual compound return / max drawdown | the single worst peak-to-trough decline | everything but one episode; it is the noisiest of the three and grows with track length | long records only, and never precise enough to rank strategies |
Everything in this lesson assumed the return series in front of you honestly happened. For your own live trades that's true by construction. For backtests it's the assumption most likely to be false, because a backtest is a return series manufactured under the researcher's control, and there are half a dozen standard ways to manufacture one that looks brilliant and means nothing. The next lesson goes through them one by one: look-ahead bias, survivorship, overfitting, and the multiple-testing problem that makes "I tested fifty configurations and one worked" a statement of failure rather than discovery.
10.5 Backtests and how they lie
A backtest is a claim about a counterfactual: "if I had run this rule over the past ten years, here is what would have happened." It is not evidence the way a live track record is. It is a simulation, built by someone who already knows how the story ends, on data that has been cleaned, adjusted, and filtered by people who also knew how the story ended.
That matters because nearly every error you can make in a backtest pushes the result in the same direction: up. Look-ahead bias inflates. Survivorship inflates. Ignoring costs inflates. Excluding flat days inflates. Testing fifty variants and keeping the winner inflates. No equally large family of mistakes makes backtests look worse than reality. So when you see a backtested Sharpe of 2.1, the right prior is not "this strategy has a Sharpe of 2.1." It is "this number is an upper bound, and my job is to figure out how far below it the truth sits."
This lesson walks through each way the number gets pushed up, with enough mechanical detail that you can spot the problem in your own code and in other people's pitch decks. The previous lesson covered what Sharpe and its cousins measure and what they hide. This one covers how the inputs to those measures get corrupted before the ratio is ever computed.
10.5.1 Look-ahead bias: trading on information you did not have
The purest form of look-ahead is a signal that uses today's close to decide today's position and then books today's return. "Buy when the daily return is positive" backtests beautifully: it captures every up day and skips every down day. It also cannot be traded, because at the moment you would need to place the order, the close does not exist yet.
Nobody writes that bug on purpose. It sneaks in through indexing. The fix, everywhere in quant code, is the one-bar lag: the signal for bar t must be computed from data through bar t-1, and only then multiplied by the return from t-1 to t. In pandas that is a .shift(1) on the signal series before it meets the returns. In a loop, it is using prices[:i] to decide position i, never prices[:i+1]. That off-by-one is the most common bug in amateur backtests, and it hides in the equity curve because the curve it produces looks fantastic.
The subtler versions pass a casual code review.
Normalization over the full sample. Say your signal is a z-score: today's funding rate minus the mean, divided by the standard deviation. If you compute that mean and standard deviation over the entire history, every historical z-score contains information about the future, because the future data shaped the mean it's being compared against. A funding rate that looked extreme in 2021 relative to 2019-2021 data may look ordinary relative to 2019-2025 data. The fix is expanding or rolling windows: on each date, the statistics use only data available on that date. Same trap applies to anything fit on the full sample: regression coefficients, percentile ranks, volatility estimates used for sizing.
Restated and revised data. Fundamental data gets restated. Economic prints get revised, sometimes heavily. If your backtest uses the final revised GDP number on the date of the initial release, it's trading on a figure that didn't exist yet. Positioning data has a built-in version of this: COT data is reported as of Tuesday but published Friday afternoon, a lag we covered back in the futures lessons. A backtest that acts on Tuesday's positioning on Tuesday is three days ahead of anything you could have done. Point-in-time databases exist precisely to solve this, and they're expensive precisely because it matters.
The unsettled bar. If your data pipeline runs before a session closes, the latest bar in your database is provisional. The close is still forming, the daily volume is still accumulating, and a nightly job may overwrite a flag. Acting on that bar is look-ahead in disguise: you are trading a number that does not exist yet in final form. The discipline is to act only on settled bars, and to make the backtest respect the same publication timing the live system faces. If the live signal is computed at 6am on yesterday's close, the backtest must use yesterday's close too, not a same-day value that would have arrived hours after the trade.
Adjusted prices and dollar filters. Split-adjusted prices are fine for computing returns; that's what they're for. But applying a raw dollar filter to adjusted history, "only trade stocks above $10," misfires, because a stock trading at $300 today that split 10-for-1 twice shows adjusted historical prices far below what the tape actually printed. Your filter includes or excludes names based on prices nobody ever saw. Filters need to run on the prices that existed at the time.
The diagnostic for all of these is the same question: on the morning this trade would have been placed, was every number feeding the decision already published, final, and computed only from the past? Walk the data lineage of your signal and ask it of every input. A single input that looks fine but is not is enough to make the whole curve fiction.
There is also a smell test on the output side. Genuine edges in liquid markets are small. A daily-frequency strategy on a major index backtesting at a Sharpe of 4 is a bug to hunt, not a discovery. When a result looks too clean, the usual and usually correct assumption is leakage, and the burden of proof is on the code.
10.5.2 Survivorship bias: testing on the winners' roster
Take any equity strategy and test it on the current constituents of a major index over the past twenty years. The result is inflated before you write a single line of signal logic, because the current constituents are, by construction, companies that survived and grew enough to be in the index today. The names that went bankrupt, got delisted, or shrank into irrelevance aren't in your universe. Your backtest bought stocks in 2008 knowing, in effect, which ones would still exist in 2026.
The effect is large. Over any multi-decade window, a large fraction of listed US stocks delist: bankruptcies, acquisitions, going private, dropping below listing standards. A momentum or trend strategy tested only on survivors never has to live through the names that trended straight to zero. A value strategy on survivors buys every cheap stock that recovered and none of the cheap stocks that were cheap because they were dying. Both look better than any real portfolio could have.
Crypto is worse. The coins in today's data feeds are the ones still listed, still liquid, still clearing whatever volume or open interest threshold the data provider applies. The graveyard of dead projects, delisted pairs, and outright frauds is enormous, and it is invisible in the dataset. A cross-sectional crypto strategy backtested on currently listed coins has quietly assumed you would never have held any of the ones that vanished. In an asset class where going to zero is a routine outcome rather than a tail event, that assumption does most of the work.
The cure is point-in-time universe construction: on each historical date, the tradeable set is exactly what was listed, liquid, and index-included on that date, with delisted names carried through to their actual delisting return (often close to total loss). Proper point-in-time data is expensive and much of the free data you'll actually use doesn't have it. So the practical stance is honesty rather than purity: know whether your universe is point-in-time, and if it isn't, treat the result as a ceiling. Futures backtests suffer far less here. The major contracts have existed for decades and don't delist the way stocks do; a strategy on a fixed set of index, rate, and commodity futures dodges most of the problem. Single-name equity and crypto backtests carry it in full.
Survivorship applies to strategies as much as to assets. The fund databases used for performance comparisons quietly drop funds that shut down, and dead funds don't close because they were doing well. Any average return computed over currently reporting funds is skimmed from the top of the true distribution. The same mechanism operates inside your own research folder, which is where the multiple testing section below picks it up.
10.5.3 Costs: the fantasy of frictionless fills
A backtest with zero costs is a research artifact, not a trading result. The real costs are commissions, the bid-ask spread, slippage beyond the spread when your size moves the market, borrow fees on shorts, and, for perpetuals, funding paid or received every few hours simply for holding.
The damage comes from turnover, not any single cost. Suppose a strategy turns over its full portfolio once per day and pays 10 basis points round trip in spread and fees, which is optimistic for anything outside the most liquid instruments. That's 0.10% times roughly 250 trading days, a 25% annual drag. A signal that gross-earned 30% a year in the frictionless backtest nets 5%. The same signal traded weekly pays roughly a fifth of that toll. This is why slow strategies are slow on purpose: rebalancing a monthly momentum signal daily buys you almost no extra edge and charges you full freight in costs. Cadence is a cost decision.
The mid-price fantasy is worst in options. An option spread quoted 0.90 bid, 1.10 ask has a mid of 1.00, and a backtest that sells it collects 1.00 every time. You won't. Depending on liquidity you'll collect 0.92, 0.95 on a good day, and on wide illiquid strikes far less. On a structure you sell for a 1.00 credit and manage to a 0.50 debit, giving up 5 to 8 cents on each side of the round trip consumes a fifth or more of the theoretical edge. A serious short-vol backtest applies a punitive haircut to every mid-quoted credit, on the order of a third of the credit for retail-accessible spreads, and if the strategy still works after that, it might be real. If the edge only exists at mid, it doesn't exist.
Perpetual futures add funding. A long position in a perp during a euphoric stretch can pay tens of percent annualized in funding, and a backtest that models the price series without the funding series is modeling an instrument that doesn't exist. Funding cuts both ways, sometimes you receive it, but a backtest must include it either way, because for carry-flavored strategies funding is most of the P&L, not a small cost adjustment.
Slippage beyond the spread scales with your size relative to the market's depth, which connects back to the microstructure lessons: your market order eats levels of the book, and the deeper it eats, the worse your average fill. For small retail size in liquid products this is minor. For anything larger, or anything in thin markets, assume your fills are worse than the backtest by an amount that grows with size, which is one reason strategies degrade as capital scales.
The audit question is blunt: find the cost model and read it out loud. If the answer is "fills at mid, no commission, no funding," the Sharpe isn't tradeable and should be labeled research-only.
10.5.4 Accounting tricks: the flat-day Sharpe and friends
Some inflation happens after the returns are generated, in how they're summarized. The most common trick, sometimes deliberate and often innocent, is computing Sharpe only over the days the strategy held a position.
Sharpe is mean over standard deviation, annualized. Drop the flat days and you've removed a pile of zero-return observations. Removing zeros raises the mean per remaining day a lot and raises the standard deviation only somewhat, so the ratio jumps. If the strategy is in the market a fraction p of the time, the trade-days-only Sharpe overstates the calendar-time Sharpe by a factor of roughly 1 over sqrt(p). A strategy in the market a quarter of the time gets its Sharpe roughly doubled by this one accounting choice, with not a dollar of extra profit anywhere.
The comparison it corrupts is the one you care about: strategy versus benchmark. Buy-and-hold is measured over every calendar day. A strategy measured only over its active days is playing a different game with a smaller denominator. The rule is simple and non-negotiable: the return series feeding a Sharpe includes every calendar trading day, with flat days entered as zeros. If someone hands you a Sharpe, ask whether the flat days are in it. If they can't answer, you have your answer.
Two relatives of the same trick. Cherry-picked windows: a backtest that starts in March 2009 or January 2019 has been positioned, consciously or not, to begin at a generational low. Ask what the result looks like started two years earlier or later; a real edge does not depend on the start date. And compounding presentation: showing a log-scale equity curve when the linear one would reveal that 80% of the profit came from one three-month window in one instrument. Concentration of P&L in a single episode is not automatically damning, but it changes the sample size of the evidence from "ten years" to "one event," which loops back to everything the sample-size lesson said about luck.
10.5.5 Overfitting: the model that memorized the past
Every backtest fits the past to some degree. Overfitting is when the fit captures noise instead of structure, and the tell is that performance collapses on data the rule never saw.
The mechanism is degrees of freedom. Every parameter you tune, every filter you add, every special case you code in ("skip December 2018, that was weird") gives the strategy another way to contort itself around historical accidents. With enough knobs, you can fit anything. A rule with two parameters that made money across thirty markets is telling you something about markets. A rule with nine parameters that made money on one market is telling you something about your optimizer.
The practical test is the plateau. Take whatever lookback or threshold the strategy uses and nudge it. If a 14-day lookback works, do 12 and 16 work? If entry at a z-score of 2.0 works, does 1.8? Does 2.2? A real effect is a broad plateau: a whole neighborhood of parameter values that all make money, some a bit more, some a bit less, because the underlying behavior (trend persistence, premium harvesting, positioning extremes mean-reverting) doesn't care about your exact number. A spike, where 14 days prints a Sharpe of 1.6 and 12 days prints 0.3, is the signature of noise. Nothing in market structure changes that abruptly between a 12-day and a 14-day window; only noise does.
Plateau vs Spike
A related discipline is preferring rules with a reason. A signal that says "buy when large speculators are at a three-year positioning extreme against commercials" has an economic story: someone has to pay to shed risk, crowding resolves, and you know who the counterparty is. A signal that says "buy when the 17-day average crosses the 43-day average but only on Wednesdays" has no story, only a fit. Stories can be wrong, but a rule with a mechanism behind it has a chance of persisting, because the mechanism constrains what parameters even make sense before you touch the data. A rule discovered by search has only the data, and the data contains mostly noise.
10.5.6 Multiple testing: why the best of fifty means nothing
This problem survives even clean code and honest costs. You test fifty configurations. Forty-nine are mediocre. One prints a Sharpe of 1.1. You trade the winner.
You haven't found an edge. You've run a lottery and picked the winning ticket after the draw.
Here is the arithmetic. A backtested Sharpe is an estimate with sampling error, and the error is bigger than intuition suggests. For a strategy with no true edge, the standard error of an annualized Sharpe estimated from daily data is roughly 1 over the square root of the number of years. Ten years of data: standard error around 0.32. That means a genuinely worthless strategy, tested once over ten years, will usually print a Sharpe somewhere between roughly -0.6 and +0.6 just from luck, and about one test in twenty lands outside even that band.
Now test many worthless strategies and keep the best. The expected maximum of n draws from a normal distribution grows like the square root of 2 ln n. Run fifty independent zero-edge configurations over ten years and the expected best Sharpe among them is around 0.7. Run a few hundred and the best of the batch can plausibly print near 1.0. Nothing worked. Nothing had any edge. The maximum of many noisy estimates is high by construction, and it's exactly the number you selected for.
Best of Fifty Worthless Strategies
The honest counting is harder than it sounds, for two reasons pulling in opposite directions. Correlated trials count for less than one each: fifty variants that are all the same trend rule with slightly different lookbacks might amount to five effective independent tests, not fifty. And people undercount what they tried. The count that matters isn't the trials in your final notebook, it's every variant you looked at and discarded along the way, including the ones you abandoned after a glance at the equity curve, including the ideas you tested last year on the same data and forgot. The market data you keep re-mining doesn't reset between your projects. This is the research version of the file-drawer problem: the losers go in the drawer, the winner goes in the deck, and the deck says "backtested Sharpe 1.1" with a straight face.
Which is why "I tested 50 configs and one worked" isn't evidence; it's the null hypothesis behaving exactly as expected. If anything it's mild evidence against the idea: if the underlying effect were real, you'd expect many of the fifty variants to work, a plateau across the batch, not one lucky spike.
The defenses are procedural, not mathematical, because the math can't save you after the fact if you didn't keep count.
- Pre-commit the hypothesis. Write down the rule, the parameters, and the reason it should work before you run the test. One pre-committed test on ten years of data is worth more than a hundred searched ones, because its Sharpe means what it says.
- Keep the parameter count brutal. Two or three, chosen for a reason, tested at round values. Not a grid search over the integers.
- Demand the plateau, per the previous section. A batch where most variants work is evidence; a batch where one works is noise.
- Keep a research graveyard. A log of every idea tested and rejected, with a note not to re-test it. It stops you from re-mining the same noise and re-counting an old lucky draw as a fresh discovery, and it keeps your trial count honest.
- Test across markets. A rule that made money on twenty futures contracts with the same parameters has faced twenty semi-independent juries. A rule fit to one instrument has faced one, and you chose that instrument after looking.
10.5.7 The deflated Sharpe: raising the bar for the number of tries
The multiple-testing logic has a formal version, worth understanding even if you never compute it exactly. It's called the deflated Sharpe ratio, and it answers a precise question: given how many strategies were tried, how varied their results were, how long the sample is, and how non-normal the returns are, what's the probability that the best observed Sharpe exceeds what pure luck would have produced?
The procedure, in plain steps. First, from the number of effectively independent trials and the spread of their Sharpes, compute the expected maximum Sharpe under the assumption that every trial was noise. That's the hurdle, and as the previous section showed, with hundreds of trials it can sit near 1.0 rather than at zero. Second, ask how many standard errors the winning strategy's Sharpe sits above that hurdle, where the standard error accounts for the sample length and gets wider when returns are skewed and fat-tailed, which, per the fat-tails discussion earlier in this part, they always are, and especially so for short-vol strategies whose smooth curves hide occasional violence. The output is a probability that the edge is real rather than selected noise.
Two things follow from the shape of that calculation. The hurdle rises with the number of trials: a Sharpe of 1.0 from a single pre-committed test can be strong evidence, while the same 1.0 as the best of five hundred configurations is roughly what noise predicts. And the penalty for skew means strategies with occasional large losses (short volatility, short gamma, carry) need a higher observed Sharpe to clear the same bar as a symmetric strategy, because their standard errors are wider than the normal-distribution math assumes. The strategies most likely to seduce you with a smooth backtest are precisely the ones the correction hits hardest.
You don't need to run the formula on every idea. You need its two reflexes: every reported Sharpe should arrive with a trial count attached, and a Sharpe without a trial count is uninterpretable, not merely weak evidence.
10.5.8 Out-of-sample, walk-forward, and the only test that cannot be gamed
The standard defense against overfitting is the train-test split: build the rule on the first seven years, test it untouched on the last three. If performance holds on data the rule never saw, that's real evidence. The mechanics matter: the out-of-sample period must be genuinely untouched, including by your eyeballs, and the universe, costs, and timing conventions must be identical across both periods.
The weakness is human, not statistical. The first time you test on the holdout, it's out-of-sample. Then the result disappoints, you tweak the rule, and test again. And again. After five iterations, the holdout isn't out-of-sample anymore; you've fit to it through the feedback loop of your own decisions, just more slowly than a direct optimization would have. Each peek spends the holdout's evidential value, and it doesn't regenerate. The walk-forward variant, refitting on a rolling window and always testing on the next unseen chunk, is sturdier because it simulates the actual experience of running the strategy through time, but it too can be silently iterated into an in-sample exercise if you keep adjusting the process after seeing the results.
Which leaves the one test that can't be gamed even in principle: the forward record. Fix the rule, put it live (real money or a rigorously honest paper account with realistic fills), and from that day on, every return is out-of-sample by the arrow of time. The future hadn't happened when the rule was frozen, so nothing about the rule can have been fit to it. This is why the live record and the backtest are different kinds of object: the backtest is the hypothesis, the live curve is the experiment.
Reading the comparison between them is a skill of its own. Live performance modestly below backtest is the normal, honest outcome: it's the survivorship discount, the cost reality, and the mild overfitting all showing up on schedule. Live performance far below backtest, with no regime excuse, means the backtest was more corrupted than you knew. And live performance well above backtest isn't good news, it's a warning: the most common explanation is that something in the comparison is broken, and leakage somewhere in the pipeline is high on the suspect list. The two curves should live in the same neighborhood, and you should know why they differ where they do. Expect months of forward data before the comparison says much at all; the sample-size math from earlier in this part applies to live records too, and it isn't kind.
10.5.9 Auditing a backtest in practice
All of the above compresses into a short interrogation you can run on any backtest, yours or anyone else's. It takes ten minutes and it's worth more than any amount of admiring the equity curve.
Where did the universe come from, and is it point-in-time? If the answer involves the word "current," survivorship is in and the number is a ceiling.
Where does the signal meet the return? Find that exact line of code and confirm the lag. Confirm every input was published and final at the moment of the decision, including revisions and reporting lags.
What's the cost model? Read the numbers. Mid fills and zero commission mean research-only. For options, look for a credit haircut. For perps, look for funding. Check that the backtest's rebalance cadence matches what would actually be traded live.
Is the Sharpe calendar-time? Zeros for flat days, every trading day in the denominator, a start date that wasn't chosen for effect.
How many things were tried? Ask for the graveyard. A researcher who can't list their failed variants isn't hiding them from you, they're hiding them from themselves, which is worse. Discount the headline accordingly.
Is it a plateau or a spike? Nudge the parameters and watch what happens. Ask whether the same rule works on neighboring markets.
Is there a forward record? If yes, weight it far above the backtest. If no, everything above determines how much benefit of the doubt the simulation deserves, and the default answer is: some, never much.
None of this makes backtesting useless. A carefully built backtest is how you separate ideas worth risking money on from ideas that merely sound good, and the discipline of building one forces you to specify a strategy precisely enough to criticize. The point is calibration: a backtest is a hypothesis with a number attached, the number is biased upward by construction, and knowing exactly which biases are present is what lets you discount it intelligently instead of either worshipping it or throwing it away.
A clean, honest, survivorship-discounted, cost-realistic, multiplicity-adjusted backtest still leaves the biggest question untouched: how big to trade it. A true Sharpe of 0.8 can compound into wealth or blow up an account depending entirely on sizing, and sizing starts with measuring risk in the right units. That's the next lesson: why risk is measured in volatility rather than dollars, and what follows from taking that seriously.
10.6 Risk in volatility units
Here are two positions. Position one: $10,000 of a regulated utility stock that moves about 15 percent a year. Position two: $10,000 of a mid-cap semiconductor name that moves about 45 percent a year. Same dollars. Same line on your broker statement. If you think of these as the same size, you're measuring the wrong thing, because the second one is three times the bet.
Most retail sizing runs on dollars. "I put 10k into it." "I never risk more than 5k per trade." Dollars are what the broker shows you, what the margin call is denominated in, and what you eventually eat or spend, so it feels natural to count risk in them. But dollars measure exposure, not risk, and the gap between those two words is where a lot of accounts quietly die. This lesson builds the alternative: measuring every position in units of its own volatility, so a bet on natural gas, a bet on a bond ETF, and a bet on an altcoin can all be compared on one scale and sized so each one hurts about the same amount when it goes against you.
Everything in the rest of this part sits on top of this idea. Vol targeting in the next lesson, the Kelly discussion after that, and the full sizing chain later all assume you've already stopped thinking in dollars and started thinking in volatility units. So this is the lesson to get right.
10.6.1 The problem with dollar sizing
Exposure answers the question "how much money is in this position." Risk answers "how much is this position likely to move my account per day, per week, per month." Those are different questions because instruments differ in how violently they move. A dollar of short-term treasury ETF and a dollar of a leveraged crypto perp are the same exposure and wildly different risk.
Put numbers on the opening example. Back in the realized volatility lesson you met the rule of 16: annualized volatility divided by 16 gives you daily volatility, because there are about 256 trading days in a year and the square root of 256 is 16. The utility at 15 percent annualized vol moves about 15 / 16, call it 0.9 percent, on a typical day. On a $10,000 position that's roughly $90 of daily wobble. The semiconductor name at 45 percent annualized moves about 2.8 percent a day, roughly $280 on the same $10,000. Hold both and the semiconductor position dominates your daily P&L three to one, even though your statement says you're "equally invested" in each.
Now scale that up to a whole book. A trader running ten equal-dollar positions believes they're diversified ten ways. If two of those names are high-vol growth stocks and eight are boring dividend payers, the honest description is that they're running a concentrated two-name book with some low-vol filler attached. The dollar allocation lies about where the risk lives. We'll quantify exactly how badly it lies later in this lesson, but in a mixed-vol equal-dollar portfolio the most volatile sleeve routinely accounts for the majority of the portfolio's variance while holding a minority of the capital.
There's a second, sneakier failure of dollar thinking: it makes your risk drift over time without any decision on your part. The same $10,000 in the same stock is a different bet in a calm summer than in the week after an earnings warning, because the stock's volatility changed while your dollar count didn't. Dollar sizing means your actual risk is set by the market's mood rather than by you. Vol-based sizing hands that dial back.
10.6.2 Volatility as the common currency
The fix is to measure every position in the same unit: expected movement. Two definitions do most of the work.
Instrument volatility is how much the thing itself moves, expressed as a percentage of its price. You can quote it daily or annualized; they convert through the rule of 16 (or the square root of 252 if you want the exact trading-day count, the difference is cosmetic). A stock with 32 percent annualized vol moves about 2 percent on a typical day. This number belongs to the instrument, not to you. It's the same whether you own one share or ten thousand.
Cash volatility (or dollar volatility) is what that movement means for your account: cash_vol = exposure x instrument_vol. In plain terms, take how many dollars you have riding on the thing and multiply by how much it moves in percent, and you get how many dollars your position swings. A $20,000 position in the 2-percent-a-day stock has a daily cash vol of about $400. That $400 is the number that matters to you. It's the typical size of the daily mark-to-market swing this position feeds into your account.
Once every position is translated into cash volatility, they all live on one scale. $400 a day of Apple risk, $400 a day of crude oil risk, and $400 a day of ETH risk are comparable quantities in a way that "$20,000 of Apple, two crude contracts, and 5 ETH" never will be. The instruments have nothing in common; their cash volatilities are the same kind of number. Risk measured in volatility units is portable across asset classes, and a swing trader touching equities, futures, and crypto (which is exactly who this platform is built for) needs a portable unit more than anyone.
A note on what "typical day" means here. Daily volatility is a standard deviation, so roughly two thirds of days land within one cash vol of zero and about 95 percent within two, if returns were normal. They aren't normal, as the fat-tails discussion earlier in this part made clear, and crypto in particular produces days that a normal distribution says should never happen. Treat cash vol as a good estimate of ordinary conditions and a floor, not a ceiling, on what a bad day can do. We come back to this caveat at the end, because it's the limit of everything in this lesson.
10.6.3 Measuring it: standard deviation and ATR
You need a number for instrument volatility before you can size with it. Two estimators cover practically all real usage.
10.6.3.1 Standard deviation of returns
The default: take the last N daily returns (20 trading days, about one calendar month, is the common window), compute their standard deviation, and that's your daily vol. Multiply by 16 for the annualized figure. This is the same realized vol you met in the options lessons, doing double duty. The number the platform shows as RV 20d for an equity is exactly this quantity, which means the sizing input for a stock is already sitting on its volatility page.
The choice of window is a tradeoff you should make consciously. A short window (10 to 20 days) reacts fast: when a stock's behavior changes, your size adjusts within a couple of weeks. It's also noisy, so your position sizes jump around, which costs you commissions and slippage in rebalancing. A long window (6 months, a year) is stable but slow, and slow is dangerous in the one situation that matters most: it keeps telling you an instrument is calm well after it has stopped being calm. A 20 to 30 day window, or a blend of a short and a long window, is where most systematic sizing lands. The key point: reacting too slowly to rising vol is the expensive mistake, and reacting too slowly to falling vol just costs you some upside.
One subtlety: close-to-close standard deviation only sees where the market ended each day. A stock that gaps down 4 percent and recovers to close flat registers as a quiet day. For sizing purposes that quiet day wasn't quiet, which brings us to the second estimator.
10.6.3.2 Average true range
ATR asks a more physical question: how many dollars (not percent) does this thing travel in a typical day, counting gaps?
The building block is the true range for a single day, which is the largest of three distances: high minus low, the absolute distance from today's high to yesterday's close, and the absolute distance from today's low to yesterday's close. In plain terms, it's the full span the price covered since yesterday's close, so an overnight gap counts even if the day's own range was narrow. ATR is then a moving average of true range, classically over 14 days, though 20 works the same way and lines up with the monthly window used for return vol.
ATR comes out in price units. A $50 stock with an ATR of $1.50 travels about $1.50, or 3 percent, on a normal day. Because it's denominated in dollars per share (or points per contract), ATR plugs directly into share-count arithmetic without a units conversion, which is why swing traders like it: stop placement and position sizing both come out in one step.
10.6.3.3 Which one to use
Mostly it doesn't matter, and anyone telling you one is dramatically superior is selling something. For a liquid instrument, ATR expressed as a percentage of price and the standard deviation of daily returns track each other closely, and sizing built on either produces nearly the same positions. ATR runs slightly higher on gap-prone instruments since close-to-close vol misses the gaps, and slightly different in trending periods, but these are second-order effects.
Pick based on convenience. If your workflow is stops and share counts on individual stocks, ATR is already in the right units. If your workflow spans asset classes and percent returns (which is where this course is heading), standard deviation of returns is the cleaner primitive, because percent vol composes: it annualizes with the rule of 16, it feeds portfolio math, and it's what every vol figure on this platform is quoted in. A practical note if you build your own tools: when you only have closing prices and no intraday data, price times daily return vol is a serviceable stand-in for ATR. It gives the same relative sizing across instruments, and relative sizing is all the risk-parity logic below actually needs.
10.6.4 Sizing from volatility
With a vol estimate in hand, sizing becomes one line of algebra. There are two equivalent framings; use whichever fits the instrument.
10.6.4.1 The dollar-vol framing
Decide how much daily cash volatility you want a single position to contribute, then solve for size:
Divide the daily dollar swing you want by the daily dollar swing of one unit, and you get how many units to hold.
Worked example. You run a $100,000 account and decide each position should contribute about $150 of daily cash vol. Stock A trades at $80 with a 2 percent daily vol, so one share swings about $1.60 a day. Size: 150 / 1.60 = 93.75, call it 93 shares, about $7,400 of exposure. Stock B trades at $80 with a 0.8 percent daily vol, so one share swings $0.64. Size: 150 / 0.64 = 234 shares, about $18,700 of exposure. Same price, wildly different dollar allocations, identical risk. The calm instrument gets the bigger dollar slice because it needs more dollars to matter.
The $150 choice implies 15 basis points of equity per position per day at the account level. Ten such positions, if they were independent, would give the account a daily vol somewhere near 0.5 percent (risks add in quadrature, not linearly, and correlation pushes the true figure around). Choosing that account-level number properly is the vol targeting problem, and it gets its own lesson next. Here we only need the per-position mechanics.
10.6.4.2 The ATR framing
Identical logic, ATR units:
where risk_factor is the fraction of equity you assign per ATR of movement. A widely used value in systematic equity momentum is 0.001, ten basis points: on a $100,000 account, each position is sized so that a one-ATR move changes the account by about $100.
Worked example at that setting. Stock A: $50 price, $1.50 ATR. Shares = 100,000 x 0.001 / 1.50 = 66 shares, about $3,300 of exposure. Stock B: also $50, but a $4.00 ATR. Shares = 100 / 4 = 25 shares, $1,250 of exposure. A typical day moves either position about $100. The wild stock gets less than half the capital of the calm one, and your P&L stops depending on whichever holding happens to be jumpiest.
| Position | Price | Vol measure | Shares | Dollar exposure | Daily cash vol |
|---|---|---|---|---|---|
| Dollar-vol A | $80 | 2.0% daily vol | 93 | ~$7,400 | ~$150 |
| Dollar-vol B | $80 | 0.8% daily vol | 234 | ~$18,700 | ~$150 |
| ATR A | $50 | $1.50 ATR | 66 | ~$3,300 | ~$100 |
| ATR B | $50 | $4.00 ATR | 25 | ~$1,250 | ~$100 |
Within each framing the two positions carry the same daily cash vol despite very different dollar exposures: the calm instrument gets far more capital to land at identical risk. That equality of the last column, not of the exposure column, is the whole point of sizing in volatility units.
Two practical notes. Round share counts down, not to the nearest integer; systematic sizing errs small. And rerun the calculation on a schedule (weekly is plenty for swing horizons) rather than in a panic, because vol estimates move every day and chasing them daily just burns commissions.
10.6.4.3 Stop-based sizing and its limits
The most common sizing rule in retail trading looks superficially similar: risk a fixed fraction of the account to the stop. "I risk 1 percent, my stop is 5 percent away, so the position is 20 percent of my account." This is better than nothing and much better than pure gut feel, but it has two structural problems that vol sizing doesn't.
The stop distance is a choice, and the formula rewards bad choices. Tighten the stop from 5 percent to 2 percent and the same rule now tells you to put 50 percent of the account into the position. Your measured "risk" stayed at 1 percent while your actual exposure went up two and a half times. The market doesn't care where your stop is; a 3 percent overnight gap hits the 50 percent position for 1.5 percent of your account regardless of the stop sitting 2 percent away. Stop-loss risk and position risk are different quantities, and only one of them is under your control.
And stops aren't guaranteed exits. Gaps, halts, weekend crypto moves, and limit-down futures sessions all deliver fills well beyond the stop price. The 1 percent number is the minimum loss conditional on the stop being hit cleanly, not the maximum loss.
Set stop distances in volatility units rather than in percent-of-price or chart feel. A stop 2 ATRs away scales automatically with the instrument's behavior, wide on wild things and tight on calm things, and then stop-based sizing and vol-based sizing collapse into the same calculation. Fixed-percent stops on instruments with different vols are either too tight (you get shaken out by ordinary noise on the volatile ones) or too loose (you give back too much on the calm ones). The mechanics of trailing those stops through a trade belong to the sizing chain lesson later in this part; here the point is only that the stop distance itself should be a vol quantity.
10.6.5 Instrument vol versus position vol
So far every example was unleveraged stock, where exposure and capital committed are the same dollars. Derivatives break that equality. One more distinction is needed: the volatility of the instrument is not the volatility of your position in it.
Instrument vol, as defined above, is the percent volatility of the underlying's returns. Position vol is the volatility of your account equity caused by the position, and it scales with leverage:
Take the instrument's own volatility and multiply it by how many times your capital you've deployed. Unleveraged, the ratio is at most 1 and position vol is at most instrument vol. With derivatives the ratio can be 5, 10, 50, and position vol inflates in exact proportion.
Futures make this concrete. Take an equity index future at 6,000 points with a $50 multiplier, so one contract controls $300,000 of notional. Suppose the index runs 15 percent annualized vol, roughly 0.9 percent a day. One contract therefore swings about $2,800 on a typical day (300,000 x 0.15 / 16). On a $100,000 account, that single contract is a position vol of 45 percent annualized: the instrument is a placid 15-vol index, but your position in it is three times your capital, so your equity experiences it as a 45-vol asset. Meanwhile the margin requirement is a small fraction of notional. Margin is what the exchange demands as a deposit, and it has nothing to do with risk. Plenty of traders size futures by "how many contracts can my margin support," which is like sizing a mortgage by the minimum down payment. The sizing question is not whether you can open the position. It is what the position does to your equity per day, and only notional times vol answers that.
Crypto perps are the same arithmetic with bigger numbers and a UI that actively encourages the mistake. The leverage slider on a perp exchange sets margin, not risk. A coin running 60 percent annualized vol moves about 3.75 percent a day. At 10x leverage, your position vol is 600 percent annualized, which in daily terms means your margin swings about 37 percent on an ordinary day, not a crash. The liquidation engine you met in the perpetuals lessons exists precisely because position vol at those ratios routinely exceeds the collateral behind it. When you hear that someone got liquidated on a 5 percent move, the instrument did nothing unusual; their position vol was simply set to a level where a one-and-a-bit sigma day was fatal.
The clean way to run leverage is to invert the formula. Decide the position vol you want, then let leverage fall out:
If you want a 20 percent position vol on that 60-vol coin, notional is capital x 0.20 / 0.60, one third of your capital, no leverage needed at all. If you want 20 percent position vol on a 5-vol short-term bond future, notional is four times capital, and leverage is the tool that gets you there. Leverage is not a return amplifier you dial up when confident. It is the mechanism that lets low-vol instruments reach a useful position vol. High-vol instruments never need it. The instruments where exchanges offer the most leverage are exactly the ones where using it is least defensible.
- Instrument vol
- 60% annualized
- Leverage
- 1.0x (unlevered)
- Notional
- $33,000
- Instrument vol
- 5% annualized
- Leverage
- 4.0x
- Notional
- $400,000
10.6.6 The case for equal risk over equal dollars
Everything above sized one position at a time. The strongest argument for volatility units appears when you put several positions side by side, because equal-dollar allocation fails in a way you can compute exactly.
Take a $100,000 account split into four equal $25,000 positions: a bond ETF at 5 percent annualized vol, a large-cap stock at 18, a gold miner at 35, and a crypto position at 70. To keep the arithmetic honest and simple, assume the four are uncorrelated (real correlations make equal-dollar allocations look better in calm times and worse in crashes, which is a story for the drawdown lesson).
Each position's contribution to portfolio variance is its weight times its vol, squared: (w x sigma)^2. Variance (vol squared) is the quantity that adds across independent positions, so squaring each position's cash vol and summing tells you how much of the total each one owns. Run the numbers:
| Position | Weight | Vol | (w x sigma)^2 | Share of portfolio variance |
|---|---|---|---|---|
| Bond ETF | 25% | 5% | 0.000156 | 0.4% |
| Large-cap stock | 25% | 18% | 0.002025 | 5.0% |
| Gold miner | 25% | 35% | 0.007656 | 18.9% |
| Crypto | 25% | 70% | 0.030625 | 75.7% |
The crypto position holds a quarter of the dollars and three quarters of the risk. The bond ETF is functionally not in the portfolio: 0.4 percent of the variance means that if you deleted it entirely, the equity curve would barely notice. What the owner of this book believes is a diversified four-asset portfolio is, in risk terms, mostly a crypto account. Every equal-dollar portfolio containing mixed volatilities has this shape to some degree; the highest-vol sleeve eats the risk budget because the contribution goes up with the square.
Equal Dollars, Unequal Risk
Now rebuild the same book with inverse-volatility weights: each position weighted proportional to 1 / sigma, then normalized so the weights sum to one. The calm asset gets a big slice, the wild one a small slice, and each ends up contributing the same cash vol. For these four instruments the weights come out near 67 percent bonds, 19 percent large-cap, 10 percent miner, and 5 percent crypto, and each position's w x sigma lands around 3.3 percent, dead equal by construction. Now every position matters and none dominates. A crypto drawdown hurts in proportion to its risk share, a quarter of the book's variance, instead of three quarters of it.
The crypto slice is about $5,000 against the $25,000 the equal-dollar version held. This is the answer to a question every multi-asset trader eventually asks: "how much crypto should I hold next to my equities?" In dollars the question has no principled answer. In risk units the answer is mechanical: enough that its cash vol matches the risk you want it contributing, which at crypto volatilities is always far fewer dollars than intuition suggests.
Equal risk is the right default for a second reason beyond balance: it's an honest statement about what you know. Weighting positions by conviction assumes you can rank your ideas by future performance, and the backtesting lesson earlier in this part should have left you skeptical of that. Weighting them by dollars assumes nothing and delivers an accidental concentration. Weighting them by inverse vol assumes only that your vol estimates are roughly right, and vol is by far the most forecastable property of financial returns: far more predictable than direction, more stable than correlation. Vol clusters (calm follows calm, storm follows storm, per the realized vol lesson), which is exactly the property an input needs to be worth sizing on. Equal risk is what "I don't know which of these will do best" looks like when written as an allocation.
There's a legitimate refinement where conviction does enter sizing, scaled and capped so it can't blow the framework up, and that's the forecast machinery covered in the sizing chain lesson later in this part. But conviction there modulates a risk-based size; it never replaces the risk measurement. Get the default right first.
10.6.7 What volatility sizing does not fix
Sizing in vol units is the largest single upgrade available to a discretionary trader's risk process, and it would be dishonest to leave you thinking it's sufficient. Four limits, each of which later lessons pick up.
Volatility is an estimate, and it's an estimate of the recent past. A 20-day window through a quiet market says the instrument is calm, and the sizing formula obediently hands you a large position, maximum size arriving exactly when the market has been at its most sedate. Calm periods precede violent ones often enough that this is a real failure mode, not a technicality. The practical mitigations are unglamorous: don't let low vol estimates push size past a hard cap, consider blending a longer window in as a floor, and never let the formula override the max-risk-per-trade rule that the semi-systematic lesson later in this part treats as non-negotiable.
Volatility is symmetric and tails are not. Standard deviation charges the same price for upside and downside movement and assumes tomorrow resembles the recent sample. Fat tails, which this part opened with, mean the worst days are much worse than the vol estimate implies, and for short-vol and carry-type positions the return distribution is skewed so that the vol looks low right up until the one day that defines the year. Vol sizing handles ordinary movement, not the rare extreme. Tail events get handled by structure (defined-risk trades, hard caps, diversification across return streams), not by a bigger lookback window.
Position-level risk is not portfolio-level risk. Everything here treated positions one at a time, and the four-asset example leaned on an independence assumption that real markets violate on the worst days, when correlations lurch toward one and every position becomes the same position. Sizing each position to equal risk is the foundation; deciding what the whole book's risk should be, and scaling everything to hit it, is the vol targeting problem, and correlation is its failure mode. Both are next lesson's subject.
And vol says nothing about edge. A perfectly risk-balanced portfolio of bad trades loses money smoothly. Volatility units tell you how big; the rest of the course is about what and when.
One question remains open: how much total volatility the whole book should run, and what happens to that machinery when vol spikes and every position starts moving together. Sizing each position to equal risk builds the parts, not the machine. That's volatility targeting, next.
10.7 Volatility targeting
The last lesson ended with a rule: size every position in volatility units, so a quiet bond future and a violent crypto perp each contribute a similar daily swing to your account. That rule fixes the relative sizes. It says nothing about the absolute level. You can hold ten positions, each carefully equalized in risk, and still be running the whole book at triple the volatility you can stomach, or at a third of the volatility your edge deserves. Something has to set the overall dial.
Volatility targeting is that dial. You decide, in advance and in writing, how volatile your account should be, expressed as an annualized standard deviation of returns. Then you scale your gross exposure up when markets are calm and down when they're wild, so the account's realized volatility stays near the number you chose. That's the entire idea. The rest of this lesson is what the number should be, how the scaling works mechanically, why the resulting equity curve is better than the one you get from fixed dollar sizing, and the specific way this machinery fails in a crisis, because it does fail, and the failure has a shape you can prepare for.
Most traders never make this decision at all. They trade one contract because they always trade one contract, or they put 10 percent of the account in each idea because ten positions felt diversified. Under fixed sizing like that, the market decides how much risk you run. When realized volatility triples, your risk triples, at exactly the moment you least want it to. Volatility targeting takes that decision away from the market and gives it back to you. It is a policy rather than a technique: the account has a risk budget, and positions spend from it.
10.7.1 What a vol target is
A vol target is a single number: the annualized standard deviation you want your account returns to run at. Write it as a percentage of account equity. A 20 percent vol target on a $100,000 account means you intend the account's yearly return to have a standard deviation of about $20,000.
The rule of 16 from the realized volatility lesson converts that into something you can feel. Annual volatility divided by 16 is roughly daily volatility, because there are about 252 trading days in a year and the square root of 252 is just under 16. A 20 percent annual target is a typical daily move of about 1.25 percent, so around $1,250 on the $100,000 account. Ordinary days will be smaller, bad days will be two or three times that, and the occasional horror will exceed even the bad days, because returns have fat tails and every sigma-based statement in this part carries that asterisk. But as a planning number: 20 percent annual, 1.25 percent daily, and a monthly standard deviation of about 5.8 percent (divide the annual figure by the square root of 12).
The target is not a loss limit and not a guarantee. It's a statement about the width of the distribution you're choosing to sit inside. If your strategy has positive expectancy, a wider distribution means faster growth and deeper drawdowns; a narrower one means slower growth and shallower drawdowns. The target is where you pick your point on that tradeoff, explicitly, instead of inheriting whatever point your position count happens to imply.
The word gets used at two levels. You can vol target a single position (scale a BTC position so it contributes 10 percent annualized to the account) and you can vol target the whole book (scale everything so the combined account runs at 20 percent). The mechanics are the same division at both levels. This lesson mostly works at the book level, because that's where the decision lives, and the lesson near the end of this part picks up the multi-sleeve version, where combining imperfectly correlated strategies lets you run more gross exposure for the same book-level target.
10.7.2 Picking the number
The target has to clear three tests at once: your edge can support it, your account can survive it, and you can live with it.
Start with the edge. There's a hard relationship, derived properly in the next lesson, between the quality of a strategy and the maximum volatility it makes sense to run it at. For a strategy with roughly normal returns, the growth-optimal vol target equals the strategy's Sharpe ratio expressed as a percentage. A Sharpe of 0.5 supports, at the theoretical maximum, a 50 percent vol target. Beyond that point, more risk actively reduces long-run compound growth: past the peak, added volatility subtracts from your compounded return instead of adding to it, so the extra pain now buys negative reward. That theoretical maximum is also a cliff edge you should stay far away from, for reasons the Kelly lesson makes concrete, but it sets a ceiling. A trader who believes, after honest accounting, that their edge is a Sharpe around 0.5, and who runs a 60 percent vol target, is over the cliff even if every input to that belief is correct. And the belief is never correct: Sharpe estimates from backtests and short live records are noisy and biased upward, which is one more reason the ceiling isn't a target.
Then the survival test, which is about drawdowns. Drawdown depth scales with the vol target. Run the same strategy at double the vol and the drawdowns roughly double while arriving in the same places. Useful rough calibration for strategies with a modest, realistic edge: expect to see drawdowns comparable to your annual vol number as a routine event, and treat a drawdown approaching twice the vol number as unwelcome but unremarkable over a decade of trading. At a 20 percent target that means 20 percent drawdowns are part of the deal and something near 40 percent is possible in a bad stretch. If reading that sentence made your stomach drop, your target is lower than 20. The drawdown lesson later in this part does the recovery arithmetic in full; here the point is only that the vol target is where you buy your future drawdowns, in advance, at a price you set.
Then the living test. Convert the target to a daily dollar figure with the rule of 16 and ask whether you can watch that number appear with a minus sign in front of it on a normal Tuesday without changing your behavior. The test is not whether you can survive it financially, but whether you can see it and still follow the system. A 25 percent target on $200,000 is a typical day of about $3,100, with days of $6,000 to $9,000 against you arriving several times a year. Plenty of traders who could afford those numbers can't trade through them, and a target you abandon in the first real drawdown was never your target; it was a number in a spreadsheet.
Where does that leave the actual choice? For an individual trading their own money with a real but modest edge, something in the 10 to 25 percent range covers almost everyone. For calibration: a plain long position in a broad equity index runs at roughly 15 to 20 percent volatility in normal times, so a 20 percent target means "as bumpy as holding stocks," which most people can picture. Institutional multi-strategy books often run 10 to 15 percent. Numbers above 30 percent are for strategies with demonstrated high Sharpe, or for small accounts whose owners have explicitly decided the account is risk capital they can lose. If you have no strong basis for the choice, I would start at 15 percent: high enough that a real edge produces meaningful returns, low enough that the inevitable bad year is survivable and the estimate errors in your Sharpe belief don't put you over the cliff.
10.7.3 Scaling a position to the target
The mechanics are one division. To run a single instrument at your target:
In plain terms: how many dollars of the thing you must hold so that its typical wiggle, applied to that many dollars, produces your target wiggle on the account. If the instrument is exactly as volatile as your target, hold one times capital. If it's half as volatile, hold twice capital, which means leverage. If it's four times as volatile, hold a quarter of capital.
Run the numbers for a $100,000 account with a 20 percent target:
| Instrument | Typical annualized vol | Notional to hit 20% target | Exposure as multiple of capital |
|---|---|---|---|
| 10-year note future | 5% | $400,000 | 4.0x |
| Broad equity index | 16% | $125,000 | 1.25x |
| Crude oil | 35% | $57,000 | 0.57x |
| Bitcoin | 60% | $33,000 | 0.33x |
| Small-cap altcoin | 120% | $17,000 | 0.17x |
The table restates the last lesson's point in dollar form: equal-risk positions are wildly unequal in dollars. It also shows something new: hitting a sensible vol target on a quiet instrument requires leverage, and that's fine. Leverage isn't risk; volatility is risk. Four times capital in treasury futures is a calmer position than one times capital in bitcoin, and the entire derivatives complex from Part 2 exists partly so that traders can dial notional up and down independent of capital. This should cure the reflex that reads "4x levered" as reckless and "unlevered bitcoin" as prudent. Measured in the units that matter, it's the other way around.
The same division runs the whole book. Compute the realized volatility of your account returns (not of any instrument, of the account itself, since that number already contains all your positions and the correlations between them), then:
and multiply every position by the scale. A book targeting 25 percent that's realizing 12.5 percent scales everything to 2x. The same book realizing 50 percent scales to 0.5x. The signal driving each position hasn't changed; the size riding on the signal has. This is what stops a strategy from silently tripling its risk just because the market got choppy, and equally what stops it from wasting a low-vol regime by running at half its budget.
Two guards belong in the mechanics from day one, because the raw formula misbehaves at both extremes.
Cap the scale. In a dead-calm market the division asks for enormous leverage: a book realizing 5 percent against a 25 percent target wants 5x gross. Refuse it. Something like a 2x or 3x maximum on the scaling multiplier is standard, and the reason isn't squeamishness about leverage in general. It's that a very low vol estimate is exactly the estimate most likely to be wrong soon. Volatility has a floor around which it compresses during long calms and from which it jumps, and the jump arrives faster than any estimate can track. Uncapped, the formula maximizes your leverage at the precise moment a jump would do the most damage. The cap is you telling the formula that you know something about vol dynamics that a trailing standard deviation doesn't.
Refuse thin estimates. A vol estimate computed from two weeks of returns is noise. If you don't have enough observations to estimate the book's volatility with any confidence, run at scale 1.0 (or below) until you do. Never let a formula lean hard on a number it barely knows.
10.7.4 Estimating the volatility you divide by
Everything above divides by a volatility, and that volatility is a forecast, not a fact. You're not asking what the vol was; you're asking what it will be over the horizon you'll hold the scaled position. The good news, established back in the realized volatility lesson, is that this is the most forecastable object in trading. Volatility clusters: turbulent days follow turbulent days, calm follows calm, and the persistence is strong enough that "tomorrow will look like the recent past" is a genuinely good forecast. Nothing comparable is true of returns. Vol targeting works at all because you're steering by the one gauge on the dashboard that actually predicts its own future.
The estimation choice is a tradeoff between speed and stability. A short lookback (10 to 20 days) reacts quickly when the regime changes but jumps around on noise, and a single wild day distorts it for its whole window. A long lookback (6 to 12 months) is stable but stale, still reporting calm weeks into a storm. The standard resolution is an exponentially weighted estimate, which weights recent days most and fades older ones smoothly, reacting in days rather than months without the cliff effects of a short fixed window. A pragmatic alternative that captures most of the benefit: blend a short-window and a long-window estimate, and when they disagree sharply, trust the higher one. Vol spikes are fast and vol declines are slow, so an estimator that's quick to raise its number and slow to lower it errs on the survivable side.
What you shouldn't use is anything approaching the instrument's full-history volatility. Bitcoin's lifetime vol tells you about bitcoin's lifetime. The position you're sizing lives in next month, and next month resembles last month far more than it resembles 2017.
Then there's the question of how often to act on the estimate. Recompute daily; that costs nothing. But don't trade every twitch of the number, because a scaling system that adjusts positions on every 3 percent drift in estimated vol will bleed its edge into commissions and spread. The standard fix is a buffer: leave the position alone until it drifts meaningfully from the freshly computed ideal, something like 10 percent of the position size, then trade back to the ideal. For a swing trader the practical cadence is a weekly resize plus an immediate one whenever the vol estimate moves by a large step. You resize weekly in normal times and daily in crises, which is exactly the cadence the situation calls for.
10.7.5 What vol targeting does to the equity curve
Hold a strategy's signals fixed and switch only the sizing, from fixed dollars to vol targeted, and the equity curve changes character in three ways.
The swings become uniform. Under fixed sizing, an equity curve alternates between dead stretches where nothing you do seems to matter and stretches of terror where every day is a week's worth of P&L. Those aren't changes in your edge; they're changes in market volatility passing straight through constant notional. Vol targeting flattens that passthrough. Your P&L distribution stops being a mixture of regimes and becomes roughly one distribution, which means drawdowns now come from your strategy being wrong rather than from being wrong while accidentally huge. It also makes your own statistics readable: a bad month at constant risk is information about the edge, while a bad month under fixed sizing might just be information about VIX.
In most risk assets, scaling down when volatility rises also dodges the worst returns. High volatility and bad returns arrive together in equities, credit, and crypto. The mechanism runs both directions: falling prices make markets more volatile (deleveraging, panic, forced selling), and volatile markets frighten holders into selling. The result, visible across long histories of index data, is that the most volatile periods carry the poorest average returns, so mechanically scaling exposure down as volatility rises has historically improved the risk-adjusted returns of plain equity exposure rather than merely smoothing them. You give up some participation in sharp recoveries and get paid back by being small through the worst clusters of losses. It's a free-ish lunch with a real cost, itemized in the failure section below. But the direction of the historical evidence is clear: for assets where vol spikes coincide with sell-offs, vol targeting has been additive.
And there's the compounding arithmetic. The performance lesson introduced volatility drag: compound growth is roughly the average return minus half the variance, so episodes of extreme volatility eat growth out of proportion to their length. A strategy that spends most of its life at 12 percent vol and occasional months at 60 percent suffers drag dominated by those few months. Holding vol near a constant removes the episodes that do the most compounding damage. Two return streams with the same average return, one at steady 20 percent vol and one averaging 20 percent through wild swings between 8 and 50, don't grow the same: the steady one compounds faster. Vol targeting converts the second stream into something closer to the first.
A side benefit: once everything you run is vol targeted, comparison becomes honest. Two strategies at the same vol target differ only in the quality of their returns, so their equity curves can be read against each other directly, and the Sharpe comparisons from two lessons ago stop being confounded by size.
10.7.6 Where it fails
Everything above is true and none of it survives first contact with a crisis unmodified. Vol targeting has a failure mode, it's well understood, and you should be able to recite it before you run the system, because the system will meet it.
10.7.6.1 The estimate lags the spike
Volatility doesn't glide from 15 to 45. It jumps. In early 2018 the main equity volatility index more than doubled in a single session. In the pandemic crash of 2020, equity index volatility went from the low teens to crisis levels inside three weeks, with individual days that exceeded entire normal months. No trailing estimate, however cleverly weighted, sees a jump before it happens; the estimate reflects the past and does not see ahead. So the sequence in every genuine vol shock is: you take the first hit at full size (or worse, at capped-leverage size, since spikes tend to arrive out of calm regimes where your scale was at its maximum), your estimate catches up over the following days, and you cut. Vol targeting protects you in the second week of a crisis, not on the first day. Sizing against the first day is what the leverage cap, the modest target, and the tail-risk material elsewhere in the course are for. If you sized your book such that only the vol-targeting machinery stands between you and ruin on a gap, you sized it wrong, and no estimator tuning fixes that.
10.7.6.2 You sell weakness, mechanically
During a sell-off in a long book, the scaling rule instructs: vol rises, so you cut, which means selling into falling prices, near what may be the lows. If the market then V-bottoms, you ride the recovery at reduced size and buy your exposure back at higher prices. That round trip is a real cost, not a bug: it's the premium you pay for the smoothing. Some crashes keep crashing, and in those the same mechanical cut is what saves the account. You can't have the protection in the crashes that continue without paying the cost in the ones that reverse. What you can do is refuse to let the machinery whipsaw you at high frequency (the trade buffer earns its keep here) and refuse the temptation to override the cut because "it always bounces." The one time it doesn't bounce is the time the rule existed for.
10.7.6.3 Correlations converge
This layer connects forward to the portfolio lesson at the end of this part. A book's volatility isn't the sum of its positions' volatilities; it's dampened by every imperfect correlation between them. A book of ten positions, each individually sized to contribute 8 percent, might realize only 15 percent as a book because the positions disagree with each other often enough to cancel. Your vol targeting operates on that 15.
In a crisis, two things happen at once: every instrument's own volatility rises, and the correlations between risk assets converge toward one. Equities, credit, commodities, crypto, and most carry trades become a single trade called "risk," and the cancellation your book-level vol estimate was built on evaporates. The book's realized volatility therefore jumps by more than any single instrument's does: you get hit by the vol spike and the correlation spike multiplied together. A book that honestly measured 15 percent in the old regime can be a 45 percent book in the new one before you've changed a single position. The diversified book was, in part, a bet on calm continuing, and the vol number it reported was a property of the regime as much as of the book.
The practical consequence: in stress, cut faster and deeper than the instrument-level arithmetic suggests, because the book-level number is deteriorating on two axes at once. If you track one leading indicator for this, make it the vol term structure from the regime lessons: when the front of the curve inverts above the back, the market is pricing the regime break before your realized estimates can see it, and pre-emptively pulling scale below what the formula says is a defensible override. It's one of the few overrides I endorse in this course, because it's itself a rule, not a feeling.
10.7.6.4 Everyone else is doing it too
Vol targeting isn't a niche retail trick. Enormous pools of institutional money run some version of it: volatility-control funds, risk parity allocations, annuity products with embedded vol caps, dealer hedging programs. When volatility spikes, all of them receive the same instruction from the same arithmetic at the same time: reduce. Their selling raises volatility further, which generates more reduction instructions, and the loop feeds itself until the leverage that accumulated during the calm has been flushed. The 2018 episode mentioned above was substantially this loop running at full speed, and the blowup lesson at the end of this part dissects it properly. For your purposes the implication is that vol spikes are sharper and faster than an innocent reading of history suggests, because the machinery reacting to them is now a large fraction of the market. The lag problem from the first subsection is worse than it used to be, and calm regimes end more abruptly. Set your leverage cap accordingly.
None of this argues against vol targeting. Every alternative fails worse: fixed sizing takes the same first-day hit and then keeps full size through the entire crisis; discretionary de-risking does whatever your adrenaline says. The argument is against believing the target. Your realized vol won't equal your target vol; it'll hum near the target in normal times and overshoot badly for short stretches around regime breaks. Choose a target such that the overshoot stretches are survivable, cap the leverage the calm regimes tempt you into, and treat the machinery as plumbing that keeps risk roughly constant most of the time, not a guarantee that it's constant all of the time.
10.7.7 Running it in practice
Here is the full system, small enough to actually operate, for a swing trader running the kinds of strategies in Part 9.
Fix the target once, in writing, using the three tests: below the ceiling your honest Sharpe estimate implies, drawdowns you can survive, daily swings you can watch. Suppose 15 percent on $100,000. That's a typical day near $940 and a routine bad stretch that may approach a $15,000 drawdown.
Each position enters at a size that contributes its planned share of the budget, using the instrument-vol division from this lesson and the per-trade risk mechanics from the last one. Once a week, compute the account's realized vol from your own daily equity changes, exponentially weighted or short-and-long blended. Divide target by realized, cap the result at 2x, and if the answer differs from your current gross by more than the buffer, adjust every position proportionally. If your equity history is too short to trust the estimate, run at 1x or below and let the record accumulate.
Add the two crisis clauses: any day the vol estimate steps up sharply, resize immediately instead of waiting for the weekly pass, and when the vol term structure inverts, take scale below the formula's answer until it normalizes. Both clauses are rules. Write them down with the target, because the moment they trigger is the moment you'll least feel like inventing them.
Realized Vol Hums, It Does Not Lock
10.7.8 When a simpler stop is enough
Everything in this lesson is the advanced version of position risk, and it is worth being honest that most traders should not start here. The strategies earlier in the course used a plainer tool: a stop placed a fixed multiple of the daily ATR away from price, trailed as the trade moves your way, with the exit taken only on a daily close through the level. That is cruder than scaling a whole book to a volatility target, but it is often enough, and for a concrete reason. Volatility targeting needs inputs a newer trader simply does not have yet: an honest Sharpe estimate, a realized-volatility history of your own equity curve, a feel for how your positions move together under stress. Guess those numbers before you have lived them and the machinery just launders the guesses into false precision. An ATR stop needs none of it. It caps the loss on each trade in the market's own units and asks nothing about your long-run statistics.
So the honest sequence is to trade the ATR-stop version first, keep the records this part keeps asking for, and graduate to full volatility targeting only once you have enough live data to estimate the inputs it consumes, or when you are running a fully systematic book where a volatility target is the natural sizing rule from the start. The advanced tool is better when its inputs are real. Until then, a simple stop you actually follow beats a sophisticated sizing rule fed with numbers you made up.
The target you chose in this lesson came from rules of thumb: a ceiling set by your Sharpe estimate, a floor set by usefulness, and a survival check in between. A more precise claim is available: a formula that takes an edge and a variance and returns the exact size that maximizes long-run growth. It's elegant, it's correct under its assumptions, and almost nobody who runs it at full strength holds on through the drawdowns it produces. The next lesson works through the Kelly criterion and why the right answer in practice is a deliberate fraction of it.
10.8 The Kelly criterion
Here's a game. I flip a fair coin. Heads, your bet grows by 60 percent. Tails, it loses 50 percent. You can play as many rounds as you like, and you must choose what fraction of your bankroll to stake each round before we start.
The expected value of one round is positive: 0.5 x (+60%) + 0.5 x (-50%) = +5% per flip. The tempting move is to bet everything, every time, since every flip is a positive-EV proposition. So run it. Bet 100 percent of your bankroll each round and flip the coin many times. Half the flips multiply your wealth by 1.6, half multiply it by 0.5, and after one of each you hold 1.6 x 0.5 = 0.8 of what you started with. Two flips, 20 percent gone, and the order doesn't matter. Keep playing and your median outcome grinds toward zero at about 10.6 percent per flip (the per-flip growth factor is sqrt(1.6 x 0.5) = sqrt(0.8), about 0.894). A game with positive expected value ruins almost everyone who plays it at full size.
That's the puzzle the Kelly criterion resolves. Somewhere between betting nothing (no growth) and betting everything (guaranteed decay) there's a fraction that makes your money compound as fast as it possibly can. Kelly is the formula for that fraction. This lesson covers what it is, what it actually optimizes, why nobody with functioning nerves runs it at full size, and how to use a de-rated version of it as the theoretical backbone of every sizing rule you've met so far in this part.
10.8.1 Arithmetic returns lie to compounders
The coin game fails at full size because of one distinction: the arithmetic mean of your returns is not the rate your wealth compounds at.
Arithmetic mean answers "what is the average return of a single round." Geometric mean answers "at what steady rate does repeated play grow my money." For anyone who reinvests, and every trader managing an account reinvests by default, the geometric mean is the only one that pays. The two are connected by an approximation worth memorizing:
Growth rate equals average return minus half the variance. In plain terms, volatility is a direct tax on compound growth. Two strategies with the same average return don't grow your account at the same speed; the choppier one grows it slower, and if the chop is bad enough, a positive-average strategy compounds to nothing. The coin game puts numbers on that: mu is +5% per flip, but the variance of a full-bankroll bet is enormous, and the sigma^2 / 2 penalty eats the entire edge and then some.
Betting a fraction instead of everything changes the balance. Staking f of your bankroll scales the return of each flip by f, which scales mu by f but scales variance by f squared. The edge shrinks linearly while the volatility tax shrinks quadratically. Small bets keep most of their edge and shed almost all of their drag. That asymmetry is why an intermediate fraction wins: as f rises from zero, growth first rises (edge accumulating faster than drag), peaks, then falls (drag accumulating faster than edge), and eventually goes negative. Kelly is the top of that hill.
The Kelly Hill
10.8.2 The formula
For a simple repeated bet where you win b times your stake with probability p and lose your stake with probability q = 1 - p, the growth-maximizing fraction is:
which rearranges to the more memorable form:
In plain terms, take your expected profit per dollar staked (the edge) and divide by what a win pays per dollar staked (the odds). The formula came out of mid-century work on information theory, got proven in practice at blackjack tables and racetracks, and migrated into finance from there. It answers exactly one question: what fixed fraction of my bankroll, bet repeatedly, maximizes the long-run compound growth rate of my wealth. Equivalently, it maximizes the expected logarithm of wealth. Those are the same statement, and the log framing matters later.
Work it. You have a bet that wins 55 percent of the time and pays even money (b = 1). Then f* = (1 x 0.55 - 0.45) / 1 = 0.10. Bet 10 percent of the bankroll per round. Run the growth arithmetic and full Kelly earns about 0.50 percent of compound growth per bet: g = 0.55 x ln(1.10) + 0.45 x ln(0.90), which comes out just over 0.005. Half a percent per flip doesn't sound like much until you remember it compounds without limit; that's the fastest this particular edge can be turned into wealth by any betting scheme whatsoever.
Second example, closer to a trade. A setup wins 50 percent of the time, winners pay 2R, losers cost 1R. Then b = 2, p = q = 0.5, and f* = (2 x 0.5 - 0.5) / 2 = 0.25. Kelly says risk 25 percent of your account on this trade. That number is your first hint of the problem with full Kelly. A coin-flip trade with a respectable 2:1 payoff, the kind of setup swing traders describe as bread and butter, and the growth-optimal stake is a quarter of everything you have, on one position. Every risk rule you've ever been given screams at that number, and the rules are right, for reasons the rest of this lesson makes precise.
A few properties of the formula are worth noting. No edge means no bet: if b x p = q, then f* = 0, and if the edge is negative, Kelly says the optimal stake is zero (or the other side, if you can take it). Kelly never tells you to bet without an advantage, which already makes it more disciplined than most traders. Because you always stake a fraction of your current bankroll, losses automatically shrink your bets and wins grow them. After a losing streak a Kelly bettor is risking fewer dollars, not more. Compare that to the doubling-down instinct, which does the exact opposite and converts losing streaks into ruin. And with divisible stakes Kelly can never take you to literal zero, since it always leaves 1 - f* on the table. Strict ruin is impossible. Drawdowns that feel indistinguishable from ruin are not, as you'll see.
10.8.3 The trading version
Trades aren't binary bets. Returns are continuous, so the formula gets restated in return space. For a strategy or instrument with expected excess return mu (over cash) and volatility sigma, the growth-optimal fraction of capital is:
In plain terms, divide the strategy's edge by its variance, and that's how much of your capital to deploy into it, where numbers above 1 mean using leverage. Same shape as the binary formula: edge in the numerator, a risk measure in the denominator.
Put numbers in. A strategy earns 10 percent a year over cash at 20 percent annualized vol. Then f* = 0.10 / 0.04 = 2.5. Kelly says run it at two and a half times leverage. At that setting the portfolio's volatility is 2.5 x 20% = 50 percent a year, and its expected growth rate works out to about 12.5 percent.
There is a cleaner way to say this. Divide through and you find that at full Kelly, your portfolio volatility equals your Sharpe ratio:
And the growth rate you earn there is Sharpe squared over two. Kelly converts the vol targeting question from the previous lesson into a statement about strategy quality. A strategy with a Sharpe of 0.5, which is a genuinely good systematic strategy run over real costs, has a full-Kelly vol target of 50 percent annualized. A Sharpe of 1.0, which almost nobody sustains at scale for long, justifies 100 percent. Set those numbers against the 10 to 25 percent vol targets that serious traders actually run: everyone sane operates far below full Kelly, and the rest of the lesson explains why.
One technical footnote. The continuous formula assumes you rebalance back to the target fraction continuously and that returns are roughly normal. Real rebalancing is periodic and real returns have fat tails, so treat f* = mu / sigma^2 as the idealized ceiling, not a setting to dial in. Both caveats get their own sections below.
10.8.4 What full Kelly actually feels like
Full Kelly maximizes long-run compound growth. That statement is mathematically airtight and psychologically almost meaningless, because "long run" is doing an enormous amount of work.
Start with the drawdown arithmetic. Under the idealized continuous model, a full Kelly bettor's probability of ever seeing their bankroll fall to a fraction x of its starting value is simply x. Probability of a 50 percent drawdown at some point: one half. Probability of a 90 percent drawdown: one in ten. Not in a catastrophe, not because the edge died, but as the routine operating experience of the growth-optimal strategy working exactly as designed. A full Kelly account spends a large share of its life deep underwater relative to its own high-water mark, and the swings scale with the wealth: the drawdowns in year ten are as violent, proportionally, as the ones in year one.
Then there's the dispersion of outcomes. Kelly wins the long run with probability one, meaning that over an infinite horizon the full Kelly bettor ends up richer than any other fixed-fraction bettor. Over the horizons a human career actually contains, the distribution of outcomes is brutally wide. Two traders running identical full-Kelly strategies for a decade can end that decade with wealth differing by an order of magnitude, purely on path. The median outcome is strong; the experience of getting there involves stretches, sometimes years, where a half-Kelly or quarter-Kelly version of the same strategy is beating you, and you have no way to know from inside the drawdown whether you're living bad variance or a dead edge. That ambiguity matters. Back in the thinking-in-distributions lesson you saw how long bad variance can run, and full Kelly maximizes your exposure to it.
There is a subtler point. Kelly is optimal for a bettor whose satisfaction with wealth is logarithmic: someone who genuinely feels that going from 100k to 200k is worth the same as going from 1M to 2M, and who would accept a coin flip between those doublings and a matching halving with indifference at the right odds. Almost nobody is actually built that way. Most people feel losses far more than the log says they should, need to withdraw money to live, have careers and investor relationships that don't survive 80 percent drawdowns, and can't reinvest with the frictionless perfection the model assumes. If your true tolerance for pain is lower than logarithmic, and it is, then full Kelly is over-betting relative to your own preferences even when the inputs are perfect. The formula isn't wrong; it's answering a question about a person who doesn't exist.
10.8.5 Fractional Kelly
The fix is to scale the framework, not abandon it. Bet a fixed fraction c of the Kelly stake, with c somewhere between a tenth and a half, and the math turns sharply in your favor.
Under the standard approximation, betting c times Kelly earns you (2c - c^2) of the maximum growth rate at c times the volatility. The numbers deserve a table:
| Fraction of Kelly | Share of max growth | Share of full Kelly vol |
|---|---|---|
| 1 (full) | 100% | 100% |
| 3/4 | 94% | 75% |
| 1/2 | 75% | 50% |
| 1/4 | 44% | 25% |
| 1/8 | 23% | 12.5% |
Half Kelly keeps three quarters of the growth for half the volatility. That trade is so lopsided that half Kelly is close to a professional consensus as an upper bound for anyone betting real money on estimated edges. The drawdown math improves even faster than the table suggests: at half Kelly, the probability of ever losing half your bankroll drops from 50 percent to about 12.5 percent (the general result under the idealized model is x raised to the power 2/c - 1, so moving from full to half Kelly turns a drawdown probability of x into x cubed). Quarter Kelly makes deep drawdowns rarer still while keeping nearly half the growth. The region between quarter and half Kelly is where the growth-versus-pain tradeoff becomes tolerable enough to run as a business.
The growth curve is asymmetric, and that asymmetry matters more than anything else here in practice. Around the peak, the curve is roughly symmetric: betting 80 percent of Kelly and betting 120 percent of Kelly cost you about the same small slice of growth. But the risk isn't symmetric at all. The underbettor gets less growth and less volatility; the overbettor gets less growth and more volatility, strictly worse on both axes. Push further and it gets uglier: at exactly twice Kelly, expected compound growth falls all the way back to zero (check it on the 55 percent coin: betting 20 percent instead of 10 gives g = 0.55 x ln(1.20) + 0.45 x ln(0.80), which is a hair below zero). Beyond twice Kelly you are in the coin game from the top of the lesson, grinding a positive edge into a shrinking bankroll. A trader with a real, persistent edge who sizes at two and a half times Kelly will go broke slowly and be genuinely baffled about why, because every individual trade was positive EV.
So the operating rule follows from the geometry: when uncertain, err low, because the cost of underbetting is mild and the cost of overbetting compounds. Every reason in the next section says you're more uncertain than you think.
10.8.6 Estimation error
Everything so far assumed you know p and b, or mu and sigma. You don't, and the gap between the edge you estimate and the edge you have is the strongest argument for deep-fractional Kelly.
Volatility is the friendly input. Vol is persistent and reveals itself quickly; a few months of data pins sigma down well enough for sizing, which is why the vol-unit lessons could be so confident. The mean is the hostile input. The standard error of an estimated annual return is sigma divided by the square root of the number of years observed. A strategy with 20 percent vol observed for 25 years gives you a standard error on the mean of 20 / sqrt(25) = 4 percentage points. Twenty-five years of clean data, and your 10 percent estimated edge is honestly "somewhere between about 2 and 18." Nobody has 25 years of clean data on their own strategy. On the two or three years a live track record typically spans, the confidence interval on your edge comfortably includes zero. Your Kelly fraction inherits every bit of that uncertainty, because mu sits right there in the numerator.
It gets worse, because on top of the noise your estimate carries an upward bias. Back in the backtesting lesson you saw why: strategies get promoted to live trading because they tested well, and testing well is partly luck, so the edge that survived selection is on average smaller than the edge that was measured. Feed a selection-inflated mu into f* = mu / sigma^2 and you compute a full Kelly stake for a strategy better than the one you own. You believe you're at full Kelly; you're actually beyond it, on the wrong side of the peak, in the region where more aggression means less growth. This single mechanism, honest Kelly math on top of dishonest inputs, is a plausible description of how a large fraction of confident, hardworking traders destroy accounts.
The response can be formalized cleanly. If you treat your edge as a distribution rather than a point estimate and maximize growth over that distribution, the optimal stake comes out below the naive Kelly of the point estimate, and the wider your uncertainty, the further below. Parameter uncertainty and fractional Kelly lead to the same place. Betting half Kelly is roughly what an honest bettor does when they take seriously the possibility that their edge is half what they measured.
10.8.7 Fat tails and the short-vol trap
The second silent assumption in f* = mu / sigma^2 is that sigma describes the risk. For fat-tailed return streams it understates it, and back in the statistics lesson you saw that nearly everything traded on this platform is fat-tailed, with crypto as the extreme case. When the tails are fatter than normal, the true growth-optimal fraction is lower than the formula's output, because the sigma^2 / 2 drag term underestimates what rare large losses do to a compounding account. Same direction as estimation error: the naive number is too high, shade down.
The trap is sharpest for negatively skewed strategies, and this course spends a lot of time on exactly those. The concave sleeves from the strategy lessons, short VRP, short earnings vol, funding capture, share a signature: high win rates, small steady gains, occasional violent losses. Run naive Kelly arithmetic on a short-vol strategy's track record and the output is absurdly aggressive, because the sample mean is flattered by a string of wins and the sample sigma is calm precisely because the catastrophic tail hasn't shown up in the data yet. The strategy's realized history systematically understates both inputs' danger until the one month that repairs the record. Kelly sized off the calm years is maximum leverage into the crash. The asset-class rule: whatever Kelly fraction you'd tolerate on a symmetric strategy, cut it again, hard, for anything short vol. The performance-measurement lesson made this point about Sharpe ratios flattering short vol; the same distortion flows straight through Sharpe into Kelly, since full Kelly vol equals Sharpe.
One more real-world subtraction: simultaneous positions. Kelly fractions aren't additive across bets unless the bets are independent. Ten positions that are 60 percent correlated are closer to one big bet than to ten small ones, and running each at its individual Kelly fraction puts the aggregate book far beyond full Kelly at the portfolio level. Correlations across your book, especially their habit of converging in stress (the drawdown lesson coming up makes this concrete), mean portfolio Kelly is well below the sum of position Kellys. Yet another argument pointing the same direction.
10.8.8 What Kelly is actually for
After all that demolition you might conclude Kelly is useless. It is not. Kelly is close to useless as a literal position-size calculator, and close to indispensable as three other things.
As a ceiling: whatever sizing method you use, compute the rough Kelly fraction for the bet and check you're comfortably under it. The 2R coin-flip trade from earlier had a full Kelly of 25 percent risk per trade. The classic fixed-fraction rule of risking 1 to 2 percent per trade is, for setups in that quality range, somewhere around one-tenth Kelly or less. That's not timidity; given estimation error, fat tails, and correlated positions, a true tenth of naive Kelly is probably a quarter to a half of honest Kelly, which is exactly the sensible band. The old 1 percent rule turns out to be deep-fractional Kelly in practical form, and it's reassuring when folk wisdom and theory converge on the same number from opposite directions.
As the theory underneath your vol target: the previous lesson had you choose a portfolio vol target and admitted the choice was partly judgment. Kelly sharpens it: full Kelly vol equals Sharpe, so a trader who honestly expects a long-run Sharpe around 0.5 has a Kelly ceiling of 50 percent annualized vol, and quarter-to-half Kelly puts the sensible target between about 12 and 25 percent. If you're running a 15 percent vol target on a strategy you believe has a Sharpe of 0.5, you're implicitly at about 30 percent of Kelly, which is a defensible, even conservative, place to live. If you're running 40 percent vol on the same belief, you're at 80 percent of Kelly and betting that your Sharpe estimate has no error in it. The framework won't pick your number, but it tells you what your number claims about your edge, and most traders have never once checked whether their sizing and their honest edge estimate are even in the same universe.
And as a direction-of-error rule, the one the growth curve's asymmetry hands you for every sizing decision: the penalty for betting too small is linear and mild, the penalty for betting too big is compounding and fatal, so every doubt about your edge resolves toward smaller. This is also the quantitative backbone of the overbetting point the semi-systematic lesson will hammer shortly: the danger of overbetting is mathematical before it's psychological, because it's the one sizing error capable of turning a winning strategy into a losing account.
Kelly gives you the theoretical maximum and the shape of the penalty for exceeding it; what it doesn't give you is a way to translate a specific trade idea, with its own conviction level and stop placement, into a live position in a real account. That translation layer, from forecast strength through volatility units to contracts and shares, with rules for trailing stops and adding to winners, is a framework of its own. That's the next lesson.
10.9 The position sizing chain
The last three lessons handed you the pieces. You know why risk is measured in volatility units rather than dollars. You know what a portfolio vol target is and why scaling to it smooths your equity curve. You know what Kelly optimizes and why nobody sane runs it at full strength. This lesson bolts those pieces into a single machine: a complete, number-in-number-out procedure that takes your account size, your risk appetite, the instrument's volatility, and your conviction, and hands back an exact position size and an exact exit rule. Nothing left to feel out in the moment.
The framework is built for a specific kind of trader, and there's a good chance it's you. You pick your own trades. You read positioning data, you watch the momentum indicator, you spot the setup on the chart, and you decide to be long crude or short a tech name. What you don't have is a backtested, fully automated system that generates entries for you. Traders in this position are usually strong at trade selection and terrible at everything after the entry: they size up when they feel sure, hold losers out of hope, cut winners out of fear, and go too big after a hot streak. The semi-automatic approach splits the job cleanly. You keep the part you're good at (choosing what to trade and which direction) and hand the parts that destroy accounts (sizing and exits) to fixed rules.
The deal has three clauses. You decide what to trade, which direction, and how strong your conviction is. The system decides how large the position is and when you exit. Neither of you overrides the other. Ever. The moment you override the system once, you've got a different system, one with an untested discretionary exception built in, and the exception will always fire at the worst possible time because that's exactly when you'll want to use it.
One more point: this framework needs no backtest. Most sizing advice assumes you have a historical track record of your exact strategy, which a discretionary trader never has. Everything here works from first principles instead: the instrument's measurable volatility, a risk budget you choose deliberately, and a conviction score you assign honestly. Those three inputs are enough.
10.9.1 Quantifying conviction: the forecast
Every trade starts with a number called the forecast. Before you enter anything, you assign your conviction a value on a fixed scale from -20 to +20. Positive means long, negative means short, and the magnitude says how strong the setup is. Zero means no trade.
The scale is anchored so that +10 (or -10) is your average trade. Not your weakest, not your best: the ordinary, decent setup you see regularly. +20 is reserved for the handful of trades a year where everything lines up: multiple timeframes agree, positioning is at an extreme, there's a catalyst, and the setup is one you rarely see. +5 is a speculative punt with limited confirmation. The same ladder runs on the short side with negative numbers.
| Forecast | Meaning |
|---|---|
| +20 | Maximum conviction long. Everything lines up, rare setup |
| +15 | Strong conviction long. Clear setup with confirming signals |
| +10 | Average conviction long. Decent setup, some confirmation |
| +5 | Weak conviction long. Speculative, limited confirmation |
| 0 | No position |
| -5 to -20 | Same ladder, short side |
Why bother scoring conviction at all instead of just trading fixed size? Because your conviction genuinely does contain information, and throwing it away is wasteful. A setup where COT positioning is at a three-year extreme, seasonality agrees, and momentum has just turned is a better bet than a lone chart pattern, and it deserves more capital. The forecast is how that judgment enters the math in a controlled way, instead of leaking in as "I feel really good about this one so I'll triple my size." The scale converts an emotion into a bounded input. At +20 you hold exactly twice the average position, never five times it.
Four rules keep the forecast honest.
Most of your trades should be +10 or -10. If you look back at a month of trades and see a string of +18s and +20s, you're not finding exceptional setups every week; you're grade-inflating your own homework. A useful calibration question: how often do you see this setup? If the answer is every week, it's a +5 to +10. If the answer is a few times a year, it might earn +15 to +20.
Once set, the forecast never changes for the life of the trade. You don't bump +10 to +15 because the trade is working, and you don't cut it to +5 because you're nervous. If you allow mid-trade forecast changes, you've reintroduced emotional sizing through the back door, which is the exact disease this framework exists to cure.
The forecast affects position size and nothing else. It doesn't change your stop. A +5 trade and a +20 trade on the same instrument use the identical trailing stop rule; the only difference is how many units you hold. Conviction buys you size, not room.
And be honest at entry, because there's no reward for optimism. A forecast that's too high doesn't make the trade work more often. It just makes the loss bigger when it doesn't.
10.9.2 Measuring the instrument's volatility
The forecast is your input. The market's input is volatility, and you already know from the realized volatility lesson back in the options part how to measure it. Here you need it in a specific, practical form: the expected dollar move of one unit of the instrument on a typical day.
Start with annualized realized volatility, which the platform shows you directly for the instruments it covers. A refinement worth using: blend a short window with a long one, so your estimate reacts to current conditions without being whipsawed by them.
In plain terms: weight the last month or so of volatility at 70 percent because recent vol is the best predictor of near-term vol (vol clusters, as covered earlier), but keep a 30 percent anchor to the long-run level so a freakishly quiet or freakishly wild month doesn't fully take over your sizing.
If you want a quick estimate by hand instead, take the last 25 daily absolute returns, average them, and multiply by 16 to annualize. This runs a little below a proper standard-deviation estimate, since averaging absolute moves understates the spread of a fat-tailed series, but it's close enough to size from and fast to redo.
From annual vol, get daily vol by dividing by 16:
The 16 is the rule of 16 from the realized volatility lesson: roughly 256 trading days in a year, and sqrt(256) = 16. Purists use sqrt(252) = 15.87; the difference is noise.
Then convert to dollars:
This number is the workhorse of the whole system. It says: on an ordinary day, one unit of this thing moves about this many dollars.
Take BTC at $68,400 with 20-day RV at 55 percent and 252-day RV at 45 percent:
Blended vol: 0.7 x 55% + 0.3 x 45% = 52% annualized
Daily vol %: 52% / 16 = 3.25%
Daily vol $: 3.25% x $68,400 = $2,223
On a typical day, expect BTC to move about $2,223 in one direction or the other. Everything downstream (position size, stop distance) is denominated in this unit.
One operational rule: don't recalculate your vol estimate daily. Only update it when it moves more than about 25 percent from the value you're using. Vol estimates wobble day to day, and if you re-derive your position size every morning you'll generate a stream of pointless small adjustments that cost commissions and attention. Weekly checks are plenty. A 25 percent threshold means you re-size when the risk picture has genuinely changed, not when the estimator twitched.
10.9.3 The chain itself
The position sizing chain converts four inputs (capital, vol target, instrument vol, forecast) into a position, in five steps.
Step 1: Annual cash vol target = Capital x Vol target % Step 2: Daily cash vol target = Annual cash vol target / 16 Step 3: Daily vol per unit = Daily vol % x Price per unit Step 4: Vol scalar = Daily cash vol target / Daily vol per unit Step 5: Position = (Vol scalar x Forecast) / 10
Here is what each step means; the chain is only useful if you understand why each link exists.
Step 1 states your risk appetite in dollars. With $100,000 and a 25 percent vol target you're saying: I accept my account value swinging by roughly $25,000 over a year. You chose this number in the vol targeting lesson; now it goes to work.
Step 2 converts that to a daily figure by dividing by 16, the same rule of 16 running in reverse. $25,000 of annual volatility is about $1,562.50 of daily volatility. That's what your whole position should move on a typical day.
Step 3 is the instrument's contribution: how much one unit (one BTC, one share, one futures contract) moves per day in dollars. For futures, remember to multiply by the contract multiplier from the specs lessons; one ES point is $50, so a 40-point daily vol is $2,000 per contract, not $40.
Step 4 divides your daily risk budget by the instrument's daily risk per unit. The result, the vol scalar, is the anchor of the entire system: the number of units you hold at average conviction. At forecast +10, you hold this many units, and your position's daily move equals exactly your daily risk budget.
Step 5 scales the anchor by conviction. Dividing the forecast by 10 turns the scale into a multiplier: +10 gives you exactly the vol scalar, +20 doubles it, +5 halves it, -10 gives you the vol scalar short.
Run the full chain for BTC with a $100,000 account and a 25 percent vol target, using the $2,223 daily vol from above:
Step 1: $100,000 x 0.25 = $25,000 annual cash vol Step 2: $25,000 / 16 = $1,562.50 daily cash vol Step 3: 3.25% x $68,400 = $2,223 daily vol per BTC Step 4: $1,562.50 / $2,223 = 0.703 BTC (the vol scalar) Step 5 at forecast +10: (0.703 x 10) / 10 = 0.703 BTC Step 5 at forecast +15: (0.703 x 15) / 10 = 1.054 BTC Step 5 at forecast +20: (0.703 x 20) / 10 = 1.406 BTC
The calibration checks out. At forecast +10 you hold 0.703 BTC, and 0.703 x $2,223 = $1,563 of daily dollar volatility, which is your daily cash vol target to the dollar. The system is self-consistent: an average-conviction position contributes exactly the risk you budgeted, no more and no less.
The chain also self-adjusts, which is what makes it worth the arithmetic. If BTC's volatility doubles next month, the daily vol per unit doubles, the vol scalar halves, and your next position is automatically half the size. If vol collapses, positions grow. Your dollar risk stays roughly constant across regimes without you making a single judgment call. Compare that to the trader who always buys "about $50k of BTC": that trader is running triple the risk in a wild market that they run in a quiet one, and usually without noticing until the wild market teaches them.
The chain also changes cross-instrument comparisons. 0.703 BTC and 160 shares of an index ETF sound incomparable, but if both came out of the same chain at forecast +10, they carry identical daily dollar risk. Equal-risk sizing, which the vol-units lesson argued for in principle, falls out of this machinery automatically.
10.9.4 The trailing stop: your only exit
Sizing was the first half of the deal. The second half is the exit. There is exactly one exit mechanism: a volatility-scaled trailing stop on daily closes. No profit targets, no discretionary exits, no exceptions.
The rule for longs: exit if the daily close falls more than X times daily volatility (in price points) below the highest close since entry. For shorts, mirror it: exit if the daily close rises more than X times daily volatility above the lowest close since entry.
In formulas:
Stop distance = X x Daily vol $ Initial stop (long) = Entry price - Stop distance Trailing stop (long) = Highest close since entry - Stop distance
The stop only ratchets in your favor. Every new high close drags it up; nothing ever moves it down.
X sets the character of your trading. Small X means tight stops, short holds, many trades, and frequent shakeouts on noise. Large X means loose stops, long holds, few trades, and more open profit surrendered before the exit finally triggers.
| X | Approx. holding period | Approx. trades/year | Character |
|---|---|---|---|
| 2 | ~9 days | ~29 | Tight, many false exits |
| 3 | ~4 weeks | ~12 | Moderate |
| 4 | ~6.5 weeks | ~8 | The default, best balance |
| 6 | ~13 weeks | ~4 | Loose, more giveback |
| 8 | ~26 weeks | ~2 | Very loose, slow instruments only |
I would start at X = 4 and stay there unless you have a specific, articulated reason to move. At X = 4 you hold winners for weeks, which matches the swing-to-position horizon this whole course is built around, and the stop sits far enough out that ordinary daily noise almost never touches it.
Concrete example. You go long BTC at $65,000 with daily vol at $2,223:
Stop distance: 4 x $2,223 = $8,892 Initial stop: $65,000 - $8,892 = $56,108
BTC rallies and closes at $80,000:
Trailing stop: $80,000 - $8,892 = $71,108
You've now locked in $6,108 of profit per BTC ($71,108 minus your $65,000 entry) while still giving the trade room to keep running. The position can't round-trip back to a loss.
The Trailing Stop Ratchets
The giveback is the part that bothers people first. By construction, you'll always surrender X times daily vol from the peak before you exit. Suppose the rally runs to $90,000 and then rolls over. Your stop sits at $90,000 - $8,892 = $81,108. You exit there, capturing $16,108 of the $25,000 move, about 64 percent, and handing back $8,892 from the top.
That giveback is a cost of the method, not a flaw in it. The only way to avoid it is to sell at the top, and selling at the top requires predicting the top, which you can't do reliably. What the trailing stop guarantees instead is that you capture the majority of any large move without ever needing to call a turn. Across many trades, catching 60-plus percent of every big trend beats occasionally nailing an exit and routinely cutting winners at 20 percent of the move because the P&L felt too good to risk.
Why a volatility-scaled stop rather than a fixed price stop? A fixed stop is really a random-width stop. A $5,000 stop on BTC is 1.7 times daily vol when vol is $3,000 (you'll be stopped out by ordinary noise) and 10 times daily vol when vol is $500 (absurdly wide, huge loss if hit). The vol-scaled stop is the same statistical distance from price in every regime: tight in calm markets, wide in wild ones. That adaptation is what a dollar-based stop can't give you.
Two operational details. Stops are evaluated on daily closes only. Intraday spikes through your level don't count; the market gets until the close to make up its mind, which filters out a surprising number of stop-hunting wicks (recall the liquidity-run mechanics from the technical analysis lessons). And when a close does breach the stop, you exit the next trading day, at the open or with a limit order. No "one more day to see if it recovers."
What about gaps? Sometimes price gaps straight through your stop, especially in crypto over weekends or in futures across session breaks. Your stop was $71,000, price closed at $72,000, and the next session opens at $68,000. You exit at $68,000 and eat the slippage. Three layers of the framework keep this survivable: the vol target caps your total exposure, X = 4 puts the stop far enough away that gapping through it is rare, and the hard ceiling on vol targets (coming below) exists precisely because gap risk can't be engineered away.
10.9.5 Managing the trade day by day
The daily routine is simple, and that is deliberate.
Before entry: get current RV (short and long window), blend it, compute daily vol per unit, compute your vol scalar, assign the forecast, compute the position, compute the stop distance, enter, and write down the entry price, forecast, and stop level.
Every day the trade is open: check the daily close. Update the highest close since entry (lowest for shorts). Compute the stop as that extreme minus the stop distance. If the close is above the stop, do nothing. If the close is below it, exit the next day. That's the whole job, and it takes about two minutes per position.
Just as important is the list of things you don't do while the trade is open. You don't check the P&L and debate taking profits. You don't exit because a momentum indicator rolled over. You don't trim because you're nervous, and you don't add because it's working (there's a correct way to add, next section). You don't close on a news headline, and you don't close because you want the money for something else. Every one of those is a discretionary exit, and discretionary exits are the specific failure mode you signed away when you took the deal at the top of this lesson.
If this feels too passive, it frees you up. All the attention you used to burn managing open positions goes back into the thing you're actually good at: finding the next trade.
10.9.6 Pyramiding: adding to winners without breaking the rules
Suppose you entered at +10, the trade is working, and now the market hands you a fresh setup in the same direction: a clean breakout, a new positioning extreme, whatever your process recognizes. Your conviction has genuinely increased. The forecast rule says you can't touch the original trade. So what do you do?
You open a new, separate bet on the same instrument. It gets its own entry price, its own forecast, its own position size run through the chain at current vol, and its own trailing stop anchored to its own entry. The original bet is untouched.
Three rules govern this. Each bet is fully independent: one hitting its stop doesn't close the others. The total absolute forecast across all open bets on one instrument must stay at or below 40, which caps your maximum exposure at four average-conviction units of size no matter how euphoric the trend gets. And every added bet must be a real setup that would justify a trade on its own, not "it went up so I bought more."
Example. You are long 0.70 BTC from $65,000 at forecast +10. BTC breaks $75,000 convincingly and consolidates above it, a setup you would take cold. You open bet two: 0.70 BTC at $75,000, forecast +10, its own stop trailing from its own highest close. Total forecast is 20 of the 40 allowed; total position 1.40 BTC.
Now BTC pulls back after topping at $78,000. Both stops sit at $78,000 - $8,892 = $69,108. A close at $70,000 leaves both alive. A close at $68,500 kills both. In a different path, where bet one had built a big cushion before bet two entered, the pullback might stop out only the newer bet while the older one rides on. Each bet lives and dies on its own terms, which is what makes pyramiding safe: you're never averaging into one giant undifferentiated position.
This structure forbids adding to losers. There's no version of "my forecast was +10 and it dropped, so it's a better price now, add more." A new bet needs a new setup, and a position moving against you isn't one. Averaging down is the most reliable account destroyer in discretionary trading, and the framework has no rule that allows it.
10.9.7 Choosing your vol target
The chain treats the vol target as a given, but choosing it is the most consequential decision in the whole framework. It scales everything: position sizes, stop distances in dollar terms, drawdowns, and how bad the worst week of your year feels.
The vol targeting lesson covered the general logic; this is the specific calibration for a semi-automatic trader. The right vol target depends on the Sharpe ratio you can realistically achieve, and the mapping comes straight from the fractional Kelly logic of the last lesson. Kelly says optimal risk scales with edge; running at half Kelly, the practical shortcut is:
So an expected Sharpe of 0.5 supports a 25 percent vol target, 0.4 supports 20 percent, 0.3 supports 15 percent. The more edge per unit of risk you actually have, the more risk you can afford to run, and half Kelly keeps you far enough from the cliff edge that estimation error doesn't push you over.
The question is what Sharpe to assume. Without a backtest, a discretionary trader has no measured number, so use the ceiling that experience with this style of trading supports: about 0.5 at best for a skilled semi-automatic trader over the long run. That caps the vol target at 25 percent. If you're newer or unsure, start at 15 to 20 percent and earn your way up with live results. Nothing is lost by starting small except a little upside; everything is lost by starting too big.
One adjustment matters enough to be a rule. Trade type changes the math because skew changes the math. Trend-following trades (buying breakouts, riding momentum) have positive skew: many small losses, occasional huge wins, and the trailing stop caps the downside naturally. Half Kelly is appropriate. Counter-trend trades (fading extremes, buying dips against the trend) have negative skew: many small wins and the occasional trade that keeps going against you. For those, cut to quarter Kelly, meaning halve your usual vol target. If your standard target is 25 percent, a counter-trend trade runs at 12 to 15. This connects directly to the skew discussion in the performance measurement lesson: negative-skew strategies look smooth right up until they don't, and the correct response is to pre-commit to smaller size, not to hope you'll be quick enough on the exit.
What do different targets feel like in practice? With $100,000, rough expectations look like this:
| Vol target | Daily cash vol | Rough worst day in a month | Rough worst week in a year |
|---|---|---|---|
| 10% | $625 | ~$1,000 | ~$2,800 |
| 15% | $938 | ~$1,500 | ~$4,200 |
| 20% | $1,250 | ~$2,000 | ~$5,600 |
| 25% | $1,563 | ~$2,500 | ~$7,000 |
| 40% | $2,500 | ~$4,000 | ~$11,200 |
| 50% | $3,125 | ~$5,000 | ~$14,000 |
These are ordinary-conditions figures. In a genuine crisis, the kind of week where correlations converge and vol explodes (2008, March 2020, the May 2021 crypto crash), losses can run three to five times the worst-week column. Apply that multiplier to the 25 percent row: a $20,000 to $35,000 week on a $100,000 account. If that number would end you, financially or psychologically, your vol target is too high, whatever your Sharpe estimate says.
There is a hard ceiling: never exceed a 50 percent vol target, under any circumstances, with any track record. Even a Sharpe of 2.0, which you don't have and won't have, doesn't justify going past it, because fat tails, estimation error, and gap risk all live beyond the reach of the daily-close machinery. The drawdown math coming two lessons from now will make the case in full; for now, take the ceiling as law.
10.9.8 The chain end to end: worked examples
Three complete runs, from setup to stop, so the whole procedure is concrete.
Start with a counter-trend BTC long. BTC has fallen from $108,000 to $65,000. The platform's momentum indicator is deeply negative and has just crossed above its signal line, and options skew has stretched to an extreme. You expect a relief rally. Capital $100,000. This is a counter-trend, negative-skew trade, so the disciplined vol target is 12 to 15 percent, half your standard 25. The numbers below run at the full 25 percent so you can compare them directly against the trend example that follows; treat that as the undisciplined version and halve everything for the target you'd actually use. Forecast +10: against the primary trend, moderate conviction only, no matter how oversold it looks.
Blended vol: 0.7 x 55% + 0.3 x 45% = 52%
Daily vol %: 52% / 16 = 3.25%
Daily vol $: 3.25% x $65,000 = $2,113
Daily cash vol target: ($100,000 x 0.25) / 16 = $1,563
Vol scalar: $1,563 / $2,113 = 0.740 BTC
Position at +10: 0.740 BTC (notional $48,100)
Stop distance: 4 x $2,113 = $8,450
Initial stop: $65,000 - $8,450 = $56,550
Max initial loss: 0.740 x $8,450 = $6,253, about 6.3% of the account
At the disciplined 12.5 percent counter-trend target, everything halves: 0.37 BTC, max initial loss around 3.1 percent. That's what quarter Kelly buys you: a negative-skew trade that can't hurt you much when the trend reasserts itself, which it often will.
Next, a trend-following ETH long. ETH breaks above its 50-day high, fast trend measures agree, volume confirms. It's with the trend, so the full 25 percent target applies, and the setup is strong: forecast +15. ETH at $3,400, blended RV 70 percent.
Daily vol %: 70% / 16 = 4.375%
Daily vol $: 4.375% x $3,400 = $148.75
Vol scalar: $1,563 / $148.75 = 10.5 ETH
Position at +15: (10.5 x 15) / 10 = 15.75 ETH (notional $53,550)
Stop distance: 4 x $148.75 = $595
Initial stop: $3,400 - $595 = $2,805
Max initial loss: 15.75 x $595 = $9,371, about 9.4% of the account
The initial risk is larger than the BTC trade. That's the forecast doing its job: a +15 trade carries 1.5 times the risk of a +10 trade, by design, because you judged the setup 1.5 units of conviction better. The system didn't get excited. You expressed measured conviction once, at entry, and the math translated it.
And last, a sedate one for contrast: an index ETF position trade. Price $520, blended RV 18 percent, conservative 15 percent vol target, forecast +10.
Daily vol %: 18% / 16 = 1.125%
Daily vol $: 1.125% x $520 = $5.85
Daily cash vol target: ($100,000 x 0.15) / 16 = $937.50
Vol scalar: $937.50 / $5.85 = 160 shares
Position at +10: 160 shares (notional $83,200)
Stop distance: 4 x $5.85 = $23.40
Initial stop: $520 - $23.40 = $496.60
Max initial loss: 160 x $23.40 = $3,744, about 3.7% of the account
Three instruments with wildly different prices and volatilities, one procedure, and every position calibrated to a known slice of your risk budget. The notional values look arbitrary ($48k of BTC, $53k of ETH, $83k of an ETF) but the risk isn't. Notional is the wrong measure, as the vol-units lesson argued, and the chain is what makes the right one work in practice.
10.9.9 The rules, stated once
Everything above compresses into a short list, and the list is the system. Breaking any single rule quietly disables the whole framework, because each rule exists to block a specific, well-documented way traders sabotage themselves.
At entry: assign a forecast between -20 and +20 before entering, size with the chain without rounding up, and record entry, forecast, and stop.
During the trade: never change the forecast, never override the trailing stop (not for news, not for indicator signals, not for feelings), never add to a loser, check stops at the daily close only, and execute triggered exits the next day.
At exit: the trailing stop is the only exit. No profit targets, no "this is enough," no extra day of grace.
At the risk level: total absolute forecast per instrument stays at or below 40, the vol target never exceeds 50 percent, vol estimates update only on a 25 percent change, and counter-trend trades run at half your normal vol target.
Every rule is a pre-commitment made while you're calm, designed to bind you at a future moment when you won't be. That's not specific to this framework. It's the whole design philosophy, and the next lesson takes it head on: why the trader who systematizes everything after the entry, even while staying fully discretionary about the entry itself, survives the moments that end everyone else.
10.10 The semi-systematic mindset
The previous lesson handed you a complete machine: a forecast scale, a sizing chain, a volatility-based trailing stop, a pyramiding rule, a vol target. This lesson is about why you'll be tempted to break that machine, why the temptation can't be reasoned with, and how to build your trading so the machine wins anyway.
Most trading education treats psychology as its own subject, with its own shelf of advice about mindfulness, emotional control, and knowing yourself. This course doesn't, on purpose. The entire behavioral content of this course reduces to one claim and one consequence. The claim: overbetting kills more traders than bad analysis does, and no amount of self-awareness prevents it at the moment it happens. The consequence: since the problem can't be solved inside your head, it has to be solved outside it, with rules that remove the decision entirely. A trader who systematizes sizing and exits while keeping entries discretionary has fixed the thing that actually kills accounts. A trader who masters their emotions but keeps negotiating position size in the moment has fixed nothing.
That split, discretion in trade selection and rails everywhere after entry, is what this course means by semi-systematic. The previous lesson gave you the rails. This one is about respecting them.
10.10.1 Why self-awareness fails
The standard advice fails, because if self-knowledge worked, none of this machinery would be necessary.
The standard advice says: learn your biases, journal your emotions, recognize when you're tilting, and correct in real time. The problem is a gap that shows up wherever humans make decisions under arousal. The state you make plans in isn't the state you execute them in. When you calmly decide, on a Sunday, that you'll never risk more than 1 percent per trade, you're a different decision-maker than the person watching a position rip on Wednesday afternoon with the feeling that this is the trade of the quarter. Cold-state you sets policy. Hot-state you holds the mouse. And hot-state you doesn't experience the moment as temptation; it experiences it as insight. The override never announces itself as a mistake. It arrives as information: "This setup is different." "The stop is obviously too tight here." "I've never been more sure." Every trader who blew up a good strategy by oversizing one trade heard some version of those sentences from the inside, and from the inside they sounded like analysis.
Knowing about this gap doesn't close it. People taught about a bias remain roughly as subject to it as before; they mostly get better at spotting it in others. You can recite the disposition effect from memory and still feel, in the moment, that this particular loser deserves more room. The knowledge and the impulse live in different systems, and under P&L stress the impulse system has the faster connection to your hands.
So willpower isn't a component you can build on. It fails exactly when the load peaks, which is the worst failure profile a component can have. Every other field that faces this problem has reached the same answer. Aviation and surgery don't handle high-stakes moments by asking professionals to be more self-aware; they hand them checklists and procedures precisely because experienced, intelligent people skip steps under pressure. Casinos are built top to bottom around hot-state decision-making, and the house takes the winning side of that trade. The fix for a decision that fails under pressure is to stop making it under pressure. Make it once, in the cold state, write it down, and then arrange your trading so the hot state has nothing left to decide.
10.10.2 Overbetting is the one that kills you
Of all the errors available to a trader, why single out size?
Because the penalty structure is different in kind, not just in degree. Bad analysis, a genuinely worthless signal, costs you your edge; sized sanely, you bleed slowly toward the market's average minus costs, and you have months or years to notice and stop. Overtrading costs you fees and slippage, a steady leak. Bad entries cost you a few tenths of expectancy per trade. All of these hurt, and all of them are survivable long enough to be diagnosed. Overbetting is the only common error that converts a winning strategy into a losing one while every individual decision still looks defensible.
You've already seen the mechanism twice in this part, so this assembles it rather than introducing it. The Kelly lesson showed that compound growth equals average return minus half the variance, g = mu - sigma^2 / 2, and that scaling a bet up scales the edge linearly but the variance quadratically. Push size past the growth-optimal point and growth falls; push it to roughly twice that point and growth hits zero; push past that and a strategy with genuinely positive expected value compounds your account toward nothing. There's a size at which a real, honest, positive edge becomes a wealth destroyer, and nothing about the trades themselves changes. Same signal, same win rate, same average return per trade. The only variable is how much you bet, and it alone decides whether you compound up or down.
The drawdown lesson coming next does this properly, but one piece belongs here: losses and recoveries aren't symmetric. A 20 percent drawdown needs 25 percent to recover. A 50 percent drawdown needs 100 percent. An 80 percent drawdown needs 400 percent. Overbetting is what turns the routine losing streak that every strategy produces (and the sample-size lesson showed you how long those streaks run even with a real edge) into a hole with that geometry. The trader who risks 1 percent per trade and hits eight losers in a row is down about 8 percent and mildly annoyed. The trader risking 10 percent on the same eight trades is down roughly 57 percent and now needs to more than double the account just to reclaim the high-water mark, with wounded confidence and, usually, a fresh urge to size up and get it back, which is how a drawdown becomes a spiral.
Overbetting is a behavioral problem, not just a math one, because the urge to overbet arrives precisely at the moments when your judgment about size is least reliable. After a winning streak, when the profits feel like the house's money and your self-assessed skill has inflated past anything the sample supports. After a losing streak, when getting back to even starts to feel like a goal that justifies extra risk. And on the trades that feel most certain. That last one is the most expensive feeling in trading, and it gets its own paragraph.
Conviction feels like information about the trade. Mostly it's information about you: how recently similar setups paid, how clean the chart looks, how many confirming opinions you consumed this week. Calibration is a skill, and untrained calibration is poor; the overwhelming majority of people rate themselves above average on skills they care about, and traders rate their sure things far surer than the outcomes justify. Worse, in markets there's a structural reason the strongest-feeling trades underdeliver: a setup that looks obviously compelling to you looks compelling to everyone, which means it's crowded, which means the entry is worse and the exit is a stampede. The trades that feel like +20s on the forecast scale are exactly the ones where the scale exists to stop you from betting like it's a +40. This is why the previous lesson's framework caps the forecast at 20 and tells you most trades should be a 10. The cap is a hard limit on the damage your best feeling can do, not modesty theater.
So the ranking holds. Bad analysis is a slow leak you can find and patch. Overbetting is a structural failure that takes the account down in one storm, it's triggered by internal states rather than market conditions, and it recruits your own conviction as its advocate. That's why it gets the course's entire psychology budget.
10.10.3 The division of labor
If the in-the-moment self can't be trusted with size and exits, the design question becomes: which decisions should stay human at all?
The answer is comparative advantage, assessed honestly in both directions. Humans are genuinely good at things that are hard to write down. They synthesize heterogeneous context: this COT extreme matters more than usual because the seasonal window agrees and the macro calendar is clear, that funding spike is less meaningful because a single venue is distorting it. They recognize when data is broken or when the regime has shifted in a way no lookback window has caught yet. They filter a screener's twenty candidates down to the three where the story, the level, and the flow line up. This is real edge, it's why you're trading discretionarily instead of buying an index fund, and the earlier parts of this course spent dozens of lessons feeding it: positioning, vol surfaces, regime, technicals as the execution layer.
Humans are terrible at the other half. They struggle with consistency: applying the same rule the same way on trade 4 and trade 400, in a drawdown and out of one. They struggle with arithmetic under stress and with indifference to sunk P&L. They struggle to sell a loser without flinching and to let a winner run without grabbing the profit early. The disposition effect, the tendency to cut winners quickly while giving losers room, is one of the best-documented behaviors in every population of traders ever studied, and it's precisely backwards, a machine for clipping your right tail and feeding your left one.
Rules have the mirror-image profile. They can't read context; a formula doesn't know that this breakout has a catalyst and that one is a Friday afternoon head fake. But they're perfectly consistent, they compute exact sizes without rounding up for excitement, and they feel nothing at all about being down 30 percent on a position.
The semi-systematic split follows directly. You keep what you're good at: what to trade, which direction, when to enter, and how strong the setup is on a fixed, bounded scale. The system keeps what you're bad at: how big, where the stop lives, when the trade dies. That was the deal stated in the previous lesson; what this lesson adds is that the deal only works if it's absolute. A rule you override on special occasions isn't a rule with exceptions. It's a suggestion, because the hot state will find that every occasion that matters is special. The value of the machine isn't that it makes better decisions than your best self. Your best self might genuinely beat it. The value is that it makes the same decision every time, including the times when the self showing up to trade is nowhere near your best.
One clarification before the inventory, because the framework contains an apparent contradiction. This lesson says fixed sizing formulas instead of conviction sizing, yet the previous lesson's framework sizes positions by your conviction through the forecast. The difference is what makes the framework safe. The forecast is bounded (nothing past 20, so your maximum-conviction trade is exactly twice your average one, not ten times), it's set once before entry, and it's frozen the moment you're in the trade. Conviction sizing in the dangerous sense is the opposite on all three counts: unbounded, decided in the heat of the moment, and renegotiated continuously while the position is open. The forecast scale is cold-state conviction, quantized and capped. What it forbids is the hot-state version: bumping a +10 to a +15 because the trade started working, which is your excitement voting, not your analysis.
10.10.4 The decision inventory
Take a trade's life from idea to flat and sort every decision into one of two bins: discretionary or on rails. Doing this explicitly, once, for your own trading, is worth more than any amount of reading about discipline, so treat this section as a template.
Before entry, discretion rules. Which markets to scan, which signals to weight, whether today's setup clears your bar, long or short, enter now or wait for the level: all yours. This is where the platform's dashboards and screeners live in your process, and where everything from the positioning, options, and technicals parts of this course gets applied. The forecast itself, the number from 5 to 20 that says how good this setup is, is the last discretionary act of the trade, and it doubles as the handoff: the moment you write it down, the machine takes over.
From that handoff on, every decision has a known, documented human failure mode sitting on it.
Position size. The failure mode is everything in the overbetting section: house money after wins, revenge after losses, oversizing sure things. The rail is the sizing chain, computed, not estimated, with the output taken as-is. Not rounded up because it feels small. The previous lesson made this point and it bears repeating because rounding up is how the system dies by a thousand cuts: a formula whose output you adjust by feel is feel-based sizing with extra steps.
The stop. The failure mode is placing stops by pain tolerance ("I don't want to lose more than $500") or by hope ("below that support it's obviously wrong"), neither of which has anything to do with how much the instrument actually moves. The rail is the volatility-scaled stop: a multiple of daily vol, or equivalently a multiple of ATR, below the highest close since entry. It adapts to the market instead of to your feelings, sits wide enough that noise doesn't shake you out, and trails so that the exit question never reopens.
The exit. This is the decision where discretion does the most damage, because the disposition effect means feel-based exits aren't randomly wrong but systematically wrong in the worst direction. Every "I'll just take this profit before it disappears" clips a winner; every "it will come back" extends a loser. The rail is simple: the trailing stop is the only exit. No profit targets, no exiting on a bad feeling, no closing because the news turned scary. If the news is genuinely bad, price will take out the stop and the system will exit you. If price shrugs the news off, the position deserved to live, and your fear was noise.
Adding to the position. The failure mode is averaging down, which is the overbetting engine in disguise: every add to a loser raises size exactly as the trade proves itself wrong, financed by the feeling that being early isn't the same as being mistaken. The rail from the previous lesson: never add to a loser, ever, and add to winners only as a new, independent bet with its own forecast, its own stop, and a hard cap on total exposure per instrument.
Mid-trade meddling. The failure mode is everything else: trimming because you're nervous, doubling the forecast because it's working, moving the stop down "just this once," closing early because you want the win on the books before the weekend. The rail is that the daily routine contains exactly one question: did today's close breach the stop? If no, you do nothing. Not "you may do nothing." Nothing is the assignment.
The inventory has a clear shape. Discretion is front-loaded into the one phase where human judgment has an edge and emotional stakes are lowest, because you have no position yet and can walk away from any setup for free. The rails cover the entire phase where money is on the line and your state is compromised by definition. Discretion belongs in trade selection. Everything after entry runs on rails.
10.10.5 Building rules that survive contact with temptation
Not every written rule holds up. Most traders have rules; most traders break them. The difference between rules that hold and rules that fold comes down to a few design properties, and they're worth engineering deliberately.
Binary compliance. A rule must be checkable with yes or no, by a stranger reading your journal. "Don't risk too much on one trade" isn't a rule, it's a mood; the hot state will happily agree that this trade isn't too much. "Risk per trade never exceeds 4 times daily vol times position size, computed before entry" is a rule. If following it requires judgment, the judgment call becomes the leak, because the hot state controls the judgment. Everything ambiguous in your rulebook will eventually be interpreted in favor of the trade you currently want to make.
No override clause. The moment a rule contains "except when," the exception swallows it. Stress is precise: it will locate the loophole you left and route every violation through it. If a rule genuinely needs an exception, that means the rule is wrong, and the fix is to rewrite the rule through the change process below, not to override it live. Between rewrites, the rule as written is the rule.
Friction on the violation path. Make breaking a rule mechanically harder than following it. Put the actual stop order in the market rather than keeping a mental stop, so that doing nothing executes the plan and intervening requires action. Pre-compute the position size before the entry trigger fires, so the number exists before the excitement does. Check stops on daily closes only, per the previous lesson, so intraday noise never gets a vote. Some traders add process friction on top: a standing arrangement that any deviation must be written down and dated before the order goes in. It's remarkable how many violations don't survive the act of writing "I am overriding my stop because I am sure" in a journal you'll reread.
Written before, not during. Every rule gets authored in the cold state, away from open positions, and the complete set fits on one page you can see while trading. A rulebook you have to remember is a rulebook the hot state gets to misremember.
A checklist at the gate. The entry checklist is where the rules become a physical act. Before any order: setup identified and named; direction and forecast written down; vol estimate current (within the update rule from the sizing lesson); position size computed from the chain; stop distance and initial stop level computed; max loss at the stop calculated in dollars and as a percent of the account; total forecast on this instrument within the cap. Seven lines, two minutes, and it converts "I checked everything" from a feeling into a fact. Checklists look insulting to skilled people, which is exactly why fields where skipped steps kill people force them on their most skilled practitioners. The checklist isn't there for the trades where you're careful. It's there for the trade you rush into at 3:47pm because it's running without you, which is precisely the trade most likely to be oversized, and the two minutes it costs are the point: a setup that can't survive a two-minute delay was a chase, not a setup.
The maximum that is never negotiated. Above all the per-trade machinery sits one number: the most you can lose on a single trade if the stop is hit, as a fraction of the account. The sizing chain usually keeps you well inside it, but the hard cap exists for the day you're tempted to run the chain with a thumb on the scale, a fudged vol estimate, an "adjusted" forecast. Whatever the cap is (the vol-target arithmetic from the previous lesson implies single digits of percent at the stop, and lower is fine), it has the property that no setup, no conviction level, no drawdown, and no opportunity changes it. Not negotiated means not negotiated with yourself, which is the only counterparty who was ever going to ask.
10.10.6 Changing the rules without cheating
A rulebook frozen forever would be its own mistake. Your vol target should evolve with evidence about your edge, a stop multiple might genuinely be wrong for the instruments you trade, and a checklist item might prove useless. The danger isn't change; it's when and why the change happens. A reliable tell separates learning from cheating: honest rule changes happen away from open positions and away from recent pain, and dishonest ones happen mid-trade or mid-drawdown, always in the direction that permits what you currently want to do.
So put the change process itself on rails. Rule changes happen on a fixed review cadence, monthly or quarterly, never intraday and never with the affected position open. Every change is written: the old rule, the new rule, and the evidence. And the evidence must be a sample, not a story. "The stop cost me money on Tuesday" isn't evidence; any rule will lose on individual trades, and a stop that never costs you a winner is a stop too wide to protect you. "Over the last 40 trades, exits at this multiple gave back an average of X while a wider multiple would have captured Y, here's the tally" is evidence. The sample-size material from earlier in this part applies to your rules exactly as it applies to your strategies: single outcomes are noise, and a rule change justified by one painful trade is curve-fitting your rulebook to your last regret.
A useful backstop is a waiting period: a proposed change sits written and unimplemented until the next review before it takes effect. If it still looks right two weeks later, with no position on and no fresh wound, it probably is. Most proposals don't survive the wait, which tells you what they were.
10.10.7 The journal as a compliance audit
The last piece of structure is measurement, because a system nobody audits degrades quietly.
Most trading journals track P&L, and P&L is the least informative thing a semi-systematic trader can track over a month of trades; the distributions lesson showed how little a small sample of outcomes says about edge. The higher-signal number is compliance: for every trade, did the entry pass the checklist, was the size the chain's output to the decimal, was the forecast left untouched, did the exit come from the stop and nothing else. Log it per trade as a simple yes or no with a note on any deviation. Over a quarter you get the statistic that actually predicts your survival: your violation rate, and its trend.
Then split your results into rule-following trades and violations, and compare. Two findings are common. Violations underperform, which is useful and motivating and roughly what you expected. The more dangerous finding: sometimes a violation makes money. You override the stop, the trade comes back, you bank a profit that the rules would have denied you. That trade is the most expensive winner you'll ever have, because what it paid you in dollars it charged you in discipline. It taught your hot state that overrides work, and the hot state generalizes from a sample of one. The tenth override, sized bigger because the first nine built confidence, is the one that ends the account. Score outcomes and compliance separately, and treat a profitable violation as a violation, full stop. The market occasionally pays people for mistakes; that's variance, not vindication.
10.10.8 Taking the objections seriously
Three pushbacks come up every time this framework meets an experienced discretionary trader, and they deserve straight answers.
"My post-entry management is part of my edge." Maybe. It's a testable claim, so test it instead of asserting it: for your next 30 trades, log the exit your feel produced and, in parallel, the exit the trailing stop would have produced, and compare the totals. Most traders who run this audit find their interventions cost money net, mostly by clipping the handful of big winners that were carrying the whole distribution; the trend-following math is unforgiving about that, because when returns are skewed, the right tail pays for everything and feel-based exits amputate the right tail. If your audit genuinely shows your overrides add value across a real sample, you've found a rare skill, and you can promote it into a written rule with defined conditions, which is the difference between an edge and a mood.
"The rules would have kept me out of my best trade ever." Probably true, and it's an argument for the rules, not against them. Memory curates. You remember the oversized trade that paid; the counterfactual archive of oversized trades that would have ended you doesn't send highlights. Any risk framework, run long enough, will cost you some individual spectacular outcome, the same way refusing to bet your house on one hand costs you the memory of the time it would have doubled. The framework isn't optimizing your best trade. It's optimizing the compound growth of the whole sequence, and the Kelly lesson already showed those two goals point in different directions.
"This much structure will strangle the intuition that makes me good." The inventory answers this one. Every input your intuition actually uses, reading the tape, weighing positioning against regime, sensing when a level will hold, lives in the discretionary zone, untouched. What the rails remove isn't intuition but arithmetic and impulse: the sizing math you were doing worse than a formula anyway, and the exit twitches that were never intuition to begin with, just fear and greed mistaken for it. Traders who adopt the split tend to report the opposite of strangulation: with size and exits off their desk, the attention that used to burn on watching open P&L goes back into finding the next trade, which is the only place it earns anything.
There's one more objection nobody says out loud: the rules make trading feel less like trading. Less action, fewer decisions, long stretches where the daily routine is checking a close against a stop level and doing nothing. That feeling is accurate, and it's the fee. The excitement you're giving up was never free; it was being financed, the whole time, by your expectancy.
Structure is what lets you take risk and stay in the game long enough for edge to show up in the results, but it doesn't make the losing stretches disappear; it only makes them survivable, and you should know in advance exactly what surviving them looks like. The next lesson does that arithmetic: how drawdowns compound, why recovering is harder than losing, and how to tell the drawdown you sit through from the one that's telling you something.
10.11 Drawdowns and ruin
A drawdown is the distance between your account and its own best day. Formally: take the running maximum of your equity curve (the high-water mark), and the drawdown at any moment is the percentage drop from that peak to where you are now. Maximum drawdown is the worst such drop over whatever window you're looking at. Every trader tracks it, most traders underestimate it, and the ones who stop trading usually stop because of it, not because their average trade was bad.
The previous lesson said that rules don't make losing stretches disappear, they make them survivable, and that you should know in advance what surviving looks like. This lesson is that arithmetic: the brutal asymmetry between losing money and making it back, why the order your returns arrive in can matter as much as the returns themselves, the actual probability math of blowing up, why the diversification protecting you in normal markets evaporates exactly when you need it, and the hardest judgment call in trading, telling the drawdown you sit through from the one that's telling you the edge is gone.
This isn't pessimism. It's the operating manual for the part of trading where you're losing, which is most of the time.
10.11.1 The arithmetic of getting it back
Losses and gains aren't symmetric, because they compound off different bases. Lose 10 percent and you need 11.1 percent to get back to even, not 10. Lose 50 percent and you need 100 percent. The general formula:
In plain terms, the money you lost is a fixed number of dollars, but you have to earn it back with a smaller account, so every percentage point of recovery is worth fewer dollars than the percentage points you lost. The deeper the hole, the worse the exchange rate.
The full table:
| Drawdown | Gain needed to recover |
|---|---|
| 5% | 5.3% |
| 10% | 11.1% |
| 20% | 25% |
| 30% | 42.9% |
| 40% | 66.7% |
| 50% | 100% |
| 60% | 150% |
| 70% | 233% |
| 80% | 400% |
| 90% | 900% |
The pattern is what matters. Down to about 20 percent, recovery costs roughly what you lost plus a small tax. Past 30 percent the curve bends, and past 50 percent it goes vertical. The function is convex in the worst possible direction: each additional point of drawdown costs more recovery than the last one did.
The Cost of Getting It Back
Now add time. Suppose you're genuinely good and compound at 15 percent a year, which over real costs and real markets would put you in rare company. Recovering a 50 percent drawdown means doubling the account, and at 15 percent a year a double takes about five years (1.15^5 is just over 2). A 90 percent drawdown needs a 10x, which at the same rate takes about sixteen and a half years. That is sixteen years of excellent trading to undo one catastrophic stretch, and it assumes your edge, your nerve, and your capital source all survive intact for the duration.
They usually don't, which is the second layer of the asymmetry. The arithmetic above assumes you keep trading at full effectiveness through the recovery. In practice deep drawdowns degrade the trader along with the account. You size down out of fear (rational or not), you second-guess signals you'd have taken cleanly at high water, and if you manage outside money, redemptions shrink the base further right when you need it. The table shows the best case. The realized recovery is almost always slower.
This is why the entire sizing apparatus of the previous five lessons exists. Vol targeting, fractional Kelly, fixed per-trade risk: every one of those tools is, at bottom, a machine for keeping you in the flat part of that table. The difference between a trader who caps drawdowns near 20 percent and one who lets them reach 50 isn't that one loses less. It's that one of them faces a 25 percent recovery problem and the other faces a 100 percent recovery problem, roughly a five-year difference in lost time at good rates of return.
10.11.2 Drawdowns are the normal state of trading
A profitable strategy spends most of its life in drawdown, a fact nobody internalizes until they live it. New equity highs are single points; everything between two highs is underwater by definition. Run the simulation yourself with any positive-edge return stream and count the days at a new high versus the days below one. Even for a strategy with a Sharpe ratio around 1, the kind of performance most traders never sustain, the account sits below its last peak far more often than it sits at one. The high-water mark is where you visit. Drawdown is where you live.
Two properties of maximum drawdown make it nastier than the metrics you met in the performance lesson.
It only ratchets up. Max drawdown is a running worst case, so a longer track record can never show a smaller one. Trade for two years and your worst drawdown might be 12 percent. Trade the identical strategy for twenty years and your worst drawdown will be deeper, not because anything changed but because you gave the bad tail more chances to show up. When someone quotes a max drawdown, the number is meaningless without the length of the window it came from.
It's also a single draw from a wide distribution. Your backtest's max drawdown is one sample of one path. Rerun history with slightly different luck and the same strategy prints a different worst stretch, sometimes much worse. Simulate a strategy many times with the same statistical properties and the spread of maximum drawdowns across runs is enormous. The practical rule that falls out of this: plan for a live drawdown roughly twice as deep as the worst one in your backtest. Not because backtests are dishonest (though the backtesting lesson covered how they can be), but because the historical path is one draw and the future is another, and you sized the backtest after seeing its luck.
That planning number should drive real decisions. If your backtest shows a 15 percent max drawdown, ask yourself whether you can financially and psychologically hold through 30. If the honest answer is no, the strategy is oversized for you regardless of what the Kelly math or the vol target says. The binding constraint on size is rarely the optimizer. It is the deepest drawdown you can pass through without breaking something: the account, the mortgage payment, or the discipline to keep taking signals.
There is a rough relationship between volatility and drawdown depth. A strategy's routine drawdowns run around one to two times its annualized volatility, and its worst drawdowns over long horizons run deeper than that, with lower-Sharpe strategies drifting toward the bad end of the range. A 20 percent vol book should expect 20 to 40 percent drawdowns as a cost of doing business, not as a crisis. If that number is unacceptable, the fix is the vol target, set before the drawdown arrives, not a hasty intervention in the middle of one. The negative skew warning from the performance lesson applies double here: strategies that sell insurance, short vol, carry, premium harvesting, understate their eventual drawdown in any sample that hasn't yet contained their bad event. Their smooth years aren't evidence of shallow drawdowns to come. They're the premium collected while the drawdown waits.
10.11.3 Path dependency
Take two years of monthly returns, shuffle them into a different order, and compound them. The ending wealth is identical. Multiplication commutes: (1 + r1)(1 + r2) equals (1 + r2)(1 + r1), so for a self-contained account with fixed-fraction sizing and no money moving in or out, the sequence of returns doesn't change where you end up. It only changes the path.
The word "only" matters, because almost nothing about a real trading account is self-contained. Four things break the symmetry, and each one converts a temporary path into a permanent outcome.
Margin is the most violent. Compounding forgives any order of returns; a margin clerk doesn't. Run a leveraged book and there's a drawdown level at which positions get liquidated whether you agree or not, and once liquidated, the recovery leg of the path happens without you. Two traders can hold the same positions through the same year, one entering the bad month with a profit cushion and one entering it fresh, and the identical market move is a drawdown for the first and a forced liquidation for the second. Same returns, different order, different lives. This is the sense in which ruin is usually path-triggered rather than edge-triggered: the strategy didn't stop working; the path just passed through a level where someone else's rules ended the trade.
Same Returns, Different Order
Withdrawals are next. If you live off the account, or investors pull capital, dollars leave at whatever the current equity is, and dollars withdrawn in a trough are gone at trough prices. An account that pays out a fixed amount per year can be ruined by a return sequence whose average is comfortably positive, if the bad years come first: the early drawdown plus the withdrawals eat the base, and the good years that follow compound off too little capital to matter. Reverse the order, good years first, and the same average return funds the same withdrawals forever. Sequence risk is the retirement-planning name for it, but it applies with full force to any trader paying rent from the P&L.
Your own rules break the symmetry too. Any sizing scheme that reacts to the equity curve, vol targeting, drawdown-based risk reduction, the mechanical de-risking discussed at the end of this lesson, makes the account path dependent on purpose. That's usually a trade worth making, but be aware you're making it: reactive rules mean the order of returns now changes your terminal wealth, where before it only changed your comfort.
And so does your behavior, the unofficial rule set. The trader who cuts size after three losing months and restores it only after two winning ones has built a path-dependent system without writing it down. Whether that improvised rule helps or hurts depends entirely on whether losing months predict more losing months, which for most strategies they don't. This is one more argument for the previous lesson's thesis: if the equity curve is going to change your sizing, decide the function in advance, at high water, rather than improvising it at the low.
10.11.4 Risk of ruin
Ruin has a clean classical model, and the model's lesson survives everything messy about real markets, so start clean.
You make repeated even-money bets, each risking one fixed unit of capital, with probability p of winning each bet, p greater than 0.5. You start with N units and you keep playing indefinitely. The probability that you ever lose all N units is:
In plain terms, your probability of total loss is a number less than 1 raised to the power of how many losses you can absorb. The edge sets the base; the sizing sets the exponent. And exponents beat bases.
With numbers: say your edge gives p = 0.55, a solid 55 percent win rate on even-money outcomes. The base is 0.45/0.55, about 0.818. Now vary only the unit size:
| Risk per bet | Units of capital (N) | Risk of ruin |
|---|---|---|
| 10% | 10 | 13.4% |
| 5% | 20 | 1.8% |
| 2% | 50 | 0.004% |
| 1% | 100 | about 2 in a billion |
Same trader, same edge, same markets. The only thing that changed across those rows is the fraction risked per bet, and the probability of ruin moved by seven orders of magnitude. This table is the most important one in this part of the course. Edge determines whether you should be playing at all. Sizing determines whether you survive long enough for the edge to matter. And because N sits in the exponent, sizing dominates: a mediocre edge at 1 percent risk outlives a great edge at 10 percent risk in nearly every world.
The formula also explains why doubling down is lethal on the math alone, before psychology even enters. Increasing bet size after losses shrinks N exactly when N is already depleted, driving the exponent down at the worst moment. Every martingale-flavored recovery scheme is a machine for converting a small probability of ruin into a large one, in exchange for smoothing the weeks when it doesn't trigger.
A necessary complication: you don't bet fixed units; you bet fractions of current equity, as every lesson since the vol targeting one has told you to. Fractional betting changes the mathematics of ruin completely: since you always risk a fraction of what remains, the account approaches zero asymptotically but never arrives. Literal ruin, in the textbook sense, becomes impossible.
That is no comfort, because literal ruin was never the thing that ends traders. What ends traders is functional ruin: the drawdown level at which you stop, whoever makes the decision. There are at least three versions and every account has all three. The broker's version is the margin call or the auto-liquidation level. The stakeholder's version is the point where investors redeem, the partner says enough, or the money is needed for something that isn't trading. And your own version is the depth at which you can no longer take the next signal at full size, at which point the strategy you backtested is no longer the strategy you're running. For most individual traders the third threshold binds first, and it's far shallower than they think: plenty of people who claim they could sit through 40 percent quit, in practice, near 25.
So define your ruin line honestly, as the shallowest of the three, and run the ruin logic against that line instead of against zero. Fractional sizing doesn't make the exponent argument go away; it just relabels the target. Risking 2 percent of equity per trade with a functional ruin line at a 30 percent drawdown gives you on the order of fifteen to twenty consecutive-loss-equivalents of room, and the same exponential machinery from the table applies. The Kelly lesson already gave you the continuous version of this: at full Kelly, the probability of ever seeing your account at fraction x of its start is simply x, which is exactly why nobody runs full Kelly. Fractional Kelly and modest vol targets are how you buy exponent.
One more honesty check on the classical model: it assumes your bets are independent and your loss per bet is capped at the unit you risked. Real trades gap through stops, real losses exceed their planned R, and real positions fail together. The last of those deserves its own section, because it's the mechanism behind most real-world ruin.
10.11.5 When everything becomes one trade
Consider your current positions. Six trades, say, each risking 2 percent of the account, each in a different market. On a normal day that's six bets and 2 percent of risk per bet, and the ruin table above says you're bulletproof.
In a stress event it's one bet risking 12 percent, and you built it by hand.
Correlation isn't a constant. The numbers you estimate from calm data describe how markets co-move when nothing much is happening, and they systematically understate co-movement in a crisis. You've seen the pieces already, in the microstructure lessons and in the vol targeting lesson's failure modes. Leveraged players hold overlapping portfolios. When a shock hits one market, the losses force deleveraging, and the deleveraging isn't choosy: positions get cut where liquidity exists, not where the problem started. Selling begets margin pressure elsewhere, which begets more selling. Market makers widen and pull, so every sale moves price further than it would have a week earlier. Assets that share no fundamental economics suddenly share a seller, and a shared seller is enough to make them correlate.
This is a reliable stylized fact in risk management: in a crash, correlations across risk assets lurch toward 1. Equities, credit, carry currencies, commodities with a growth bid, and crypto don't go down for their own reasons on those days; they go down for the same reason. Crypto is the extreme case, and you saw it in the perpetuals lessons: alts that trade on their own stories for months move at a correlation near 1 to BTC in every violent flush, so a "diversified" book of eight alt positions is one BTC-beta position with extra fees. Long-vol positions and being flat are, on the worst days, close to the only diversifiers that keep working.
The implication for ruin math is direct: your effective bet size is set by your stress-correlation matrix, not by your position list. Price every position by asking what it does on the day the S&P drops 5 percent, funding resets violently, and spreads triple. Positions that lose money on that day belong, for sizing purposes, to a single bucket, and the bucket's total risk is what belongs in the ruin arithmetic. If the bucket sums to 15 percent of your equity, you're running a 15 percent bet whose trigger you don't control, and no per-trade discipline changes that. The fix operates at the book level: cap the summed stress-day risk, hold genuinely offsetting exposures if you can find them, and treat any addition that loses on the same day as the rest of the book as an increase in your one big position rather than a new diversifying one.
This also reframes what drawdowns look like when they come. The account-killers are rarely a single bad trade. They're a bad week in which everything turned out to be the same trade, discovered simultaneously. The blowup case studies later in this part, the leveraged convergence book that died when its uncorrelated spreads converged to one trade, and the short-vol products that all rebalanced into the same close, are this section with names and dates attached.
10.11.6 Cut or sit
Sooner or later the drawdown arrives, and it brings one question: is this variance, or is the edge dead? Sit through variance and you get paid for the pain. Sit through a dead edge and you donate the rest of the account to whoever is on the other side. Cut a live edge at the bottom and you lock in the loss and miss the recovery your own system was about to deliver. From inside the drawdown, in the moment, the two are genuinely hard to distinguish, and anyone who claims they can always tell is describing hindsight.
You can't fully solve the identification problem, so don't try to solve it in real time. Build the response in advance, in three layers.
The first layer is mechanical and doesn't care about the diagnosis: reduce size as the drawdown deepens, on a schedule written at high water. A workable version looks like full risk down to a 10 percent drawdown, three-quarters risk from 10 to 20, half risk from 20 to 30, and a hard stop for reassessment at 30. The specific breakpoints matter less than their existence and their non-negotiability. This rule buys you the same protection regardless of which world you're in. If the drawdown is variance, you recover somewhat slower than full size would have, a real but bounded cost. If the edge is dead, every step down the schedule cuts the bleed, and the hard stop guarantees a dead strategy can't take more than a pre-chosen amount of your capital. You're paying a small tax in the good world to cap the loss in the bad one, which, per the recovery table at the top of this lesson, is overwhelmingly the right trade. The convexity of the recovery arithmetic carries the argument: giving back some upside near a 15 percent drawdown is cheap, and avoiding the trip from 30 to 50 percent is worth almost any price, because 50 costs five years.
The de-risking schedule has one cost worth naming: it makes your account path dependent (this is exactly the third mechanism from the path dependency section) and it means you recover from every drawdown at reduced size, so realized recoveries lag the arithmetic. Accept the cost. The alternative, full size all the way down, is how the account finds your functional ruin line.
The second layer is diagnostic, and it runs on one question: is this drawdown on-script or off-script? On-script means the losses are the kind your strategy was always going to produce, in the size and situation you expected. You did the groundwork for this back in the distributions lesson: if you know your win rate and payoff profile, you know what your losing streaks should look like. A system that loses 45 percent of its trades has roughly even odds of printing a seven-loss streak somewhere in a couple hundred trades; when that streak arrives, it isn't evidence of anything; it's the distribution doing what distributions do. Compare the current drawdown against the planning number from earlier, twice the backtest's worst. Inside that envelope, with losses arriving in the conditions where the strategy loses by design, you sit, at whatever size the mechanical schedule currently prescribes.
Off-script is different. A short-vol book losing money in a vol spike is on-script: that's the insurance paying out, and the premium you collected in the quiet months was compensation for exactly this week. The same book bleeding steadily through a calm, low-vol market is off-script: it's losing in the conditions where it's supposed to win, which means the mechanism you thought you were harvesting isn't there anymore. That distinction, losing where you expect to lose versus losing where you expect to win, is your most useful single diagnostic. Alongside it, ask whether the world changed in a way that removes your specific edge: the flow you were fading is gone, the structural buyer or seller left the market, the regime flipped in the sense of the regime lessons, the funding or premium you were collecting has compressed to nothing. Edges are usually somebody else's predictable behavior, and when that somebody leaves, no streak mathematics will bring the edge back.
The third layer is the exit and re-entry protocol, and it exists because the moment of maximum drawdown is the moment of minimum judgment. Decide now, in writing: at what drawdown do you stop entirely, what evidence would convince you the edge is dead rather than resting, and what has to be true, and for how long, before size comes back. A reasonable pattern after a hard stop is to keep generating signals on paper, and restore capital in steps only after the paper track behaves like the strategy again. The details are yours to set. The requirement is that they be set at high water: rules written at the low aren't rules, they're negotiations with yourself, and those negotiations tend to go the wrong way.
What's never on the menu is the opposite move: sizing up in the hole to get it back faster. The ruin section already showed you the mathematics, shrinking N exactly when N is depleted, and the previous lesson showed you the psychology. The urge will come anyway, dressed as confidence ("the edge is fine, so bigger size recovers faster"), and the answer is the same one the semi-systematic lesson gave for every such urge: the decision was made at high water, and it isn't being remade today.
Survival is the precondition for everything else in this course, but it's only the precondition. Once the sizing is honest and the drawdown plan is written, the remaining question is what happens when the plan is thrown out, and the cleanest way to learn that is from traders who threw it out in public. The next lesson tours five real blow-ups, each one a sound trade that met leverage, structure, or plumbing and did not survive the meeting, and each one carrying the specific rule that would have carried you through.
10.12 When derivatives blow up
Everything in this course so far has been mechanism: how funding pins a perp to spot, how dealers hedge gamma, how a futures contract converges to delivery, how a short vol book earns its premium. This lesson runs the same material in reverse. Five real blowups, each one a course concept failing in public, each one traceable to a specific lesson you've already read. Blowups are not history for its own sake; they are the only true out-of-sample test of risk thinking. A sizing rule that only works in the backtest is worthless; the rules worth having are the ones that would have carried you through these five weeks alive, and by the end of each case you'll see exactly what that rule was.
One point before the first case. None of these disasters came from a bad prediction. LTCM's trades were mostly right in the end. Short vol in 2018 had been printing money for years for a real reason. Oil demand did recover. The GameStop shorts were correct about the business. FTX's traders may have had perfectly fine market views. Every one of these blowups was a failure of sizing, structure, or plumbing, not of analysis. That is the recurring theme of this whole part of the course, and these cases are where it stops being abstract.
10.12.1 LTCM: leverage meets correlation convergence
Long-Term Capital Management was, on paper, the least likely fund in history to blow up. It launched in the mid 1990s staffed by veterans of the most successful bond arbitrage desk on Wall Street and two future Nobel laureates in economics. Its trades were relative value convergence, not gambling. Buy the cheap off-the-run Treasury, short the expensive on-the-run one, wait for the few basis points of spread to close. Short the rich leg of a swap spread against the cheap one. Buy Italian government bonds against German ones as European rates converged ahead of the euro. Each trade had a small, well-understood expected profit and, taken alone, small risk.
Small expected profit is the problem. A convergence trade earning a few basis points does nothing for investors at normal size, so the fund ran enormous leverage to turn basis points into returns. By early 1998 it held roughly 4 to 5 billion dollars of capital against a balance sheet on the order of 125 billion, call it 25 to 30 times levered, with derivative notional beyond the balance sheet running above a trillion dollars. The arithmetic from the drawdown lesson: at 25x leverage, a 4 percent adverse move in the asset base wipes out 100 percent of equity. The fund's models said the portfolio was far too diversified for its positions to move 4 percent against it together: dozens of trades, spread across bond markets, equity volatility, merger arbitrage, and currencies, in different countries, with historically low correlations between them. The measured portfolio vol was modest and the Sharpe ratio was spectacular, on the back of net returns above 40 percent in its best years. At the end of 1997 the partners were so confident in the machine that they returned a large slice of outside capital to investors, which shrank the equity base and pushed effective leverage higher on the same positions.
Then August 1998. Russia defaulted on its domestic debt, and global markets did what they do in a panic: everyone sold whatever was less liquid and bought whatever was most liquid, everywhere, at once. Through that lens, LTCM's diversification evaporates. Every single trade, in every market, was some version of the same position: short liquidity and short the flight-to-quality premium, long the cheap illiquid thing and short the expensive liquid thing. Dozens of trades that were uncorrelated in calm markets became one giant trade the moment the world wanted liquidity, and the correlations between them ran toward 1 exactly as described in the drawdown lesson's section on correlation spikes. Spreads that "could not" widen further widened further, because the other holders of the same trades were also levered and also getting margin calls, and their forced selling pushed the spreads against everyone who remained.
The other mechanism: at high leverage, your losses cause more losses. Mark-to-market losses trigger margin calls, margin calls force liquidation, liquidation in size moves the price against the rest of your book, which triggers more margin calls. This is the liquidation cascade from the crypto lessons, and the futures mechanics lesson showed that daily mark-to-market means there's no waiting it out. The fund lost around 1.9 billion dollars in August alone, roughly 45 percent of its capital in one month, and by late September its equity had fallen toward a few hundred million against a balance sheet still around 100 billion, leverage past 100x by arithmetic rather than by choice. The Federal Reserve brokered a rescue in which a consortium of major banks injected about 3.6 billion dollars for 90 percent of the fund, not out of charity but because a fire-sale liquidation of a trillion-dollar derivatives book would have hit every one of them.
The footnote: held to maturity, most of the trades converged. The consortium wound the book down at a modest profit. The fund was right and dead anyway, because leverage decides how long you're allowed to be right.
The rule that survives this one is a sizing rule with two clauses. Size the book against the stress scenario, not the historical covariance: assume that in a crisis every position that shares a hidden risk factor (short liquidity, short vol, long carry) moves against you simultaneously, and hold enough capital to eat that day. The diversification multiplier from the portfolio lessons is a fair-weather number; it's real in normal times and it's roughly 1 in a panic, so the leverage it justifies must be leverage you can carry when it disappears. And treat rising confidence as a risk signal. Returning capital and levering up after four good years was the fund's real decision error, made months before Russia mattered. It's the exact overbetting failure from the semi-systematic lesson, executed by some of the smartest people ever to run money.
10.12.2 Volmageddon: the short vol trade that shorted itself
Back in the VIX complex lesson you learned that VIX futures spend most of their life in contango, and Part 7 explained why: sellers of volatility insurance collect a premium because buyers will pay up for crash protection. Through the mid 2010s an entire retail product category grew up to harvest that premium mechanically. Inverse VIX exchange-traded products held a short position in the front two VIX futures and rebalanced daily to maintain constant minus 1x exposure. In a calm, grinding bull market this was a money machine: the products collected the roll-down of the contango curve every day, and the flagship inverse note roughly quintupled in the two years into January 2018. By early February 2018 the inverse products held on the order of two billion dollars, next to a larger pool of levered long VIX products doing the mirror-image trade.
The fatal detail is the daily rebalance. The algebra explains the whole event. A product that promises L times the daily return of an index must trade at each close to reset its exposure. The size of that trade, as a fraction of assets, is:
where r is the day's return on the underlying futures. For an inverse product, L = -1, so L^2 - L = 2, meaning the product must buy 2 dollars of VIX futures for every 1 dollar of assets per 100 percent move in the futures. And it buys rather than sells: when VIX futures rise, a short position grows beyond minus 1x and must be covered. The levered long products: L = 2 gives L^2 - L = 2 as well, and they also buy when the index rises, because their exposure must grow with their assets. Both sides of the product complex, the shorts and the levered longs, are forced buyers of VIX futures on a day VIX futures rally. The demand is mechanical, its formula was published in every prospectus, and anyone could compute it in advance from public AUM figures. Several desks did.
On February 5, 2018, after a run of vol-suppressed months and a sharp two-day equity selloff, the S&P dropped about 4 percent and VIX had its largest one-day percentage jump on record, roughly doubling from the high teens into the high 30s. A near 100 percent futures move in the formula above meant the product complex owed the market a rebalancing buy order on the order of its entire combined asset base, concentrated into the close and the post-close settlement window, in a VIX futures market whose normal depth was nowhere near that size. The buying pushed VIX futures higher, which increased the required rebalance, which pushed futures higher still. The products were the marginal buyer of the thing they were short, at any price, on a schedule everyone knew. Front VIX futures went close to vertical in the final hour and the after-hours session.
The inverse note lost more than 90 percent of its indicative value that evening. Its prospectus contained an acceleration clause allowing the issuer to terminate the note after a loss of that magnitude, the issuer used it, and holders were cashed out near the lows a few days later. Recovery was not slow; it was structurally impossible, because the product ceased to exist. A sister fund with a different legal wrapper survived the same loss, then cut its target exposure to minus 0.5x, which tells you what its issuer concluded about the original design. Meanwhile the S&P itself fell only around 10 percent peak to trough and recovered within months. The volatility premium the products harvested was real and reappeared within weeks. The harvesters were gone.
These map back to the lessons. The performance lesson showed that short vol strategies flatter their Sharpe until the skew shows up: five years of smooth gains and one 95 percent day is exactly that distribution. Part 7 showed that the premium is payment for insurance, and insurers who write more coverage than their capital survive one claim on. The dealer positioning lesson gave the general principle at work: any large player whose hedging or rebalancing is mechanical and predictable will have that flow traded against, and when the flow is buy-as-it-rises, it amplifies the move it is reacting to.
The rule: size every negative-skew position to its worst plausible day, not its average day, and treat 100 percent as the worst plausible day for anything short volatility. Concretely, a short vol sleeve should be small enough that its total loss overnight is an acceptable drawdown for the book, because that's the actual distribution you're holding, whatever the daily P&L has looked like for the past three years. And read the documents: a product that can be terminated at the issuer's option after a crash converts your theoretical recovery into a realized loss by contract.
10.12.3 Negative oil: delivery mechanics meet trapped longs
The futures mechanics lesson made a point of settlement types: cash-settled contracts converge to an index, physically settled contracts converge to the real thing, and the real thing has logistics. In April 2020 the logistics became the entire market.
The setup: pandemic lockdowns had collapsed oil demand by tens of percent within weeks while production adjusted far more slowly. Unwanted crude has to go somewhere, and for the WTI contract that somewhere is specific: delivery is by pipeline or in-tank transfer at Cushing, Oklahoma, a landlocked tank farm with finite capacity. Through April, Cushing storage filled toward its working limits and the remaining space was already leased. Anyone taking delivery of May-contract crude would receive 1,000 barrels per contract at a facility with effectively nowhere to put them.
The trap. The May 2020 contract stopped trading on April 21, with delivery obligations following for anyone still long. As the storage math became obvious, every long with no ability to take delivery (which is nearly every financial long) had to sell to someone before expiry. But the only natural buyers at expiry of a physically delivered contract are parties who can handle the physical, and they had no tank space either, or what space they had was suddenly the scarcest asset in the market. The exchange, seeing where this was heading, had announced days earlier that its systems could handle negative prices, a detail many market participants seem not to have priced in and some retail brokerages' systems literally couldn't process.
On April 20, the day before the last trade date, the selling met no bid. The May contract fell from the high teens through 10, through 5, through zero, and settled at minus 37.63 dollars per barrel. The price wasn't saying oil was worthless; the June contract settled that same day above 20 dollars, and Brent, which is cash-settled against a seaborne market with flexible storage, never went negative. The minus 37 was the price of the delivery obligation: at that moment, holding a claim on 1,000 barrels at a full tank farm was a liability, and the sellers were paying whoever remained to take it. A long who bought that morning near 18 dollars lost over 55 dollars a barrel by settlement, 55,000 dollars per contract on a position whose worst case they probably believed was 18,000. One large retail broker disclosed losses above 100 million dollars covering customer accounts that had gone negative, and a retail structured product in Asia that rolled its longs into the final days passed enormous losses to its holders.
Much of this course was sitting in plain sight beforehand. The contract specification (physical delivery, Cushing, 1,000 barrels, last trade date) is public and was covered in the commodity futures lesson. The storage economics that drive contango, from the term structure lesson, were screaming: the May-June spread had blown out to historic width, which is the market openly paying anyone with storage and openly punishing anyone without it. The trap wasn't hidden. It was written in the fine print and priced on the curve, and the people it caught were holding an instrument whose mechanics they had never read, treating a delivery contract as a price bet.
Two rules survive this one, both structural rather than sizing. Never hold a physically delivered contract into its endgame unless you can take or make delivery: roll or exit well before last trade and first notice dates, mechanically, on a calendar, not on a view. The entire final-week price action of a physical contract belongs to the logistics players, and a financial trader in that arena is out of their depth. And delete the assumption that any price is impossible. Zero wasn't a floor for oil. Limit moves, negative rates, and negative prices all live in the category of things that can't happen until an exchange notice says they can, and your worst-case sizing math has to use the contract's real boundaries, which may be none.
Negative Oil, April 2020
10.12.4 GameStop: a gamma squeeze in the wild
The GameStop episode of January 2021 gets told as a story about a message board versus hedge funds. Underneath the narrative it is the cleanest public demonstration ever staged of three lessons from this course operating at once: dealer gamma hedging, short-market microstructure, and crowd psychology as a price force.
Start with the positioning. GameStop was a heavily shorted stock, with reported short interest exceeding 100 percent of the float, which is possible because borrowed-and-sold shares can be borrowed and sold again. The microstructure lessons cover what that means structurally: a very crowded trade with a mechanical exit, because every short is a future forced buyer if the price rises enough, and the higher it goes the more forced they become. Short interest that large is dry fuel, and it was visible in public data for months.
The spark was options flow. Retail buyers concentrated in short-dated out-of-the-money calls, which are cheap in premium terms and, per the delta and gamma lessons, carry enormous gamma near expiry. The dealer's position from the options flow lesson: the market maker who sells a call hedges by buying delta in the stock, and as the stock rises toward the strike the call's delta grows, forcing the dealer to buy more. A call bought at a 0.30 delta commits the dealer to roughly 30 shares of hedge per contract; if the stock rallies until that delta is 0.60, the dealer must buy about 30 more shares per contract, into a rising market. Multiply by hundreds of thousands of contracts across a thin float and dealer hedging becomes a structural buyer whose demand grows with the price. That is short gamma at the market level: hedging that amplifies the move instead of damping it, the amplification regime from the dealer positioning lesson made visible. Each rally forced dealer buying, which forced short covering, which pushed prices toward the next strike, where fresh call buying reloaded the loop.
Then the crowd on top. The psychology lessons described herding, information cascades, and feedback loops as the raw material of bubbles, and here the feedback loop had a real mechanical engine underneath it, which is the most dangerous kind. Rising prices generated attention, attention generated buying, buying generated mechanical dealer and short-cover flow, which generated more rising prices. The stock went from under 20 dollars at the start of January to an intraday print of 483 on January 28, a move of roughly 25x in under a month in a company whose business had not changed. One prominent short-focused fund lost more than half its capital that month and required a multi-billion dollar injection to continue operating. On the other side, several retail brokerages restricted buying at the peak, not as a conspiracy but as plumbing: clearinghouse margin requirements on a stock that volatile spiked into the billions, the market plumbing lesson intruding on the narrative at the worst possible moment for the crowd. The price collapsed as the loop ran out of fresh buyers, in the classic bubble shape from the psychology lessons, and the late-cycle FOMO buyers took the losses that the early shorts had been forced to realize on the way up.
The instructive victim here is the short seller, because their failure was a pure sizing failure. Being short a stock has bounded profit and unbounded loss, an asymmetry that sat in the derivatives fundamentals lessons from the start. A short at 20 that goes to 480 loses 23 times its initial value, and no stop-loss reliably saves you through halts, gaps, and borrow recalls in a squeeze. The shorts were right about the company, which is the recurring theme: correct analysis, fatal structure.
The rule, in two parts. Any position with unbounded loss gets a hard size cap sized to a multiple-of-entry move against you, not to the daily vol, and if the crowding data (short interest, borrow cost, options volume concentration) says the exit is crowded, the cap shrinks further or the trade is expressed in defined-risk form: long puts or put spreads instead of short stock, where the worst case is the premium and is chosen in advance. And from the other side of the trade: when you find yourself in a reflexive winner where the price is rising because it's rising, the psychology lessons apply to you now, and the only exit that works is the mechanical one you wrote down before the loop started.
10.12.5 FTX: the risk that was never on the chart
The first four cases were market risk expressed through structure: prices moved, structures amplified, accounts died. The fifth case involves no adverse price move at all, which is exactly why it belongs here.
FTX was, by late 2022, one of the largest crypto derivatives exchanges in the world, widely treated as one of the most sophisticated. Traders held collateral there, ran perp books there, and parked profits there, implicitly modeling the exchange the way the futures mechanics lesson taught you to model a clearinghouse: a neutral, capitalized intermediary whose default risk rounds to zero. That model was wrong in a specific, structural way. A regulated futures clearinghouse stands between buyers and sellers with segregated customer margin, a default waterfall, and member capital behind it. An unregulated crypto exchange is just a company holding your money, and your balance there is legally closer to an unsecured loan to that company than to custody of your own assets. Whether that loan is good depends entirely on what the company does with the assets, which you can't see.
What this company had done, it emerged, was lend billions of dollars of customer assets to its affiliated trading firm, which had lost or locked them, with the hole papered over by holdings of the exchange's own token, an asset whose value depended on confidence in the exchange itself. In early November 2022 a leaked balance sheet exposed the affiliate's dependence on that token, a rival exchange announced it would dump its own large holdings of it, and the classic run began: withdrawals surged, the exchange paid out for about two days, then halted withdrawals and filed for bankruptcy within the week. The shortfall in customer assets was measured in billions. Customers with flat books, hedged books, or no open positions at all lost the same fraction of their balances as the most levered gambler on the platform, and its founder was later convicted of fraud. Position risk was irrelevant. The only variable that mattered was where the assets slept at night.
The exchange risk lesson in the crypto part told you all of this in advance, in general form: venue choice is a risk decision, an exchange balance is counterparty exposure, and the exchange's own token as collateral is a correlation trap (the collateral dies at the same moment the counterparty does, the exact wrong-way risk pattern). This case adds that counterparty risk does not diversify the way market risk does. Twenty uncorrelated trades on one venue are one position in that venue. Your book-level vol target, your fractional Kelly discipline, your careful sleeve construction: all of it multiplies by zero if the platform holding the account fails, and no line on any chart warns you first. Runs are nonlinear; the venue looks fine until the week it doesn't, because the run itself is what reveals the hole.
The rule is structural and boring, which is the point. Cap the fraction of total capital at any single venue whose failure you cannot survive, full stop, and treat that cap with the same non-negotiable status as the max-risk-per-trade rule from the semi-systematic lesson. Sweep profits off exchanges on a schedule instead of letting balances compound where they sit. Hold long-term assets in custody you control rather than on a trading venue. Prefer venues where customer assets are demonstrably segregated, and price the convenience of a single cross-margined account at what it actually costs: your entire balance in the bad state. None of this improves your expected return on any single trade, and all of it decides whether you are still present for the next thousand trades.
10.12.6 The shape all five share
Lined up, the cases share one structure. In each, a genuine edge or a defensible view (convergence spreads, the vol premium, cheap oil, an overvalued stock, plain trading profits) was attached to a structure that contained a hidden convexity against the holder: leverage that compounds losses, rebalancing that buys its own rally, a delivery obligation with no floor, unbounded short losses in a crowded exit, an unsecured balance at a fragile counterparty. In calm conditions the structure was invisible and the edge printed steadily, which is precisely what let the positions grow to fatal size. The payment arrived in every case as compensation for a tail the holder had stopped imagining. And in each case the surviving rule was known beforehand, written in a contract spec, a prospectus, a margin agreement, or a lesson like the ones in this course, and it was a rule about size or structure, never about forecasting. You could have known everything these traders knew about direction and it would not have saved you; you could have known nothing about direction, followed the sizing and structural rules, and walked away intact. That asymmetry is the entire argument of this part of the course, delivered five times by the market itself.
What remains is to put the constructive version together: how to combine return streams, set a book-level vol target, and run a daily process so that these rules become defaults you never have to think about instead of resolutions you try to remember under stress. That comes next, and it is where the whole risk framework becomes a routine you can actually run every morning.
10.13 Portfolio construction and process
The last lesson toured five crash sites, and every one had the same wreckage at the center: a single exposure grown larger than its owner's ability to survive being wrong. This closing lesson of the part is the constructive mirror image. Instead of one position grown too big, you run a book of several return streams deliberately kept small, chosen because they disagree with each other, and scaled as a group to a risk number you picked in advance. The first half is the arithmetic of why that works and how to build it. The second half is the operating manual, because a book isn't a spreadsheet exercise. It's something you run every morning, resize every week, and audit every quarter, and the traders who keep the diversification math working over a decade are the ones who turned the running of it into a routine too boring to break.
Back in the strategies part, the book-building lesson gave you the practical recipe for combining this platform's convex and concave sleeves. This lesson is the theory underneath that recipe: where the diversification benefit actually comes from, how much extra exposure it entitles you to, why the weighting scheme is inverse volatility rather than anything cleverer, and how the vol targeting machinery from earlier in this part stacks into two layers. Then it hands you the process: a concrete morning workflow through the platform's dashboards, the journal that sits on top of the trade-level journal you already keep, and the review cadence that decides when a sleeve is broken rather than merely losing.
10.13.1 Why the book beats its best sleeve
The claim that justifies all the machinery: a collection of mediocre strategies that disagree with each other beats a single good strategy, and by a wide margin.
Take two return streams, each running at 15 percent annualized volatility with a Sharpe ratio of 0.5, meaning each earns about 7.5 percent a year over cash. Put half your risk in each. The blend's return is the average of the two returns, 7.5 percent, unchanged. The blend's volatility is not the average of the two volatilities. For an equal split of two streams with correlation rho, the blend's volatility is:
The streams' individual wiggles partly cancel whenever they disagree, and the correlation controls how often they disagree. If the two streams are perfectly correlated (rho = 1), nothing cancels and the blend runs at the full 15 percent: you have one strategy under two names. If they're uncorrelated (rho = 0), the blend runs at 15 divided by the square root of 2, about 10.6 percent. Same return, two thirds of the volatility, so the Sharpe rises from 0.5 to about 0.71. At a realistic correlation of 0.2, the blend runs at about 11.6 percent and the Sharpe is about 0.65.
Nothing about either strategy improved. No signal got better, no edge got bigger. The improvement came entirely from the fact that the two streams take their losses on different days, so the combined equity curve is smoother than either input. Diversification is the one place in markets where you get paid without anyone paying you: the benefit comes from arithmetic, not from a counterparty, which is why it doesn't decay when other people discover it.
The general version: combining N equally sized, uncorrelated streams of equal Sharpe multiplies the Sharpe by the square root of N. Two uncorrelated streams give you 1.41 times the Sharpe, four give you double, nine give you triple. That progression is the honest reason multi-strategy funds exist. The catch is the word uncorrelated, and the whole rest of the construction section is about how far short of uncorrelated real streams fall and what to do about it.
A higher Sharpe buys more than a smoother ride, which connects to everything this part has taught about sizing. The Kelly lesson showed that the volatility a strategy can safely run at scales with its Sharpe, so when diversification lifts the book's Sharpe from 0.5 to 0.7, it also raises the ceiling on how hard you're allowed to push the whole account. The smoothing and the extra capacity are the same fact viewed from two sides, and the next section turns that fact into a number.
10.13.2 The diversification multiplier
This is where the free lunch becomes a position size. Suppose each of your sleeves, standing alone, is sized to run at your book target of 15 percent. Blend them and, because of the cancellation above, the blend realizes something below 15. The book is now running under budget: you chose 15 percent as the risk level your edge deserves and your stomach tolerates, and you are getting 11. The fix is the same division you learned in the vol targeting lesson, applied one level up: scale everything by target over realized. That scaling factor has a name when it arises this way, the diversification multiplier, and it answers a question that otherwise feels reckless: how can it be safe to run more gross exposure just because you added strategies?
The answer is that volatility is risk and leverage is not, a point the vol targeting lesson made with treasury futures and bitcoin. If the blend of five sleeves realizes 10 percent volatility while your target is 15, multiplying every position by 1.5 produces a book that swings exactly as much as one sleeve at target would have, while containing five different sources of return. You took the diversification benefit and spent it on exposure instead of on smoothness. You could equally leave the multiplier at 1 and pocket the benefit as a calmer account. Both are legitimate; what isn't legitimate is failing to make the choice consciously, because an undermultiplied diversified book is quietly running at half the risk its owner signed up for, and returns scale with risk taken.
How big does the multiplier get? For N equally risk-weighted sleeves sharing an average pairwise correlation rho, the blend's volatility relative to a single sleeve is sqrt(1/N + (1 - 1/N) * rho), and the multiplier is one over that. The static math:
| Average correlation | 2 sleeves | 5 sleeves | 10 sleeves | Many sleeves (limit) |
|---|---|---|---|---|
| 0.00 | 1.41 | 2.24 | 3.16 | unbounded |
| 0.25 | 1.26 | 1.58 | 1.75 | 2.00 |
| 0.50 | 1.15 | 1.29 | 1.35 | 1.41 |
| 0.75 | 1.07 | 1.12 | 1.14 | 1.15 |
The table carries two of the most practical facts in portfolio construction.
Across the rows, the benefit of adding sleeves saturates fast once correlation is real. At 0.25 average correlation, going from five sleeves to ten moves the multiplier from 1.58 to 1.75, and an infinite number of sleeves caps out at 2.0. The limit exists because with common correlation you can never diversify away the part of the risk the sleeves share; adding the eleventh moderately correlated strategy mostly adds work, not diversification. Five to eight genuinely distinct return streams capture most of what is available, which is a relief, because five to eight is also roughly what one person can operate honestly.
Down the columns, the correlation matters far more than the count. Two truly uncorrelated sleeves (multiplier 1.41) beat ten sleeves at 0.5 correlation (1.35). The practical instruction: your research effort should go into finding streams that are different in kind, not into multiplying variations of the same idea. Five momentum systems on five equity sectors are one stream with extra steps. A trend sleeve, a volatility premium sleeve, and a carry sleeve are three, because their payers, their mechanisms, and their bad days differ, which is the whole convex-plus-concave argument from Part 7 restated as a correlation matrix.
The mandatory caution, familiar from the vol targeting lesson and from the first crash site in the blowup tour. The multiplier is computed from correlations estimated in whatever regime you estimated them in, and correlations between risk-bearing strategies converge toward one in a crisis. The cancellation you levered against partially evaporates exactly when volatilities are also spiking, so the book's realized risk jumps on two axes at once. The defense is the same cap you already apply to vol-target leverage: bound the multiplier at something like 2 to 2.5 no matter what the matrix says, and treat any calculation asking for more as evidence that your correlation estimates are too flattering, not that your book is too safe. The table's uncorrelated row is unrealistic at book level; assume your true average correlation in stress is materially higher than your measured one in calm, and size to the stressed number.
10.13.3 Which correlations you can trust
Everything above consumed a correlation number as if it were a fact. It's an estimate, and estimates of correlation are noisy in a specific, dangerous way: they're most stable when nothing is happening and least reliable at the moment they matter. Three habits keep the estimation honest.
Measure correlation between strategy returns, not between markets. Your crypto positioning sleeve and your equity VRP sleeve trade different instruments, but that's not what makes them diversifying. What matters is whether their daily P&L streams move together, and strategy returns can be far less correlated than their underlying markets (a short-vol equity sleeve and a long-trend futures sleeve can both touch the S&P complex and still disagree most days) or far more correlated (two "different" strategies that are both secretly long carry). Compute it from the sleeve return streams themselves, over a window long enough to mean something; a correlation estimated from a month of daily returns is little more than noise.
Distrust the average, interrogate the tail. Two streams can show a daily correlation of 0.1 across a calm year and still take their worst week simultaneously, because the low average was earned in the small moves and the big moves share a driver. Part 7 made this concrete: VRP, earnings selling, and funding carry all pay their claims on the same deleveraging day, whatever their calm-period correlation says. So supplement the correlation matrix with a cruder, better question: for each pair of sleeves, in which state of the world do both take their large loss, and is it the same state? A book where every sleeve's disaster scenario is "equities gap down and vol spikes" is one trade with several names, and no multiplier should be applied to it.
Prefer structural reasons to statistical ones. The most durable low correlations are the ones you can explain with a mechanism rather than merely observe in a backtest. Trend versus short vol is the canonical pair: one loses small and often while collecting rarely and big, the other collects small and often while losing rarely and big, and the trend sleeve's best months have historically clustered in exactly the extended crises that hurt the insurance sleeves. That opposition comes from the shape of the payoffs, not from a lucky sample, so it's the kind of correlation you can build a book on. A measured minus 0.05 between two things you can't explain is the kind you can't.
10.13.4 Weighting the sleeves
Given a set of sleeves, something must decide how much risk each one gets. The tempting answers are the wrong ones, so clear those first.
Equal dollars is wrong because dollars are not risk: a 40 percent volatility crypto sleeve given the same capital as a 6 percent volatility allocation sleeve contributes roughly seven times the risk, and the book's fate becomes whatever the crypto sleeve does, which defeats the entire purpose of holding anything else.
Weighting by backtested Sharpe is wrong for a subtler reason. Expected returns and Sharpe ratios are the noisiest numbers in finance; the performance measurement lesson showed how many years of data it takes to distinguish a 0.5 Sharpe from a 0.8 with any confidence, and the backtesting lesson showed how reliably the measured number overstates the true one. Feed those noisy estimates into any optimizer that rewards them and the optimizer doesn't find your best strategy. It finds your most overestimated strategy and concentrates the book into it. Formal mean-variance optimization has exactly this failure mode: it's exquisitely sensitive to inputs you can't estimate, and in practice it functions as a machine for maximizing the impact of your estimation errors. The finding, replicated to the point of embarrassment across decades of allocation research, is that naive equal-risk weighting beats optimized weights out of sample far more often than any optimizer's marketing suggests, precisely because the naive scheme doesn't pretend to know the unknowable.
So the scheme I recommend uses only the one input you can estimate well. Volatility, as the vol targeting lesson established, is forecastable; means and Sharpes are not. Inverse volatility weighting sets each sleeve's weight proportional to one over its recent realized volatility, then normalizes the weights to sum to one. A sleeve realizing 12 percent gets two and a half times the capital weight of a sleeve realizing 30, and each ends up contributing a similar share of the book's risk. Equal risk contribution is the point: no sleeve dominates the book's swings just because it's jumpy, and no calm sleeve is wasted just because it's quiet.
There is a legitimate objection here: surely you know something beyond volatility. Some sleeves have longer live records, better-understood payers, more capacity. The disciplined way to express that knowledge is a tilt: a bounded multiplier applied to a sleeve's inverse-vol weight before normalizing, something like 0.7 for a sleeve you hold with less confidence and 1.5 for one you hold with more, set in the cold state, in writing, and revisited only on the review cadence. Tilts nudge the equal-risk baseline; they don't replace it. The bound is what keeps them from becoming Sharpe-chasing through the back door, the same way the forecast cap in the sizing chain keeps conviction from becoming overbetting. And a tilt has a second honest use: a new sleeve with two years of history should carry less weight than an old one with ten, whatever their volatilities, because the short record could be luck. Least evidence, least weight.
One worked warning about the instinct to overweight the best-looking sleeve, because it recurs constantly in practice. Suppose two versions of an equity momentum strategy: one backtests at a Sharpe of 1.1 with a beta to the index of 0.8, the other at 0.85 with a beta of 0.55. Standing alone, the first is better. Inside a book that already holds long equity exposure through an allocation sleeve and short-vol equity risk through a premium sleeve, the first version's extra return is mostly extra helpings of a risk the book already has; its higher correlation to the rest of the book shrinks the diversification multiplier and raises the count of sleeves that take their big loss on the same day. After vol targeting equalizes their risk budgets, what distinguishes them is correlation, and the lower-Sharpe, lower-beta version wins. The standalone Sharpe decides almost nothing at book level. Vol targeting and correlation decide the contribution.
10.13.5 The two-layer construction
Assemble the pieces and the whole build is two applications of the same idea. Vol targeting runs twice, once inside each sleeve and once across them.
Layer one lives inside each sleeve, and you built it in the sizing lessons. Every strategy runs its own internal risk management (per-position vol sizing, its own vol target or risk-parity rule) so that what it hands up to the book is a return stream of roughly known, roughly constant volatility. It's the layer that makes the second layer possible: inverse-vol weighting only produces stable weights if each sleeve's volatility is itself stable, which is exactly what internal vol targeting delivers.
Layer two is four mechanical steps across the sleeves.
First, weight: each sleeve's raw weight is its tilt divided by its trailing volatility, using volatility known as of yesterday, never today. That one-day lag is the same look-ahead guard the backtesting lesson drilled: today's weight must not be allowed to peek at today's move, or the backtest of the book will flatter itself in a way the live book cannot match.
Second, normalize: divide the raw weights by their sum so they add to one. The result is the equal-risk blend, tilted.
Third, blend: the book's daily return is the weighted sum of the sleeve returns.
Fourth, scale: compute the blend's trailing volatility, divide the book target by it, cap the result (2x is a sane cap for the reasons rehearsed twice now), and multiply every sleeve's weight by that scale. This last division is the diversification multiplier being applied live: when the sleeves cancel well the blend runs quiet and the scale grows; when correlations creep up the blend runs hot and the scale shrinks, automatically, without anyone forming an opinion.
The output to read is the final allocation per sleeve: its normalized weight times the book scale. That number is the fraction of the account deployed in each sleeve, and because the scale can exceed one, the fractions can sum to more than one. A book of calm, diversifying sleeves might deploy 150 percent of the account across them and still swing less than a single sleeve at 100 percent would. If the previous lessons did their job, that sentence no longer reads as reckless; measured in the units that matter, the levered diversified book is the conservative one.
A small example, end to end. Two sleeves: equity momentum realizing 12 percent volatility, crypto realizing 30. You tilt crypto up by 1.8 because you want meaningful crypto exposure. The blend has been realizing 10 percent, your book target is 15, the cap is 2.
Raw weights: momentum gets 1 / 0.12 = 8.33, crypto gets 1.8 / 0.30 = 6.0. Normalized: 8.33 / 14.33 = 0.581 for momentum, 0.419 for crypto. Book scale: min(0.15 / 0.10, 2) = 1.5. Final allocations: 0.872 of the account in the momentum sleeve, 0.628 in crypto, summing to 1.5, which is the leverage.
The tilt shows its limit here. You multiplied crypto's weight by 1.8 and it still ended up with the smaller share, because its volatility is two and a half times higher and inverse-vol dominates. Tilts nudge; the risk math rules. That hierarchy is a feature: your opinions get expressed in a form that can't quietly concentrate the book, the portfolio-level version of the bounded forecast scale, and a sleeve can only grab a dominant share of the book by becoming calm, never by exciting you.
Two operating notes complete the construction. Recompute the weights daily but trade them slowly: sleeve volatilities drift over weeks, so the weights are slow-moving by nature, and a buffer (leave allocations alone until they drift meaningfully from ideal, then trade back) keeps the rebalancing from bleeding edge into costs, exactly as it did for single-position vol targeting. And refuse thin estimates at this layer too: a blend volatility computed from a few weeks of a new book's returns deserves a scale of one, not a levered bet on a number the book barely knows.
10.13.6 The daily process
A book like this runs on maybe thirty minutes a day, but they have to be the same thirty minutes, in the same order, every day. The order matters because information has a hierarchy: risk first, context second, opportunities third, and nothing gets acted on until the plan is written. The workflow below is mapped to the platform's dashboards, for a swing trader running a mix of the course's convex and concave strategies.
First, the book itself, five minutes, and the only segment that can produce mandatory action. Yesterday's closes against every trailing stop: any breach is an exit today, no reopening the question. The book's vol estimate: if it stepped up sharply, the crisis clause from the vol targeting lesson triggers an immediate resize instead of waiting for the weekly pass. Open orders and expirations: anything expiring or settling today gets flagged now, not discovered at 3pm. This segment runs on rails; there are no decisions in it, only checks.
Second, regime, because the regime lessons established that every other signal you will read this morning means something different depending on the answer here. The SPX dashboard is built to be read in one pass: the regime score's current state, the VIX term structure (a front end trading above the back is the tell that stress is being priced now), credit spreads (the bond market's risk vote, slower and often more honest than the equity market's), and breadth. You're not extracting a trade from this screen. You're setting the interpretive frame: in a risk-on frame, positioning extremes are entries and dips are bought; in a stressed frame, the same extremes are warnings and the premium sleeves' new entries get extra scrutiny against their survival rails.
Third, the calendar. The events calendar for scheduled macro prints, the earnings calendar for names you hold options on or are screening. This decides what today is even allowed to be: the macro events lesson made the case that initiating fresh risk two hours before a major print is donating edge, and a day with a central bank decision at 2pm is a day whose plan says so at 8am.
Fourth, the positioning sweep, filtered to the markets you actually trade and to each strategy's schedule. Crypto is daily by nature: the dashboard's z-scores on open interest, funding, and liquidations, with the crowding lessons deciding what an extreme means in the current frame, and the risk appetite read consulted at its extremes. Futures positioning is weekly data, so the futures screener's week-over-week columns are a weekend and Monday job, not a daily one; rereading unchanged numbers daily only manufactures the illusion of new information. The equity screeners run on their strategies' cadence: the VRP screen on the weekly entry day for the premium sleeve, the skew and dark pool screens when you're hunting convex expressions. The discipline is to let the strategy's schedule, not your boredom, decide which screens get opened.
Fifth, the momentum cross-check. For every candidate the sweep produced, the cross-asset momentum read: whether it clears the directional threshold, whether it's stretched into the elevated zone, whether the signal line has turned. Confluence between a positioning extreme and a momentum turn was the core of the convex strategies; a candidate with the first but not the second usually goes on the watchlist, not in the book.
Last, the plan, written before anything is executed. Candidates with direction, forecast, level, and size from the chain; scheduled actions from the first segment; and, most days, the sentence "nothing new today," which is a complete and honorable plan. The writing is the point: it converts the morning from browsing into a decision record, it is what the review process audits later, and it marks the handoff after which the semi-systematic rails own everything. Until the plan is written the workflow is read-only. After it's written, the checklist from the mindset lesson takes over, and the dashboards are closed.
The shape to preserve if you compress this: risk checks always run, regime always precedes signals, and the session always ends in writing. A trader who checks stops after browsing screeners has the hierarchy inverted, and inverted hierarchies are how a morning of research turns into an impulsive trade with a post-hoc thesis.
10.13.7 The review cadence
Above the daily loop sit three slower loops. Each exists because some failure only becomes visible at its timescale.
Weekly, the book gets resized: compute the blend's trailing vol from your own equity changes, divide, cap, compare to current gross, trade the difference if it exceeds the buffer. The weekly premium strategies run their entry screens on their fixed day. New positioning data lands at the end of the week, so the futures review (screener changes, extremes, the bias reads on your markets) is a weekend job that produces Monday's candidates. Total cost: an hour.
Monthly, the journal gets audited. The trade-level compliance stats from the mindset lesson (violation rate and its trend) come first, because a book run by a trader whose violation rate is climbing is broken in a way no correlation matrix will show. Then the book-level entries: was the realized vol near target, did the scale hit its cap, did any deviation between planned and actual allocations creep in. Monthly is also when tilt and rule changes proposed during the month come off the shelf and get decided, cold, per the change process you already have.
Quarterly, the sleeves themselves go under review. The hard judgment here is telling a sleeve in an ordinary drawdown apart from a sleeve that is broken. The distributions lesson showed you can't tell from a quarter of returns alone. Real edges produce losing quarters routinely, and cutting every sleeve after a bad stretch guarantees you sell every premium at its bottom, which Part 7 identified as the mechanism by which premia persist. The fix is to move the judgment out of the moment entirely. When a sleeve enters the book, write its expectation card: its vol target, the drawdown depth consistent with that target (routine at around the vol number, unwelcome but unremarkable toward twice it, from the vol targeting lesson's calibration), the environments in which it should lose money, and the specific observations that would falsify it. Then the quarterly review judges the sleeve against its card, not against your mood. A trend sleeve bleeding through a choppy, trendless quarter is doing exactly what its card says. The same sleeve missing a large sustained trend that its rules should have caught has violated its card, and that is evidence of breakage even if the quarter's P&L happens to be fine. Losses inside the card are the cost of the stream. Behavior outside the card is the review trigger, and the response to a trigger is investigation and, if confirmed, a written removal decision at the review, never a mid-drawdown mercy killing.
The quarterly pass also rechecks the assumptions the construction leans on. It compares sleeve correlations against the values the multiplier was sized to, and upward drift is a signal to lower the cap or trim the most redundant sleeve. It checks each sleeve's realized vol against its internal target, because a sleeve whose internal targeting has slipped corrupts the layer above it. This is also the only forum where new sleeves get admitted: with a card, a small tilt, and the least-evidence-least-weight rule applied until a live record exists.
The process section of this lesson is deliberately longer than the math section deserves relative to its difficulty, because the failure does not come from the math. The construction is a few divisions. The failure mode of diversified books is almost never the arithmetic; it's the operator who stopped running the loops, let the weights drift, skipped the audit for a busy month, and rediscovered concentration the way the blowup lesson's cast did: suddenly. The routine is what carries the strategy.
Run all of this and the separate strategies from the earlier parts stop being a pile of trades and start behaving as one book, sized to survive being wrong and operated on a routine boring enough to keep.
10.14 A worked example: the all-weather core
This part has been a toolbox: read a distribution, measure performance honestly, tell a real backtest from a flattering one, size in volatility units, target a book-level risk, respect drawdowns, and combine streams into a diversified book. This closing lesson spends all of it on one real strategy, start to finish. It is very close to what I run as the all-weather core of my own systematic portfolio, and it is deliberately unexciting: no forecasts, no clever timing, just a handful of durable premia collected by rule and sized with the machinery from the last dozen lessons. It is also the futures version the opening lesson pointed to, the couple-hundred-thousand-dollar strategy that trades the same idea as the small-account ETF sleeve. If the earlier lessons were the parts, this is the assembled machine.
10.14.1 What it holds and the premia it harvests
The strategy trades four futures, each chosen to earn in a different kind of macro weather.
The S&P 500 (ES) collects the equity risk premium: the compensation for holding business and growth risk, the largest and most reliable long-run premium there is, and the one that pays best when growth is steady and calm. The 30-year Treasury (ZB) collects the bond term premium, the payment for holding duration, and doubles as the flight-to-safety hedge, because long bonds tend to rally in exactly the growth scares that punish equities. Gold (GC) is the real-asset leg: it yields nothing, so it is not a premium in the same clean sense, but it is paid across long stretches as a store of value when real rates fall, inflation runs, or confidence in paper money erodes, which is again a different regime from the first two. This bucket is gold-heavy because there is no single clean commodity future to hold the way an ETF holds a broad commodity basket, so the commodity slot folds into gold. The US dollar (DX) is the defensive, liquidity leg: the dollar bids in the risk-off, dollar-shortage episodes that hurt everything else, so it is cheap insurance that occasionally pays.
The all-weather logic ties them together. Growth and inflation can each rise or fall, four rough kinds of weather, and each asset is built to earn in a different one. Hold them together and something is usually working, so the book's ride is far smoother than any single leg would be. This is the diversification benefit from the portfolio-construction lesson, applied across macro regimes rather than across strategies.
Put it in the language of the returns part. This is not alpha. Nobody on the other side is being outsmarted. It is a basket of structural risk premia, exotic beta, harvested by a rule anyone could write down. The payers are equity investors, bond holders, and frightened people buying safety, and their motives have been paid for over a century. That is exactly why it is worth building a core around: it does not decay when it becomes known, because knowing about it does not remove the reason it exists.
10.14.2 The rules
Three rules, every one of them taken straight from this part.
Fixed weights. The four legs are held at fixed strategic weights that balance the book across regimes, roughly 40 percent equity, 15 percent bonds, 35 percent gold, and 10 percent dollar, with gold carrying a large share so the inflation and stress weather is well covered. The weights are set in the calm and left alone rather than optimized on the history, which is the anti-overfitting discipline the backtesting lesson argued for.
A trend gate. Each leg is held only while it is above a medium-term trend filter, and stepped to cash when it falls below. I am deliberately not naming the exact lookback, because the specific number is the least important part of the whole design and the edge does not live in it, a claim the robustness section makes concrete. What the gate does is two things: it harvests the trend premium described in the returns part, and, more importantly, it cuts the left tail, because a leg in a sustained bear market gets sidestepped rather than ridden all the way down.
A volatility target. The assembled book is scaled to a constant 20 percent annualized volatility, which is precisely the machinery from the volatility-targeting lesson: measure the book's recent volatility, divide the target by it, lever up when the book is quiet and cut back when it turns wild. The 20 percent is a risk-budget choice, not an edge. The ETF twin of these identical rules runs at about 6 percent volatility for a small account; dialing the target is how you move between the two.
Everything is rebalanced daily on yesterday's data, so today's position never peeks at today's move, the look-ahead guard from the backtesting lesson. It trades futures because futures are what let you reach a 20 percent target with capital efficiency, and, as the opening lesson warned, that is also what turns it into a couple-hundred-thousand-dollar strategy rather than a one-share-each one.
10.14.3 The backtest
All-Weather Core on Futures, 20% Vol Target
From 2011 to 2026, on ratio-adjusted continuous futures, the futures core at a 20 percent target earned a Sharpe of 0.97, a CAGR of 16.4 percent at 17.2 percent realized volatility, with a worst drawdown of 21.3 percent, a Sortino of 1.23, and a Calmar of 0.77. It was up on 54 percent of days and finished positive in 14 of 16 calendar years, the two red years being 2015 at about minus 13 percent and 2022 at about minus 7 percent. The strongest years came in broad trends, when several legs were on at once. Set against its own small-account twin, the tradeoff from the opening lesson is exact:
| ETF version | Futures at 20% target | |
|---|---|---|
| Instruments | 6 ETFs | 4 futures (ES, ZB, GC, DX) |
| Annualized volatility | about 5.7% | 17.2% |
| CAGR | 6.4% | 16.4% |
| Sharpe | 1.13 | 0.97 |
| Max drawdown | -8.1% | -21.3% |
| Capital to run | one share of each ETF | roughly $200k and up |
Read the Sharpe honestly. A number near 1.0 is not a headline, and that is the point the returns part made: honest multi-premium books land around there, and anything advertising a sustained 2 or 3 at this horizon is either overfit or hiding a tail. What this is, is a real, boring, diversified premium harvest that made about 16 percent a year without predicting anything.
10.14.4 Does the edge survive scrutiny
A single backtest is one draw from a process. The tests that matter ask whether the edge survives when you stress that draw, and three of them run directly on this futures series. The permutation, overfitting, and deflated-Sharpe checks after them need the validation harness, which is wired to the identical-rules ETF twin, so those are the twin's numbers; because the edge is the trend-gated allocation and the futures version only swaps instruments and raises the target, they still speak to whether the edge itself is real. I say which is which.
Start by fixing the rules on the first 60 percent of the history and judging them on the last 40 percent, data the rules never saw. The out-of-sample Sharpe comes in at 1.23, higher than the 0.77 in-sample, which is the reassuring shape: a curve-fit strategy shows the reverse, a strong in-sample number that falls apart once the data turns unfamiliar.
In-Sample vs Out-of-Sample
One split is still one split, so resample the daily returns with replacement 5,000 times and recompute the Sharpe each time. The realized 0.97 sits in the middle of a distribution running from about 0.56 at the 5th percentile to 1.39 at the 95th, and exactly one of the 5,000 resamples came back negative. A range whose downside stays that far above zero is evidence the result is not an accident of the single ordering of days that happened to occur.
Bootstrapped Sharpe Distribution
The bootstrap speaks to the edge; a Monte Carlo of the same resamples speaks to the pain. Take the worst peak-to-trough drawdown of each resampled path and the distribution runs deeper than the 21 percent that actually happened, with a median near 30 percent and a 95th percentile near 45 percent. That the resampled drawdowns are worse than the realized one is the honest read, not a red flag: shuffling the day order breaks up the calm stretches and strings bad days together into holes the real sequence avoided. It is the drawdown lesson made numeric, size for the 45 percent tail, not the 21 percent you have lived through so far.
Monte Carlo Drawdown Distribution
The harness tests on the ETF twin agree with all of this. A permutation test, shuffling the history thousands of times to destroy the signal, returns a p-value near 0.00, so the real result beats essentially all of the shuffles. The probability of backtest overfitting is 0.20, meaning the configuration that looked best in training stayed above median out of sample four times in five. The deflated Sharpe, which penalizes for how many variations were tried, is about 0.99, still almost certainly above zero after the haircut. And the edge sits on a broad parameter plateau, holding across a wide range of trend-filter speeds rather than a single lucky setting, which is exactly why the precise lookback does not matter and is not worth publishing. A result that works at only one number is noise; a plateau is a signal.
None of this is proof, because nothing is. It is the strategy clearing the bar the backtesting lesson set, the bar that kills most of what gets pitched.
10.14.5 What it costs to run
A 20 percent volatility target means real drawdowns: a 21 percent one has already happened and a worse one will, so the whole thing has to be sized small enough that a 40 percent stretch, the sort the Monte Carlo drawdown test flags, is survivable and held through when it comes, which is the drawdown lesson's entire message. The Sharpe is about 1.0, not 2; the roughly tenfold compounding on the chart is what 15 years and a 20 percent target do to a modest per-year edge, not evidence of magic. Because it is exotic beta and not alpha, it will not decay, but it also has no secret, and the real edge is the discipline to hold it, not the rules, which fit in three paragraphs. The trend gate is a blunt instrument: it sidesteps sustained bear markets but whipsaws in choppy, trendless stretches and lags the sharp V-shaped recoveries where a leg drops below the filter and rallies before it re-crosses, and that lag is the price of the tail protection. The diversification is a fair-weather number too: in a genuine liquidity crash the legs can fall together for a few days before the bond and dollar hedges do their work, the correlation-converges warning from earlier in the part. And it is capital-hungry: the futures version is the couple-hundred-thousand-dollar strategy from the opening lesson, and below that the ETF twin is the same edge at a lower volatility you can actually run.
That is one strategy, built entirely from this part and the ones before it: premia identified in the returns part, expressed in instruments from the asset-class parts, sized and risk-controlled with the machinery of this part, and checked against its robustness tests. There is no prediction anywhere in it, and it is deliberately the least exciting thing in the course, which is the lesson worth leaving on. A great deal of the work it takes to run a book like this, the reading, the summarizing, the record-keeping, the daily check of stops and calendars and screens, is exactly the kind of task that has recently become much cheaper to do well. The final part is about using AI as a tool for that work: what it is genuinely good at, where it quietly fails, and how to lean on it without handing over the judgment this course spent its length teaching you to keep.