Kelly Criterion Position Sizing: Math, Fractional Kelly and a Risk-Engine Implementation
Educational systems analysis only. Nothing here is investment advice or a recommendation to trade any instrument.
Most trading systems spend months on the signal and an afternoon on the question that decides whether the signal survives: how much capital goes behind each bet. Size too small and a real edge compounds so slowly it is not worth the operational risk. Size too large and a perfectly good edge produces drawdowns that trip every human and automated stop long before the long run arrives. Kelly criterion position sizing is the one rule that gives a principled answer, because it maximises the long-run exponential growth rate of capital instead of the average outcome of any single bet.
It also has an ugly secret. The full-Kelly bet is the edge of a cliff: it assumes you know your edge exactly, and you never do. Small errors in the estimated edge push realised growth down fast, and the path to the long run runs through drawdowns that most mandates cannot tolerate. Practitioners therefore run a fraction of Kelly and wrap it in hard limits.
This article derives the formula for binary and continuous cases, shows why drawdowns are brutal, quantifies the cost of estimation error, extends it to many assets with a covariance matrix, and shows how a risk engine enforces it.
What this covers: the derivation, geometric growth, fractional Kelly, estimation error, multi-asset sizing, a runnable synthetic simulation, and a risk-engine architecture that clips, scales, and monitors the output.
Context and Background
Kelly sizing comes from information theory, not finance. John L. Kelly Jr. published “A New Interpretation of Information Rate” in the Bell System Technical Journal in 1956 (volume 35, issue 4). His setting was a gambler with a private wire that gives a noisy hint about a race result. He showed that the maximum exponential growth rate of the gambler’s capital equals the information rate of the channel, and that the strategy achieving it is to bet a fixed fraction of current capital, chosen from the probabilities. The paper is about communication channels; the betting framing was a vivid illustration.
Edward O. Thorp turned the idea into practice. He used it in blackjack card counting and later in running a hedge fund, and his survey paper “The Kelly Criterion in Blackjack, Sports Betting and the Stock Market” remains the clearest practitioner account. Leonard MacLean, Thorp, and William Ziemba compiled the formal results, including the “good and bad properties” analysis of the criterion, in their book and associated papers. Those are the sources used for the claims below that are not derived from first principles.
The academic reception was split. Paul Samuelson famously objected that maximising expected log wealth is just one preference among many, and that Kelly has no special claim for an investor with different risk tolerance. That objection is correct and important: Kelly is not “the optimal bet”, it is the bet that maximises the long-run growth rate of wealth. When your real objective is capped drawdown, regulatory capital, or a client who will redeem at minus twenty percent, the right answer is a constrained or fractional version.
In modern systematic trading the same ideas appear under other names. Volatility targeting, risk parity, and mean-variance optimisation with a risk-aversion parameter are all close relatives. A continuous-time Kelly portfolio is, in fact, a mean-variance optimiser with a very specific risk-aversion coefficient of one. Seeing that equivalence is what lets Kelly slot into an existing portfolio-construction stack instead of living as a separate gambling heuristic.
Sizing does not live alone. The signal quality you assume is whatever your research and testing pipeline reports, which is why the realism of your event-driven backtesting engine directly determines how far the sizer can be trusted. The sizer’s output then must pass through limits enforced in a pre-trade risk engine. For the original source material, see Kelly’s 1956 paper and Thorp’s survey, both listed in Further Reading.
The Core Derivation: Why Maximising Log Wealth Gives a Size
Direct answer: Kelly sizing picks the fraction of capital that maximises the expected logarithm of wealth, which equals the long-run compound growth rate. For a bet that wins with probability p and pays b-to-1, the optimal fraction is f = p – (1-p)/b. For a continuous asset it is f = mu / sigma squared, where mu is the excess return and sigma squared the variance.

Figure 1: From expected log wealth to the binary and continuous Kelly fractions, and then to a chosen fraction of Kelly.
The diagram shows one idea with two branches. Both the binary formula and the continuous formula fall out of the same objective, maximising expected log wealth, and both are then scaled by a human-chosen fraction c before they reach production.
Why the logarithm: multiplicative bets and the long run
Suppose you repeatedly bet a fraction f of current wealth W. After n bets with k wins, wealth is W0 times (1 + b f)^k times (1 – f)^(n-k) for a bet paying b-to-1 on a win and losing the stake otherwise. Wealth compounds multiplicatively, so the quantity that behaves nicely over many bets is the logarithm:
log(Wn / W0) = k log(1 + b f) + (n – k) log(1 – f).
Divide by n and let n grow. By the law of large numbers k/n converges to p, and the per-bet growth rate becomes
G(f) = p log(1 + b f) + q log(1 – f), with q = 1 – p.
This is the geometric growth rate. It is not the average return of a bet; it is the rate at which a single capital path compounds. Two strategies can have the same expected return per bet and wildly different G, because variance drags growth. That gap is the whole story behind position sizing.
The binary formula
Differentiate G with respect to f and set it to zero:
G'(f) = p b / (1 + b f) – q / (1 – f) = 0.
Solving gives p b (1 – f) = q (1 + b f), so p b – p b f = q + q b f, and p b – q = b f (p + q) = b f. Therefore
f* = (p b – q) / b = p – q / b.
Check it with a coin that wins 55 percent of the time at even odds (b = 1). Then f = 0.55 – 0.45 = 0.10, so Kelly wagers ten percent of capital per bet. If the edge is zero (p = 0.5, b = 1) the formula gives zero, and if the edge is negative it gives a negative number, meaning do not bet this side. The formula is also bounded by one: f cannot exceed 1 because p – q/b is at most p when b is large.
The second derivative of G is strictly negative for f in (0, 1), so G is concave and f* is the unique maximiser. Concavity matters operationally: it means G falls on both sides of the optimum, and by the same small amount for small moves. That symmetry is what makes over-betting so much worse than under-betting in practice, as the next section shows.
The continuous formula
Real positions are not coin flips. Consider an asset whose returns over a small interval are approximately normal, with excess return mu per unit time and variance sigma squared per unit time. If you hold a fraction f of capital in the asset (f can exceed one if you use leverage), continuous-time wealth follows geometric Brownian motion with drift f mu – f squared sigma squared / 2 in the log. This comes from Ito’s lemma: the log of wealth gains a drift that is the arithmetic drift minus half the variance of the position.
So the growth rate is
g(f) = r + f mu – f squared sigma squared / 2,
where r is the risk-free rate and mu is the excess return over r. This is a downward-opening parabola in f. Setting the derivative to zero:
mu – f sigma squared = 0, so f* = mu / sigma squared.
The maximum growth rate is then g* = r + mu squared / (2 sigma squared), which equals r plus half the squared Sharpe ratio, because the Sharpe ratio is mu / sigma. This is a remarkably clean result: the best possible compound growth above cash is half the Sharpe ratio squared, and the leverage that achieves it is the Sharpe ratio divided by volatility.
A worked example with synthetic numbers: take an asset with 8 percent annual excess return and 16 percent annual volatility. Variance is 0.0256, so f* = 0.08 / 0.0256 = 3.125. Kelly says to hold more than three times capital in this asset. The growth rate is r plus 0.08 squared / (2 times 0.0256) = r plus 0.125, or 12.5 percent above cash. The Sharpe ratio is only 0.5, yet the formula calls for over 3x leverage. That should make you suspicious immediately, and it is the first sign that raw Kelly is not a production number.
The same parabola, read three ways
Because g(f) is a parabola, three properties follow from algebra alone, with no simulation needed:
- At f = f, growth is maximal: g = mu squared / (2 sigma squared) above r.
- At half of f, growth is 75 percent of g, since g(c f) = g times c (2 – c). For c = 0.5 this is 0.5 times 1.5 = 0.75.
- At twice f*, growth is exactly r, the same as cash, since c (2 – c) = 0 at c = 2. Beyond that, expected log growth is negative: you are betting a positive-edge strategy into ruin.
The third bullet is the important one. A trader who doubles the Kelly size earns the risk-free rate with the volatility of a leveraged equity position. That is why the true practical rule is “never exceed Kelly”, and why a safety margin is not conservatism but arithmetic.
From binary to continuous
The two formulas are the same object. For a small edge, a bet with win probability slightly above one half at even odds has mean per bet of 2p – 1 and variance close to one, so the binary Kelly fraction 2p – 1 matches mu / sigma squared with the same inputs. Use the binary form when the payoff is a known discrete outcome (prediction markets, a binary option, a fixed take-profit and stop-loss bracket). Use the continuous form for a position whose return distribution is approximately normal over your rebalancing interval. Real strategies sit between the two, and the mismatch is another source of error.
Deeper Analysis: Drawdowns, Fractional Kelly and Estimation Error
The formulas above are exact for known parameters. Production systems never have known parameters, and the penalty for that is lopsided. This section covers three consequences: the path risk of full Kelly, what a fraction buys you, and how estimation error eats the edge.
Why full Kelly has brutal drawdowns
Maximising the long-run growth rate says nothing about the path. For the continuous model, a diffusion approximation gives a clean result for the full Kelly bettor: the probability that wealth ever falls to a fraction x of its starting value (for x between 0 and 1) is approximately x itself. So the chance of ever being down 50 percent is about one half; the chance of ever being down 80 percent is about 20 percent. For a fractional bettor using c times Kelly, the same approximation gives x raised to the power (2/c – 1). At half Kelly that is x cubed, so the chance of ever halving is roughly 12.5 percent. At quarter Kelly it is x to the seventh, about 0.8 percent.
This is the Thorp and MacLean-Ziemba style result and it follows from the diffusion model; it is an approximation that holds for the idealised process and ignores fat tails, jumps, and discrete rebalancing. Treat the powers as a guide to magnitude, not a guarantee. The qualitative message is robust: full Kelly will, with high probability, put you through a 50 percent drawdown, and the long run you were promised only arrives for the investor who survives the path.
Drawdown is also a duration problem. At full Kelly, a long underwater period is normal. A strategy with a respectable edge can spend years below its high-water mark. If the allocator, the compliance function, or the trader’s own nerve has a ten percent trigger, full Kelly is not an option however good the signal.

Figure 2: A Kelly sizer feeding a risk-engine pipeline. The raw Kelly weights are shrunk, scaled to a volatility target, capped, and finally governed by drawdown state.
The pipeline in Figure 2 is the answer to the path-risk problem in practice. The Kelly sizer produces a target. Every stage after it exists because the target alone is too fragile to execute.
Fractional Kelly: trading growth for survival
Run c times Kelly and the algebra gives growth g = r + g* times c (2 – c) above cash, as shown earlier. Variance of the log-wealth path scales with c squared. The ratio of growth to path volatility therefore improves as c falls, and the tradeoff is favourable at the low end. Half Kelly keeps 75 percent of the growth for half the volatility of log wealth. Quarter Kelly keeps about 44 percent of the growth (0.25 times 1.75 = 0.4375) for a quarter of the volatility.
A table makes the relation concrete. The figures follow from the parabola and the diffusion drawdown approximation above; they are analytical, not empirical. Here x is the surviving fraction of starting wealth, so halving means x = 0.5 and an 80 percent loss means x = 0.2.
| Fraction of Kelly (c) | Growth vs full Kelly | Chance of ever halving (x = 0.5) | Chance of ever losing 80% (x = 0.2) |
|---|---|---|---|
| 1.0 | 100% | 50% | 20% |
| 0.75 | 94% | about 31% | about 7% |
| 0.5 | 75% | 12.5% | 0.8% |
| 0.25 | 44% | about 0.8% | about 0.001% |
Read the full-Kelly row twice: a one-in-two chance of a 50 percent drawdown, at some point, even when every input is perfectly known. Be careful applying the power law by hand; x is the surviving fraction, not the loss.
The practical conclusion matches what experienced practitioners report: half Kelly is a common upper bound, and many run a quarter to a third. These are conventions, not theorems. Choose c from your drawdown tolerance by inverting the power-law relation, then haircut it for estimation error, discussed next.
Estimation error: the edge you measured is not the edge you have
Everything above assumed mu and sigma squared are known. In reality you estimate them from a finite sample, and mu is the hard one. The standard error of an annualised mean return estimated from T years of data is sigma divided by the square root of T. With sigma = 16 percent and ten years of data, the standard error of mu is about 5 percent, against an assumed mu of 8 percent. The estimate is barely 1.6 standard errors from zero. A true edge of 4 percent or 12 percent would be entirely plausible.
Now note how Kelly responds. Because f is proportional to mu, a 50 percent overestimate of mu produces a 50 percent over-bet. Take the earlier asset where the true mu is 8 percent but the estimate is 12 percent. The bettor holds f = 0.12 / 0.0256 = 4.69, which is 1.5 times the true Kelly. From the parabola, growth is g times c (2 – c) with c = 1.5, or 0.75 of g*. A 50 percent error in the input costs 25 percent of the growth. A 100 percent overestimate (c = 2) costs all of it.
The asymmetry is the point. Underestimating mu by half gives c = 0.5 and also costs 25 percent of growth, but it does so with far smaller drawdowns. Overestimating by the same amount costs the same growth with dramatically larger drawdowns and a nonzero risk of ruin. So an unbiased noisy estimate is not safe to feed into Kelly; the loss function is lopsided, and you should bias the input downward.
Three corrections are standard:
- Shrink the forecast. Pull the estimated mu toward a prior, often zero or a cross-sectional mean. A Bayesian posterior mean is smaller in magnitude than the raw sample estimate, and it is the right number to feed in.
- Use the lower end of a confidence interval. Size off mu minus one standard error, for example, rather than the point estimate.
- Apply a fraction. A fraction c below one is the blunt but effective hedge against all unmodelled error, including the model being wrong.
Optimism bias compounds the problem. A backtest’s reported Sharpe ratio is inflated by selection among many tried variants and by unrealistic fills, so the in-sample mu is biased upward before sampling noise is even considered. This is why the quality of the simulation underneath matters so much, and why the event-driven backtesting engine should model fills, fees, and latency realistically before any number from it reaches a sizer.

Figure 3: Turning estimation uncertainty into a Kelly fraction. A wide confidence interval cuts the fraction; a tight one allows it to rise.
Figure 3 shows a rule a risk engine can implement: compute a standard error for the edge, form a confidence interval on the Kelly fraction, and if the upper end of that interval is more than double the point estimate, drop the fraction toward a quarter of Kelly. The thresholds in the figure (2x, 0.25, 0.5) are illustrative design choices, not derived constants.
Multi-asset Kelly through the covariance matrix
A real book holds many positions, and they are correlated. The continuous derivation generalises directly. Let w be the vector of position weights (fractions of capital), mu the vector of excess returns, and Sigma the covariance matrix of returns. Log-wealth growth is
g(w) = r + w’mu – w’ Sigma w / 2.
Differentiate with respect to w and set to zero: mu – Sigma w = 0, so
w = Sigma^(-1) mu, with maximum growth g = r + mu’ Sigma^(-1) mu / 2.
This is exactly the unconstrained mean-variance optimal portfolio with a risk-aversion coefficient of one. It is also why a Kelly sizer slots into an existing portfolio stack: the only difference from a standard optimiser is the scale, which Kelly fixes and a mean-variance user would otherwise tune.
Two practical consequences follow from the matrix form.
Correlation silently multiplies the bet. Take two assets, each with mu = 4 percent and sigma = 16 percent, so each alone has Kelly weight 0.04 / 0.0256 = 1.56. If they are uncorrelated, Sigma is diagonal, the weights are 1.56 each, and the gross exposure is 3.13. If they are 0.8 correlated, the inverse covariance reduces each weight. Working it through, w = Sigma^(-1) mu gives each weight mu / (sigma squared times (1 + rho)) = 0.04 / (0.0256 times 1.8) = 0.87, so total exposure is 1.74. Kelly correctly treats the second asset as largely a duplicate. A system that sizes each asset independently, using the single-asset formula, would hold 3.13 in a book where the true Kelly total is 1.74. That is an 80 percent over-bet created purely by ignoring correlation.
The inverse amplifies noise. Sigma^(-1) is dangerous when Sigma is nearly singular, which happens when assets are highly correlated or when the number of assets approaches the number of observations. Directions in return space with tiny variance get enormous weights, because the formula sees them as nearly riskless. Estimated covariance matrices have exactly this problem: the smallest eigenvalues are biased toward zero. The fix is the same family of tools used in any portfolio optimiser: shrinkage estimators (a blend of the sample covariance and a structured target such as a constant-correlation or factor model), a factor covariance model that restricts the number of free parameters, and a hard cap on any single position. Without those, the multi-asset Kelly output is a noise amplifier.
Add constraints and the closed form disappears. With a long-only restriction, a gross leverage limit, or sector caps, the problem becomes a quadratic program: maximise w’mu – w’ Sigma w / 2 subject to linear constraints. It is convex and solves in milliseconds for hundreds of assets. The result is still “Kelly subject to rules”, which is what a production system actually wants.
Margin, financing, and the mu in mu over sigma squared
The mu in the formula is the excess return over the borrowing cost for the leveraged portion. If you hold 3x of an asset, you pay financing on 2x of it, and the financing rate rises exactly when markets stress. A fixed r in the formula understates the real cost of the leveraged tail. Also, Kelly leverage assumes continuous rebalancing. If you can only rebalance daily and the asset can gap, a large fraction of the idealised growth is lost to jumps that the diffusion model does not contain. Both effects argue for the same direction as estimation error: lower fraction.
A runnable synthetic simulation
The following script simulates daily returns from a normal distribution with the asset parameters used above (8 percent annual excess return, 16 percent annual volatility) and measures how growth and drawdown vary with the fraction of Kelly. All returns are synthetic, drawn from a Gaussian, so there are no fat tails, no autocorrelation, no gaps, no costs, and no estimation error: the strategy knows its true edge. That makes the results a best case for every fraction.
import numpy as np
rng = np.random.default_rng(42)
mu, sig = 0.08, 0.16 # synthetic annual excess return and vol
n_years, steps, N = 20, 252, 4000
dt = 1 / steps
T = n_years * steps
f_kelly = mu / sig**2 # 3.125
print("full kelly", f_kelly)
z = rng.standard_normal((N, T))
r = mu * dt + sig * np.sqrt(dt) * z # daily synthetic returns
def run(leverage):
w = np.cumprod(1 + leverage * r, axis=1) # daily rebalanced
peak = np.maximum.accumulate(w, axis=1)
max_dd = (1 - w / peak).max(axis=1)
return w[:, -1], max_dd
print("c lev growth median_final p5_final median_maxDD P(maxDD>50%)")
for c in [0.25, 0.5, 1.0, 1.5, 2.0]:
lev = c * f_kelly
final, mdd = run(lev)
g = np.log(final).mean() / n_years
print(c, round(lev, 2), round(g, 4), round(np.median(final), 2),
round(np.percentile(final, 5), 2), round(np.median(mdd), 3),
round((mdd > 0.5).mean(), 3))
Running it once with seed 42 on the author’s machine produced the following output. These numbers come from this synthetic run only and will differ with another seed or sample size.
| Fraction c | Leverage | Log growth per year | Median final wealth (20y) | 5th pct final wealth | Median max drawdown | Share of paths with drawdown over 50% |
|---|---|---|---|---|---|---|
| 0.25 | 0.78 | 5.45% | 3.00x | 1.20x | 30.9% | 4.4% |
| 0.5 | 1.56 | 9.34% | 6.60x | 1.05x | 54.3% | 65% |
| 1.0 | 3.12 | 12.42% | 12.41x | 0.31x | 83.3% | 100% |
| 1.5 | 4.69 | 9.24% | 6.58x | 0.03x | 95.7% | 100% |
| 2.0 | 6.25 | -0.28% | 0.98x | 0.00x | 99.4% | 100% |
Three things stand out, and they match the analytical results. Growth peaks at full Kelly at about 12.4 percent per year, close to the theoretical 12.5 percent; growth at 1.5x Kelly equals growth at half Kelly (9.2 to 9.3 percent) in line with c (2 – c) being symmetric around one; and at 2x Kelly, growth is essentially zero. So the over-bet side of the curve is genuinely worse than the under-bet side: the same growth as half Kelly comes with a 95 percent drawdown instead of a 54 percent one.
The second thing to notice is the drawdown column. Even with perfect knowledge of the edge, half Kelly in this setup (which is 1.56x leverage on an asset with 16 percent volatility) produced a median maximum drawdown of 54 percent over twenty years, and two-thirds of paths saw a drawdown over 50 percent. The diffusion approximation in the previous section says the chance of ever halving is about 12.5 percent at half Kelly, which is far below the simulated 65 percent. The two answers do not contradict each other: the approximation concerns the probability of ever halving from the starting value, while this measure is the drawdown from a running peak, which is a much easier threshold to hit because the peak keeps rising. Mind which drawdown definition a stakeholder is using.
The quarter-Kelly row is the realistic one. At 0.78x leverage it earns 5.5 percent per year in log terms, a median 3x over twenty years, with a median maximum drawdown near 31 percent. For a strategy with a Sharpe ratio of only 0.5, that is about what you should expect, and it is a reminder that Sharpe 0.5 does not support aggressive sizing at any fraction.
Trade-offs, Gotchas, and What Goes Wrong
Kelly sizing fails in specific, repeatable ways. Most come from asking a clean formula to carry more than it can.
The edge is not stationary. Kelly assumes the same parameters generate every future bet. Strategy edges decay as they get crowded, regimes change, and volatility clusters. A sizer calibrated on a calm window will lever up exactly before volatility spikes, because its variance estimate is low. Using a volatility estimate that reacts fast (an exponentially weighted estimator, or the maximum of short and long windows) reduces this but cannot eliminate it.
Fat tails and gaps break the diffusion. The formula mu over sigma squared depends only on the first two moments. A strategy that sells insurance (short options, carry trades, merger arbitrage) has small mu, small apparent sigma, and a rare catastrophic loss. Sizing it at Kelly from the sample moments levers up the exposure that will eventually blow through the account. For such payoffs, use the discrete formula with an explicit loss scenario, or cap leverage by worst-case loss rather than by variance.
Correlations go to one under stress. The covariance estimated in quiet markets understates the co-movement during a selloff. A book that is diversified in the sample can behave like a single position when it matters. Stress the covariance matrix with a crisis correlation regime before accepting weights.
Feedback between sizing and impact. Kelly assumes you can trade the fraction at the quoted price. When position sizes are material to the market, trading costs and market impact reduce mu with size, and the optimal size is smaller than the frictionless formula says. Discrete lot sizes and minimum tickets also constrain small accounts.
Rebalancing churn. Holding a constant fraction of wealth means buying after gains and selling after losses (to keep the ratio fixed). With daily rebalancing and transaction costs, turnover from Kelly resizing alone can be a meaningful cost. A no-trade band around the target, rebalancing only when the position drifts by more than some tolerance, is standard.
The real objective is rarely log wealth. Allocators have redemption triggers; firms have capital and margin limits; individuals have liabilities. Kelly is the right answer to a question you may not be asking. That is the Samuelson critique, and it is why the sizer is the first stage of the pipeline rather than the last.
Compounding assumes reinvestment. The growth-maximisation argument relies on the bet size being a fraction of current capital. If a mandate fixes the dollar allocation, or profits are withdrawn, the compounding logic does not apply, and the sizing question reduces to something closer to mean-variance with fixed capital.

Figure 4: The runtime loop. The sizer proposes, the risk engine clips with reasons, the router executes, and the monitor feeds realised drawdown back to scale the fraction.
Figure 4 is the runtime counterpart to the pipeline in Figure 2. The sizer never talks to the order router directly. Every proposal passes through the risk engine, which returns not only the clipped weights but also the reasons for clipping, so that operators can audit which constraint was binding.
Practical Recommendations
Treat Kelly as a ceiling and a ranking device, not a target. The formula tells you the largest size worth considering and, more usefully, how sizes should relate across positions: proportional to edge, inversely proportional to variance, and discounted for correlation. The absolute level should come from your drawdown tolerance and your confidence in the inputs.
A risk engine implements this as a sequence of clipping stages, each with a logged reason code. The sizer computes shrunk Kelly weights from forecasts and a regularised covariance. A volatility-targeting stage then scales the whole vector so that forecast portfolio volatility equals a budget, which is itself the single most important parameter because it caps the leverage that mu over sigma squared would otherwise demand. Gross, net, per-name, and per-sector caps follow, solved as a constrained optimisation if clipping would distort the relative weights. A drawdown governor sits last: as realised drawdown from the high-water mark deepens, it reduces the fraction c in steps, and restores it only after recovery, which makes the system de-risk when its own edge estimate is most likely wrong. The limits themselves belong in a pre-trade risk engine that can reject an order in microseconds, while the slower exposure analytics resemble those in a real-time crypto derivatives risk engine, where leverage and liquidation thresholds make over-sizing especially costly.
A minimal sizing stage is short enough to show:
import numpy as np
def kelly_weights(mu_hat, cov_hat, c=0.25, vol_target=0.10,
gross_cap=2.0, name_cap=0.40):
"""Shrunk, scaled, capped Kelly weights. Educational sketch."""
raw = np.linalg.solve(cov_hat, mu_hat) # Sigma^-1 mu
w = c * raw # fractional Kelly
port_vol = np.sqrt(w @ cov_hat @ w)
if port_vol > vol_target: # volatility target
w *= vol_target / port_vol
w = np.clip(w, -name_cap, name_cap) # per-name cap
gross = np.abs(w).sum()
if gross > gross_cap: # gross exposure cap
w *= gross_cap / gross
return w
The function is a sketch: it omits the covariance shrinkage, constraint solver, reason codes, and no-trade band that a production version needs. Its order of operations is the point: fraction first, volatility target second, hard caps last, so that the hard caps are the only thing that can be binding in a failure.
Checklist for a first deployment:
- Estimate edge with shrinkage and use a conservative figure (point estimate minus one standard error is a reasonable start).
- Use a regularised covariance; never invert a raw sample matrix.
- Start at a quarter of Kelly or lower, and invert the drawdown power law to confirm the fraction fits your tolerance.
- Set a volatility target and hard exposure caps independent of the Kelly output.
- Add a drawdown governor that cuts the fraction as drawdown deepens.
- Add a no-trade band to control turnover.
- Log the binding constraint for every sizing decision.
- Re-estimate parameters on a schedule, and alarm when realised volatility departs from the assumed one.
- Backtest the sizer with realistic costs, then forward test with small capital before scaling.
Frequently Asked Questions
What is the Kelly criterion in simple terms?
The Kelly criterion is a rule for choosing what fraction of your capital to risk on a bet with a known edge so that your capital grows as fast as possible in the long run. It maximises the expected logarithm of wealth. For a bet with win probability p and payout b-to-1, the fraction is p minus (1 minus p) divided by b. A zero or negative edge gives a zero bet.
Why do traders use fractional Kelly instead of full Kelly?
Full Kelly assumes you know your edge exactly and accepts very large drawdowns; in the diffusion approximation the chance of ever losing half your capital is about 50 percent. Since the edge is estimated with error and over-betting is penalised more than under-betting, traders scale down. Half Kelly keeps about 75 percent of the growth rate with much smaller drawdowns, and quarter Kelly is common when estimation error is large.
What is the Kelly formula for stocks and other continuous assets?
For a single asset with excess return mu and variance sigma squared, the Kelly leverage is mu divided by sigma squared, which equals the Sharpe ratio divided by volatility. The maximum growth rate above cash is half the squared Sharpe ratio. For several assets, the weights are the inverse covariance matrix multiplied by the vector of excess returns, which accounts for correlation among the positions.
Can the Kelly criterion lose money?
Yes. Kelly maximises long-run growth under the assumption that your probability estimates are right. If you overestimate your edge, you over-bet and growth falls; at twice the true Kelly size, expected log growth equals the risk-free rate, and beyond that it is negative. Even with correct inputs, individual paths see deep drawdowns, and a strategy with a real edge can still go bust before the long run arrives if sized too aggressively.
How does a risk engine enforce Kelly position sizing?
The risk engine treats the Kelly output as a proposal. It applies a fractional multiplier, scales to a volatility target, checks gross, net, per-name, and sector caps, and adjusts the fraction according to drawdown state. Each clipped weight is logged with the binding rule. Pre-trade checks then reject any order that would breach hard limits, so the sizing model cannot override capital or margin constraints even if its inputs are wrong.
How do you estimate the inputs for Kelly sizing?
Estimate mu and the covariance from returns with a window long enough to be stable but short enough to be relevant, then shrink them. The standard error of a mean return is volatility divided by the square root of years of data, so ten years at 16 percent volatility leaves about 5 percentage points of uncertainty. Use shrinkage, factor covariance models, or the lower end of a confidence interval, and re-estimate on a schedule.
Further Reading
Internal:
- Event-driven backtesting engine architecture for algorithmic trading: how to produce edge estimates realistic enough to size against.
- Pre-trade risk engine architecture for low-latency trading: where hard limits are enforced on the order path.
- Real-time risk engine for crypto derivatives: leverage, margin, and liquidation limits in a continuous market.
External primary sources:
- J. L. Kelly Jr., “A New Interpretation of Information Rate,” Bell System Technical Journal, vol. 35, no. 4, 1956, pp. 917 to 926. See the biographical overview of John Larry Kelly Jr. for pointers to the original.
- E. O. Thorp, “The Kelly Criterion in Blackjack, Sports Betting and the Stock Market”, a practitioner survey with derivations and drawdown results.
- L. C. MacLean, E. O. Thorp and W. T. Ziemba, “Good and Bad Properties of the Kelly Criterion”, covering fractional Kelly and drawdown probabilities.
This article is educational systems analysis, not investment advice. All simulation figures are synthetic and illustrative; past or simulated performance does not predict real results.
By Riju — about
