Triple-Barrier Labeling and Meta-Labeling: A Financial ML Pipeline in Python
Most failed financial models do not fail at the modelling step. They fail one step earlier, at the moment someone decides what the target variable is. The default choice, “will the price be higher in five days?”, ignores volatility, ignores the path the price took to get there, and quietly manufactures overlapping samples that no cross-validation scheme was built to handle. Triple-barrier labeling is the answer popularised by Marcos Lopez de Prado in Advances in Financial Machine Learning (Wiley, 2018): label each observation by whichever of three barriers the price touches first, a profit-take, a stop-loss, or a time limit.
This matters now because cheap compute and open-source libraries have made it trivial to fit a gradient-boosted model to market data, and just as trivial to fit one that looks brilliant in a backtest and is worthless live. Labels, sample weights, and cross-validation are where that gap opens.
You will leave with a working mental model and runnable code: volatility-scaled barriers, event sampling, uniqueness weights, purged and embargoed k-fold, and meta-labeling as a second-stage filter. Every number in the code comes from synthetic data, and this is educational systems analysis, not trading advice.
What this covers: why fixed-horizon labels break, how the three barriers work, how overlap contaminates samples and cross-validation, how meta-labeling separates side from size, a complete pandas/numpy pipeline, and the backtest pitfalls that remain.
Disclaimer: This article is an educational analysis of machine-learning methodology. It is not investment advice, a recommendation to trade any instrument, or a claim that any technique described here produces profit. All code runs on synthetic, randomly generated prices.
Context and Background: Why Labels Are the Weakest Link
Supervised learning needs a target. In most of applied machine learning the target is given: a photo either contains a cat or it does not. In financial machine learning the target must be constructed from a price series, and every construction embeds assumptions about holding period, risk tolerance, and what counts as a “win”. Those assumptions are rarely examined, even though they determine what the model can learn.
The textbook construction is the fixed-time-horizon method. Take a return over the next h bars, compare it with a threshold tau, and label the observation +1 if the return exceeds tau, -1 if it falls below -tau, and 0 otherwise. Lopez de Prado covers this in Chapter 3 of AFML and then spends the rest of the chapter explaining why it is a poor default. The chapter’s central contribution is triple-barrier labeling, followed by meta-labeling, which is the subject of the second half of this post.
The book is organised so that each chapter fixes a failure created by the one before. Chapter 3 (Labeling) fixes the target; Chapter 4 (Sample Weights) fixes the overlap that Chapter 3’s labels create; Chapter 7 (Cross-Validation in Finance) fixes the leakage that overlap causes between train and test folds; and Chapters 11 and 12 cover the dangers of backtesting and backtesting through cross-validation, including combinatorial purged cross-validation. This article follows that same dependency chain, because skipping any link breaks the ones after it. You can read the book’s table of contents on the O’Reilly listing for Advances in Financial Machine Learning.
A practical note on tooling. The reference implementations that accompanied AFML were later productised by commercial vendors, and several open-source reimplementations exist, for example the finmlkit package, whose triple-barrier documentation describes event timestamps, volatility-based targets, bottom and top multipliers, a vertical barrier in seconds, and an optional side argument that switches the output from side labels (-1 or 1) to meta-labels (0 or 1). I use plain pandas and numpy below on purpose. The mechanics are short enough to write yourself, and writing them is the fastest way to understand where leakage hides. If you are building the surrounding research stack, the companion article on event-driven backtesting engine architecture for algorithmic trading covers the simulator that will consume these labels and signals.
One more framing point. The pipeline in this article is a methodology for estimating whether a signal carries information, not a trading system. A model with a validated edge still needs execution modelling, transaction costs, capacity analysis, and risk limits. Treat labeling as the first of those layers, not the last.
The Core Argument: Triple-Barrier Labeling as a Path-Dependent Target
Triple-barrier labeling assigns each event a label determined by the first of three barriers that the price path touches after the event: an upper horizontal barrier (profit-take), a lower horizontal barrier (stop-loss), and a vertical barrier (a maximum holding time). It is path-dependent, volatility-scaled, and mirrors how a real position is actually managed.

Figure 1: End-to-end pipeline. Events and volatility feed the barriers, labels get uniqueness weights, and the weighted samples are evaluated with purged cross-validation before a secondary model decides which signals to act on.
The figure shows the whole system in one view: two parallel preprocessing steps (volatility and event sampling) feed the labeler, the labeler feeds both the weighting step and the primary model, and the meta-labels produced from the primary model’s side calls are what the secondary model learns. Everything downstream of the labeler inherits its time windows, which is why those windows get so much attention later.
Why fixed-horizon labels fail and triple-barrier labeling fixes them
The fixed-horizon method has three structural flaws. The first is that it ignores volatility. A threshold of 1% means something entirely different when daily volatility is 0.4% than when it is 2%. With a constant tau, calm regimes produce almost all zeros and turbulent regimes produce almost all +1 and -1, so the class balance, and therefore what the classifier learns, is a function of the volatility regime rather than of any signal.
The second flaw is that it ignores the path. Suppose a position would have been stopped out at -3% on day two and then recovered to finish +1% on day five. The fixed-horizon label says +1. No real trader with a stop-loss would have experienced that outcome, so the label describes a trade that could not have happened. The model is being trained to predict a payoff nobody can collect.
The third flaw is subtler: sampling on a fixed calendar grid. Labelling every bar at every timestamp oversamples quiet periods, where nothing happens and consecutive labels are nearly identical, and undersamples bursts of activity where information arrives. Lopez de Prado argues for event-based sampling in Chapter 2, using filters such as the symmetric CUSUM filter, so that observations are drawn when something has changed rather than when the clock ticks.
The three barriers in triple-barrier labeling
Define an event at time t0 with entry price P0 and a volatility-based target sigma(t0). Two horizontal barriers sit at P0 (1 + pt sigma) and P0 (1 – sl sigma), where pt and sl are multipliers chosen by the researcher. A vertical barrier sits at t0 + H bars. The label is a function of whichever barrier is touched first, and the touch time t1 is recorded alongside it.
For a long-only or side-agnostic setup the label is +1 if the upper barrier is touched first, -1 if the lower barrier is touched first, and, when the vertical barrier is hit, either 0 or the sign of the realised return. AFML discusses both conventions; the sign-of-return variant keeps the dataset binary and avoids a large neutral class, at the cost of labelling small, noisy returns as if they were decisive. Which one you choose should depend on whether you want the classifier to learn “direction” or “decisive direction”.
Setting pt or sl to zero disables that horizontal barrier, which gives you useful special cases. With only a vertical barrier you recover something close to the fixed-horizon method; with a stop-loss and no profit-take you label by whether a position survives its holding period. The generality is the point: one parameterised procedure spans the whole family of labeling rules.

Figure 2: First-touch logic. Three barriers are set at the event, whichever is touched first fixes the label, and the touch time t1 is stored because it defines the label’s information window.
The first-touch rule in triple-barrier labeling has a consequence that is easy to miss: each label has a duration. A label that resolves in two bars uses information over a different span than one that runs the full horizon. That span, from t0 to t1, is the interval over which the label depends on future prices, and it is the object that sample weighting and purging both operate on. Treat (t0, t1) as part of the label, not as metadata.
Volatility-scaled barriers
Because triple-barrier labeling places barriers at multiples of an estimated volatility, the same procedure adapts across regimes and instruments. AFML estimates daily volatility with an exponentially weighted moving standard deviation of returns, and the pandas call is short: np.log(close).diff().ewm(span=50).std(). The span of 50 is a convention you should treat as a hyperparameter, not a law.
Scaling by volatility has a statistical interpretation. If returns over the horizon are roughly scale-mixture normal, a barrier at k standard deviations has a similar probability of being touched regardless of regime, which stabilises class balance. On a stationary, driftless random walk, symmetric barriers with a long enough horizon produce labels that are close to balanced. Asymmetric barriers, say a profit-take at 2 sigma and a stop at 1 sigma, deliberately tilt the base rate and the payoff ratio, and you should expect the label distribution to follow.
Two practical cautions apply. First, the volatility estimate must use only information available at t0. An EWM standard deviation computed with .ewm() is causal, but a centred rolling window is not, and it is an easy way to leak. Second, volatility at entry is not volatility during the holding period, so barriers will sometimes be touched more or less often than a naive calculation suggests. That is acceptable; it is a property of the labeling rule, and the model learns the rule as defined.
Event sampling with a CUSUM filter
Rather than label every bar, label events. The symmetric CUSUM filter accumulates positive and negative return deviations separately and fires an event when either sum crosses a threshold, then resets that sum. This is a structural-break detector in the quality-control tradition: it fires when cumulative movement since the last event is large relative to the threshold, so it naturally samples more densely in volatile periods.
Event sampling does two jobs for triple-barrier labeling. It reduces the number of near-duplicate observations, and it concentrates the dataset on moments when the process has actually moved, which is where a model could plausibly add value. It does not remove overlap, though. Two events a day apart with a ten-day vertical barrier still share nine days of future price. The next sections deal with that directly.
Overlap, Uniqueness Weights, and Purged Cross-Validation for Triple-Barrier Labeling
Overlapping labels from triple-barrier labeling violate the independent-and-identically-distributed assumption that standard machine-learning tooling relies on. Two defences address this: weighting samples by their average uniqueness (AFML Chapter 4), and purging plus embargoing training samples that overlap the test fold in time (AFML Chapter 7).
Concurrency and average uniqueness
Take the label windows (t0, t1) from the previous section. At each bar t, count how many label windows are active: call this concurrency c(t). A label that spans bars where c(t) is high shares its outcome with many others, so it contains less independent information than a label whose bars are covered only by itself.
The uniqueness of label i at bar t is 1 / c(t) if the label is active at t, and the label’s average uniqueness is the mean of 1 / c(t) over its own window. A label that never overlaps anything scores 1.0. A label that always shares its bars with three others scores 0.25. In the code below, the synthetic series has a mean average uniqueness near 0.6 with a ten-bar horizon and CUSUM-sampled events, and a peak concurrency of 4. Those figures come from one seeded run on random data, so treat them as illustrative of the mechanics, not as a benchmark.
Average uniqueness serves three purposes. As a sample weight, it tells the learner that clustered, redundant labels should count for less. As a diagnostic, its mean estimates the effective number of independent observations: a dataset of 1,000 labels with mean uniqueness 0.6 behaves, loosely, like roughly 600 independent ones. And as input to sequential bootstrapping (AFML Chapter 4), it drives resampling that draws observations in proportion to how much new uniqueness they would add, which pushes bagged ensembles toward less redundant draws.
AFML also describes weighting by the magnitude of the return attributed to each label (so large moves matter more than noise) and a time-decay factor that down-weights older observations. Both multiply onto the uniqueness weight. The finmlkit documentation mentions a related quantity: a max return-to-barrier ratio that can serve as a sample weight. Whichever weights you adopt, they should be computed from label windows, never from the test period.
Why ordinary k-fold leaks
Standard k-fold shuffles observations into folds and trains on all but one. With overlapping labels that is a leak. Suppose a test label spans bars 100 to 110. A training label spanning bars 95 to 105 shares ten days of the same price path, and the model that trains on it has effectively seen the answer. Serial correlation in features compounds this: a feature built from a 20-day moving average at bar 111 is nearly identical to the one at bar 110.
Lopez de Prado’s Chapter 7 describes this in terms of a train set contaminated by information from the test set, and proposes two corrections. Purging removes from the training set any observation whose label window overlaps the test fold’s time span. Embargoing additionally drops a short block of training observations immediately after the test fold, because features computed on those bars (with look-back windows reaching into the test period) can carry test information backward in the dependence structure. The embargo size in AFML’s example is a small percentage of total bars, on the order of one percent, which is a starting point you should adjust to your longest feature look-back and horizon.

Figure 3: Purged k-fold with embargo. Training labels that overlap the test span are removed, and an embargo block after the test fold is dropped before fitting.
Purging is the more important of the two. If your labels have a ten-bar horizon, then up to ten bars on either side of the test boundary contain overlapping samples, and only purging catches the ones at the start of the test fold, because those training labels end inside the test span. The embargo handles the end-of-fold side. Both are cheap to implement and both reduce, but do not eliminate, optimism in cross-validated scores.
Purged cross-validation versus walk-forward
Walk-forward testing avoids leakage by always training on the past and testing on the future, but it produces a single path through history, uses little data for early folds, and gives one noisy estimate. Purged k-fold uses every fold as a test set, trains on both earlier and later data, and yields k estimates, at the price of letting the model see “future” regimes during training. Whether that is acceptable depends on the question. For asking “does this feature carry information?”, purged k-fold is more data-efficient; for asking “what would deployment have looked like?”, walk-forward is more faithful. AFML Chapter 12 goes further and describes combinatorial purged cross-validation, which generates many backtest paths by recombining purged train/test splits, so that a performance statistic can be examined as a distribution rather than a single draw.
Meta-Labeling: Separating the Side Decision from the Size Decision
Meta-labeling is a two-stage structure introduced in AFML Chapter 3. A primary model decides the side (long or short). A secondary model, trained on the primary model’s outputs, predicts whether acting on that side will be profitable, and its probability drives the decision to bet or pass and the size of the bet.

Figure 4: Meta-labeling. The primary model supplies direction, the triple barrier is applied along that direction to produce 0/1 labels, and a secondary model learns when the primary signal is worth acting on.
How the labels change
In the side-agnostic setup, the barriers are symmetric and the label is -1, 0 or +1. Once a primary model supplies a side s in {-1, +1}, the barrier geometry changes: profit-take and stop-loss are defined relative to the position, so the path is multiplied by s before it is compared with the barriers. The label becomes binary: 1 if the trade ended with positive return under that side, 0 otherwise. In finmlkit‘s API, the same switch appears as supplying side, which turns -1/1 outputs into 0/1 meta-labels using a minimum-return threshold.
This reframes the learning problem. Instead of asking the model to predict direction from scratch, which is an extremely low signal-to-noise task, you ask a narrower question: given that the primary model has already said “long”, will it work this time? That is a classification problem with a different prior, different features (including the primary signal’s strength, recent hit rate, volatility regime, and spread), and a metric you can reason about.
Why this structure helps
The AFML argument is practical. First, the primary model can be anything: a moving-average rule, a fundamental screen, a discretionary analyst’s calls. The meta-model sits on top without needing to replace it, so meta-labeling lets machine learning wrap an existing process rather than demanding that the process be rebuilt.
Second, it controls the precision-recall trade-off in a deliberate way. Lopez de Prado recommends a primary model with high recall, accepting many false positives, and relies on the secondary model to raise precision by filtering them. Because the secondary model’s F1 balance of precision and recall is bounded by the primary’s recall, you cannot recover trades the primary model never proposed. That is a design constraint worth stating clearly: meta-labeling filters, it does not discover.
Third, it separates two decisions that are usually tangled. Side is a question about market direction; size is a question about confidence and risk. Mapping the secondary model’s predicted probability p to a bet size, for example by converting p into a z-score and passing it through a normal CDF, as AFML Chapter 10 discusses for bet sizing, produces graduated positions rather than all-or-nothing ones. Fewer, better-sized bets also mean fewer transactions, which matters when costs are real.
Fourth, it makes the model more interpretable to risk teams. A secondary model that predicts “probability this signal pays” can be inspected, calibrated, and challenged with ordinary diagnostics such as reliability curves, in a way a monolithic direction predictor often cannot.
A caution about what meta-labeling cannot do
It cannot manufacture an edge. If the primary signal has no information and the features carry none, the secondary model will fit noise, and a purged cross-validated AUC near 0.5 will tell you so. The synthetic example in the next section is built to demonstrate exactly that: a random-walk price with a moving-average primary model produces a cross-validated AUC barely above chance. Read that as the method working correctly, because it refused to find signal where none was planted.
Walk-Through: A Runnable Triple-Barrier Labeling and Meta-Labeling Pipeline
The following code is a self-contained pandas/numpy implementation. All prices are synthetic: a seeded random walk with slowly varying volatility, not market data. It needs only numpy and pandas. I implemented the secondary model as a weighted logistic regression in numpy so the example has no further dependencies; in practice you would substitute a regularised tree ensemble or similar and tune it inside purged folds.
Step 1: synthetic prices, volatility, and CUSUM events
import numpy as np
import pandas as pd
rng = np.random.default_rng(42) # synthetic data only
n = 2500
vol = 0.008 * np.exp(0.5 * np.cumsum(rng.normal(0, 0.05, n)) / np.sqrt(np.arange(1, n + 1)))
ret = rng.normal(0.0002, vol)
idx = pd.bdate_range("2016-01-04", periods=n)
close = pd.Series(100 * np.exp(np.cumsum(ret)), index=idx, name="close")
def daily_vol(close, span=50):
r = np.log(close).diff()
return r.ewm(span=span, min_periods=span).std() # causal: uses only the past
def cusum_events(close, threshold):
r = np.log(close).diff().dropna()
t_events, s_pos, s_neg = [], 0.0, 0.0
for t, x in r.items():
thr = threshold.loc[t] if isinstance(threshold, pd.Series) else threshold
s_pos, s_neg = max(0.0, s_pos + x), min(0.0, s_neg + x)
if s_neg < -thr:
s_neg = 0.0; t_events.append(t)
elif s_pos > thr:
s_pos = 0.0; t_events.append(t)
return pd.DatetimeIndex(t_events)
The volatility function is deliberately causal. Its min_periods argument means the first 49 values are NaN, and those bars are dropped from the event set so that no event has an undefined target. The CUSUM threshold here is itself volatility-scaled, so the filter fires roughly when cumulative movement reaches one daily standard deviation since the last event.
Step 2: the triple-barrier labeling function
def triple_barrier(close, events, trgt, pt_sl=(1.0, 1.0), horizon=10, side=None):
pos = {t: i for i, t in enumerate(close.index)}
rows = []
for t0 in events:
i0 = pos[t0]
if i0 + 1 >= len(close) or np.isnan(trgt.loc[t0]):
continue
i1 = min(i0 + horizon, len(close) - 1)
s = 1.0 if side is None else float(side.loc[t0])
path = s * (close.iloc[i0 + 1:i1 + 1] / close.iloc[i0] - 1.0)
up, lo = pt_sl[0] * trgt.loc[t0], -pt_sl[1] * trgt.loc[t0]
hit_pt = path[path >= up].index.min() if pt_sl[0] > 0 else pd.NaT
hit_sl = path[path <= lo].index.min() if pt_sl[1] > 0 else pd.NaT
cands = {"pt": hit_pt, "sl": hit_sl, "vb": path.index[-1]}
which = min((k for k in cands if pd.notna(cands[k])),
key=lambda k: (cands[k], k != "sl")) # tie: stop-loss first
t1 = cands[which]
rows.append((t0, t1, path.loc[t1], which))
out = pd.DataFrame(rows, columns=["t0", "t1", "ret", "barrier"]).set_index("t0")
out["bin"] = np.sign(out["ret"]).astype(int) if side is None else (out["ret"] > 0).astype(int)
return out
Three implementation choices deserve attention. The path is multiplied by the side, so a short position’s profit-take is a fall in price; this is the single change that converts side labels into meta-labels. Ties between barriers cannot happen on bar data except within a single bar, where the intrabar order is unknowable, so the code resolves them conservatively as stop-loss first; the finmlkit documentation, for its part, does not specify tie handling, so state your own rule explicitly. And the label is recorded with its t1, which the next steps require.
The conservative tie rule matters more than it looks. On daily bars a large move can cross both barriers within one candle. Assuming the profit-take came first systematically flatters the backtest, while assuming the stop came first systematically penalises it; the honest approach is to pick the pessimistic rule, or to use finer-grained data to resolve the order.
Step 3: the primary model and the meta-labels
vol_s = daily_vol(close)
events = cusum_events(close, vol_s.reindex(close.index).bfill().fillna(0.01))
events = events[events.isin(vol_s.dropna().index)]
# primary model: trend side from a 10/40 moving-average crossover
ma = close.rolling(10).mean() - close.rolling(40).mean()
side = np.sign(ma).replace(0, 1).reindex(events).dropna()
events = side.index
lab = triple_barrier(close, events, vol_s, pt_sl=(1, 1), horizon=10, side=side)
print(len(events), lab.barrier.value_counts().to_dict(), round(lab.bin.mean(), 3))
With seed 42 this run produced 1,027 events; 514 resolved at the profit-take, 501 at the stop-loss, and only 12 at the vertical barrier, with a meta-label base rate of 0.507. That near-balance is what symmetric, volatility-scaled barriers on a driftless walk should give, and it is a useful smoke test: if your own data gives a base rate of 0.9, either your barriers are badly placed or your primary model has remarkable skill or leaks. The few vertical-barrier exits also show a property worth knowing: with barriers at one daily volatility and a ten-bar horizon, almost every path resolves before the time limit.
Note the bfill in the CUSUM threshold call. It is applied only to the threshold series used to decide whether an early bar fires, and events from the warm-up period are removed immediately afterwards by the isin(vol_s.dropna().index) filter. A back-fill of volatility into a label target would be a leak; back-filling a throwaway threshold that is discarded is not. This is the kind of detail worth documenting in your own code review notes.
Step 4: uniqueness weights and purged k-fold
def avg_uniqueness(close_idx, t1):
count = pd.Series(0, index=close_idx, dtype=float)
for t0, e in t1.items():
count.loc[t0:e] += 1
u = {t0: (1.0 / count.loc[t0:e]).mean() for t0, e in t1.items()}
return pd.Series(u), count
def purged_kfold(t1, k=5, embargo_pct=0.01):
idx = t1.index
n = len(idx); emb = int(n * embargo_pct)
for fold in np.array_split(np.arange(n), k):
test_start, test_end = idx[fold[0]], t1.iloc[fold].max()
train = np.ones(n, bool)
train[fold] = False
# purge: drop train samples whose label window overlaps the test span
overlap = (t1.values >= np.datetime64(test_start)) & (idx.values <= np.datetime64(test_end))
train &= ~overlap
# embargo: drop the next `emb` samples after the test block
train[fold[-1] + 1: fold[-1] + 1 + emb] = False
yield np.where(train)[0], fold
u, cnt = avg_uniqueness(close.index, lab.t1)
print(round(u.mean(), 3), cnt.max())
folds = list(purged_kfold(lab.t1, 5, 0.01))
print([(len(tr), len(te)) for tr, te in folds])
On this run the mean average uniqueness was 0.598 and peak concurrency 4. The five folds each tested on about 205 events and trained on roughly 810 to 820, which is a modest loss from the 822 training samples a plain 80/20 split would give. That small gap is typical of ten-bar horizons on event-sampled data; with longer horizons or denser events, purging removes much more, and you should check how much it removes rather than assume.
The purge condition is the compact heart of the method. A training sample overlaps the test span if its label ends at or after the test start and it begins at or before the test end. Both conditions are needed, since a training label that began long before the test fold but ends inside it is just as contaminating as one that begins inside it.
Step 5: the secondary model and its out-of-fold score
feat = pd.DataFrame({"vol": vol_s, "ret5": np.log(close).diff(5),
"ma": ma / close, "side": side}).reindex(lab.index)
feat["ret5s"] = feat.ret5 * feat.side # signed by the primary side
feat["mas"] = feat.ma * feat.side
X = feat[["vol", "ret5s", "mas"]].dropna()
X = (X - X.mean()) / X.std() # NOTE: global scaling is a mild leak, see below
y = lab.bin.loc[X.index].values
w = u.loc[X.index].values
def fit_logit(X, y, w, lr=0.1, it=500, l2=1.0):
Xb = np.c_[np.ones(len(X)), X]; b = np.zeros(Xb.shape[1])
for _ in range(it):
p = 1 / (1 + np.exp(-Xb @ b))
b -= lr * (Xb.T @ (w * (p - y)) / w.sum() + l2 * np.r_[0, b[1:]] / len(y))
return b
def predict(b, X):
return 1 / (1 + np.exp(-(np.c_[np.ones(len(X)), X] @ b)))
def auc(y, s):
r = pd.Series(s).rank().values; n1 = y.sum(); n0 = len(y) - n1
return (r[y == 1].sum() - n1 * (n1 + 1) / 2) / (n1 * n0)
t1 = lab.t1.loc[X.index]
ps = np.full(len(X), np.nan)
for tr, te in purged_kfold(t1, 5, 0.01):
b = fit_logit(X.values[tr], y[tr], w[tr])
ps[te] = predict(b, X.values[te])
print("AUC", round(auc(y, ps), 3), "base rate", round(y.mean(), 3))
The out-of-fold AUC on this run was 0.516 against a base rate of 0.507, and the model would have kept about 57 percent of signals at a 0.5 threshold. On a random walk that is exactly the expected result: no real edge, so a score within noise of 0.5. The pipeline did its job by not inventing one. If you replace the synthetic series with data that has real structure, the same code gives you a purged, uniqueness-weighted estimate of how much of it the features capture.
Note the comment on scaling. Standardising with the full-sample mean and standard deviation lets test-fold statistics influence the training features. The effect is small for z-scoring, but a strict pipeline fits the scaler inside each training fold. Treat every transformation that estimates a parameter from data, including scalers, imputers, feature selectors, and hyperparameter searches, as part of the model that must be fitted inside the purged loop.
Deeper Analysis: Choosing Barriers, Horizons, and Thresholds
Triple-barrier labeling has four free parameters: the profit-take multiple, the stop-loss multiple, the horizon, and the volatility span. Each one is a modelling decision with consequences, and each one is a degree of freedom that can be overfit.
The barrier ratio sets the payoff profile
With symmetric barriers (pt = sl = 1), the base rate is near one half on a driftless walk, and a classifier needs a genuine edge to move it. With pt = 2 and sl = 1, the profit-take is twice as far away as the stop, so on a driftless walk the profit barrier is hit less often. Under a simple continuous random-walk approximation, the probability of reaching +2 before -1 is about one third, while the payoff ratio is two to one, giving zero expected value before costs. That gambler’s-ruin result is a property of the approximation, not of any market; it shows why base rates and payoff ratios must be read together. A model that raises the hit rate above one third at that geometry has found something, and the label distribution, not the accuracy score, is where you should look first.
For meta-labeling, the asymmetric case is the natural one, because the secondary model’s job is to estimate whether the trade will pay under the strategy’s actual exit rules. Use the barriers your execution will use. If the production stop is 1.5 sigma and the target is 3 sigma, label with those. A label produced with barriers you do not trade is a label for a different strategy.
Horizon length and overlap in triple-barrier labeling
The vertical barrier controls both the label’s information window and the amount of overlap. Doubling the horizon roughly doubles the concurrency for a given event density, which lowers average uniqueness and widens the purge. It is a lever with a cost on both sides: a short horizon yields noisy labels dominated by one or two bars, and a long horizon yields few effectively independent samples. A useful habit is to log three numbers per labeling run: the barrier-touch mix, mean average uniqueness, and the fraction of training samples purged per fold. If purging removes a quarter of your training data, your horizon and event density are in tension.
Class imbalance and the meta-label
Meta-labels are often imbalanced, because a high-recall primary model emits many signals that fail. When positives are, say, a third of the data, a classifier can achieve high accuracy by predicting zero for everything. Evaluate with metrics that respect the prior: precision and recall at the chosen threshold, AUC, log-loss, and calibration. Accuracy is the wrong yardstick, and so is raw F1 without reference to the trade economics.
Threshold selection deserves its own protocol. Choosing the probability cut-off that maximises backtest return on the same data used to evaluate the model is a form of overfitting. Choose the threshold inside the purged training loop, or fix it from economics (for instance the probability at which expected payoff after costs is zero) before looking at test results.
Reading the secondary model’s probabilities
A well-behaved secondary model produces probabilities that mean what they say: of signals scored 0.6, about 60 percent should pay. Check this with a reliability table binned out-of-fold. If the probabilities are systematically over-confident, bet sizing derived from them will over-bet, and the loss will show up as drawdown, not as an obviously wrong accuracy number. Calibration (isotonic or Platt-style scaling, fitted inside the loop) is cheap insurance, and it matters most exactly when you use the probabilities to size positions rather than merely to filter them.
Trade-offs, Gotchas, and What Goes Wrong
Triple-barrier labeling, with its companions, is a set of defences against specific failure modes, and it introduces a few of its own.
Label look-ahead through parameters. Barrier multiples, horizon, volatility span, and CUSUM threshold are all hyperparameters. If you tune them to maximise out-of-fold performance and then report that performance, you have run a search and reported its best result. Every configuration you try is a trial, and the number of trials belongs in your evaluation. Bailey and Lopez de Prado’s work on the deflated Sharpe ratio and on the probability of backtest overfitting addresses exactly this selection bias, and both are worth reading before you trust any single best backtest.
Purging is not a cure for everything. It removes label-window overlap. It does not remove regime leakage (the model seeing a volatility regime in training that also appears in test), cross-sectional leakage (a market-wide shock appearing in several instruments at once, so purging one instrument’s folds leaves correlated neighbours in training), or feature-construction leakage (a normalisation computed across the full sample). Multi-asset datasets need purging that is aware of time across all assets, not per asset.
The tie and gap problem. Daily bars hide intrabar order, and overnight gaps can jump through a stop. The label records the barrier as touched, but the realised fill would be worse. If your strategy trades through gaps, label returns should use the realised exit price, not the barrier level; the code above does this by recording the path return at the touch bar, not the barrier value.
Meta-labeling inherits the primary model’s blind spots. Because it can only filter, it cannot trade the opposite side or find opportunities the primary missed. If the primary model’s recall is low in some regime, the secondary model never sees that regime’s winners, and the combined system has a hole there. Measure primary recall explicitly, ideally against a looser label definition.
Stacked overfitting. Two models mean two sets of hyperparameters and two chances to overfit, and the secondary model trains on a smaller, filtered dataset. Keep it simple: few features, strong regularisation, shallow trees, and out-of-fold probabilities for everything.
Anti-patterns worth naming. Shuffled k-fold on overlapping labels; reporting accuracy on imbalanced meta-labels; computing volatility with a centred window; choosing the threshold on the test fold; fitting the scaler on all data; evaluating a single backtest path as if it were a distribution; and treating a flawless backtest as evidence rather than a warning. AFML’s Chapter 11 is titled, in part, around the idea that a flawless backtest is probably still wrong, and that is a sound default posture. Another structural limit is sample size: an event-sampled, ten-bar-horizon dataset for one instrument over a decade may hold only a thousand or so labels, and with a mean uniqueness of 0.6 the effective sample is smaller still. That is not enough to support many features or complex models.
Finally, none of this settles whether a signal survives costs. Labels built from barrier-touch returns are gross of commissions, spread, slippage, and market impact. Subtract realistic costs from the returns before you decide what counts as a “1”.
Practical Recommendations
For triple-barrier labeling, start with the question the model is meant to answer, then choose the labeling rule that matches how the position would really be managed. Build the pipeline in the order the dependencies run: events and volatility first, then barriers, then uniqueness, then purged splits, then models. Resist the temptation to jump to the model, because nothing downstream can repair a bad label.
For a first implementation of triple-barrier labeling, use symmetric barriers at one daily volatility, a horizon of five to twenty bars, an EWM volatility span you can defend, and a CUSUM threshold close to the barrier width. Look at the barrier mix and base rate before fitting anything. Use sample weights from average uniqueness, purge by label end-time, and embargo on the order of one percent of bars, adjusted for the longest feature look-back. Treat the meta-model as a small, heavily regularised classifier and judge it on out-of-fold AUC, log-loss, and calibration rather than accuracy.
When you move beyond a toy, add combinatorial purged cross-validation so that you look at a distribution of outcomes, record how many configurations you tried, and test the final system in an event-driven simulator that models costs and latency, such as the architecture described in the event-driven backtesting engine guide.
A short checklist before you trust any result:
- Labels carry their own
(t0, t1)window, and every split uses it. - Volatility and every feature use only data available at
t0. - Scalers, selectors, and thresholds are fitted inside the training fold.
- Barrier tie handling is explicit and pessimistic.
- Barrier multiples match the strategy’s real exit rules.
- Average uniqueness, purge fraction, and barrier mix are logged for every run.
- The number of configurations tried is recorded and used to discount the result.
- Returns are net of realistic costs before they are labelled.
Frequently Asked Questions
What is triple-barrier labeling in financial machine learning?
Triple-barrier labeling assigns each trading event a class based on which of three barriers a price path touches first: an upper profit-take, a lower stop-loss, or a vertical time limit. Barriers are normally scaled by estimated volatility. It was proposed by Marcos Lopez de Prado in Advances in Financial Machine Learning as a path-dependent alternative to fixed-horizon return labels, because it reflects how positions are actually exited.
What is meta-labeling and how does it differ from primary labeling?
Meta-labeling adds a second model on top of a primary model that already chooses trade direction. The primary model’s side is applied to the triple-barrier path, producing a binary label of whether the trade paid. The secondary model predicts that outcome, filtering weak signals and sizing bets from its probability. It improves how a signal is used; it does not create new signals the primary model never produced.
Why does ordinary k-fold cross-validation fail on financial data?
Financial labels typically span several bars, so neighbouring labels share future price information. Random k-fold puts one label in training and an overlapping label in test, which leaks the answer and inflates scores. Features built from rolling windows add serial correlation. Purged cross-validation removes training labels overlapping the test span, and an embargo drops observations just after it, reducing that leakage.
How do I choose the vertical barrier and the profit-take and stop-loss multiples?
Match them to how you would actually trade. Choose multiples of volatility equal to your real target and stop distances, and a horizon equal to your maximum holding period. Then inspect the outcome: barrier-touch mix, class balance, and mean average uniqueness. Treat these parameters as hyperparameters, record every combination you test, and discount results for the search, since tuning them to maximise out-of-fold score is itself a form of overfitting.
What are sample uniqueness weights and why do they matter?
Uniqueness weights measure how much independent information each overlapping label contains. For each bar you count how many labels are active, take the reciprocal, and average it over a label’s own window. Isolated labels score near 1 and heavily overlapped labels score low. Using the result as a sample weight, or to drive sequential bootstrapping, stops clustered, redundant observations from dominating training and inflating confidence.
Does meta-labeling guarantee a profitable strategy?
No. Meta-labeling can only filter or scale signals produced by the primary model, and its results depend on honest validation. On random data it should return roughly chance-level scores, as the synthetic example here does. Real strategies also face transaction costs, slippage, regime change, and selection bias from many trials. Treat the method as a way to measure and manage a signal’s reliability, not as evidence of profit.
Further Reading
- Event-driven backtesting engine architecture for algorithmic trading: the simulator layer that consumes the labels, signals, and bet sizes built here.
- Marcos Lopez de Prado, Advances in Financial Machine Learning (Wiley, 2018): Chapter 3 (Labeling, including triple-barrier labeling and meta-labeling), Chapter 4 (Sample Weights), Chapter 7 (Cross-Validation in Finance), Chapters 11 and 12 (backtesting). Table of contents on O’Reilly.
- D. H. Bailey and M. Lopez de Prado, “The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality” (SSRN).
- D. H. Bailey, J. Borwein, M. Lopez de Prado and Q. J. Zhu, “The Probability of Backtest Overfitting” (SSRN; Journal of Computational Finance).
- finmlkit triple-barrier API documentation: one open-source implementation, useful for comparing parameter conventions.
By Riju — about
