Features
Machine learning labels
Create triple-barrier, trend, and pivot labels, and backtest model predictions
How do you create targets for a trading model, and how do you backtest what it predicts? VBT generates machine learning labels in Python, including fixed-horizon returns, triple-barrier outcomes, trends, and pivots, and runs model predictions through the same backtester as any other signal, with the look-ahead in the labels kept explicit.
from sklearn.ensemble import RandomForestClassifier
data = vbt.YFData.pull("BTC-USD", start="2018-01-01", end="2025-01-01")
labels = vbt.BOLB.run(
data.high, data.low, window=10, up_th=0.05, down_th=0.05
).labels.rename("label")
labels.value_counts(dropna=False).sort_index()label
-1.0 1089
0.0 151
1.0 1317
Name: count, dtype: int64returns = data.close.pct_change()
features = pd.DataFrame({
"rsi": vbt.RSI.run(data.close, window=14).rsi,
"ret_5": data.close.pct_change(5),
"ret_20": data.close.pct_change(20),
"vol_20": returns.rolling(20).std(),
"dist_sma_50": data.close / data.close.rolling(50).mean() - 1,
})
proba = pd.Series(np.nan, index=data.index)
for year in [2021, 2022, 2023, 2024]:
test = data.index.year == year
train_end = data.index[test][0] - pd.Timedelta(days=10)
train = (
(data.index < train_end)
& features.notna().all(axis=1)
& labels.notna()
& (labels != 0)
)
model = RandomForestClassifier(n_estimators=200, max_depth=4, random_state=42)
model.fit(features[train], labels[train] == 1)
proba[test] = model.predict_proba(features[test].fillna(0))[:, 1]
kwargs = dict(
price="nextopen",
tp_stop=0.05,
sl_stop=0.05,
td_stop=10,
time_delta_format="rows",
fees=0.001,
)
test_data = data.loc["2021":]
model_pf = vbt.PF.from_signals(test_data, (proba > 0.55).loc["2021":], **kwargs)
always_pf = vbt.PF.from_signals(test_data, True, **kwargs)
metrics = ["total_trades", "win_rate", "total_return", "sharpe_ratio"]
pd.DataFrame({
"model": model_pf.stats(metrics),
"every bar": always_pf.stats(metrics),
}).round(2) model every bar
Total Trades 145 302
Win Rate [%] 51.03 47.84
Total Return [%] 9.14 -42.96
Sharpe Ratio 0.24 0.02The model's entries won 51% of the time against 48% for entering on every bar, and turned a loss into a small gain over four out-of-sample years. A modest edge, measured honestly: labels from the future, features from the past, and predictions only on years the model never saw.
The classifier is trained on the upward and downward outcomes, leaving out zero labels. Its probability threshold selects entries, and the portfolio simulation measures what those entries actually earn after fees. You can change the target, model, or threshold and keep this same research workflow.
Choose what the model should predict
A label defines the question you want the model to answer. Use a future return when the size of the move matters, a barrier outcome when the order of price moves matters, or a trend label when you want to study broader market moves. These targets can serve regression or classification models in your preferred Python framework.
Label generators
Every label generator is an indicator that looks forward instead of back:
| Indicator | Label | Useful for |
|---|---|---|
vbt.FIXLB | Return from each bar to the bar a fixed number of bars ahead | Comparing prediction horizons |
vbt.MEANLB | Return to the average of a future window | Predicting an average future price level |
vbt.BOLB | Which threshold, up or down, is reached first within a future window | Classifying future breakout direction |
vbt.PIVOTLB | Whether each bar is a high or a low pivot, judged with hindsight | Learning turning-point targets |
vbt.TRENDLB | Trend between pivots, as a direction, a position within the move, or a return | Classifying trends or predicting progress within a move |
Like other indicators, they take parameter lists, so a grid of horizons or thresholds produces a
grid of labels, and they plot over price for a quick check. vbt.PIVOTLB and vbt.TRENDLB read
highs and lows and take two positive thresholds: up_th, the rise from a trough that confirms it,
and down_th, the fall from a peak. Equal percentages are not symmetric, since undoing a 50% fall
takes a 100% rise.
For breakout targets, BOLB offers two reference choices. The default "swing" mode measures
upward moves from the current low and downward moves from the current high. "range" mode puts the
thresholds beyond the current high and low, which suits questions about breaking out of the bar's
range. Thresholds can also vary across observations, for example using a volatility-based series you
calculate from past data.
Compare prediction horizons
The horizon is part of the research question. This fictional series produces different return targets one and three bars ahead, with percentages shown below:
close = pd.Series([100.0, 110.0, 105.0, 120.0, 108.0],
index=pd.date_range("2025-01-01", periods=5))
labels = vbt.FIXLB.run(close, n=[1, 3]).labels
(labels * 100).round(2)fixlb_n 1 3
2025-01-01 10.00 20.00
2025-01-02 -4.55 -1.82
2025-01-03 14.29 NaN
2025-01-04 -10.00 NaN
2025-01-05 NaN NaNFIXLB leaves a target missing when the required future price is unavailable. Select the horizon
you want to model and train on rows where both its target and your features are available.
Labels use the future
Labels use future observations. Use them as training targets or for retrospective analysis, not as information available to a trading decision. Exclude training labels whose future observations overlap the test period. A gap covers the fixed horizon in the example above, while pivot and trend labels need the time at which their outcomes became known. See Walk-forward and cross-validation.
Forward statistics
vbt.FMEAN, vbt.FMAX, vbt.FMIN, and vbt.FSTD compute the mean, maximum, minimum, and standard
deviation of a future window, with a wait before the window starts. They answer questions such as
"how far did the price rise in the next 20 days" and serve as continuous targets.
Meta-labeling
A model does not have to find trades on its own. In meta-labeling, a rule-based strategy proposes
trades, each trade is labeled by its outcome, for example profitable after fees or not, and a second
model learns which proposals to take or how much to put into them. The labels come from the
strategy's own trade records, such as pf.trades.returns, and the model's decision becomes a filter
on the entries or a size for each one. Keep the features aligned to the original entry opportunity,
then compare the filtered strategy with the original under the same fees and sizing rules.
Trade analytics helps inspect which trades the filter
accepts and rejects.
Backtesting predictions
Predictions are arrays like any other input. Thresholded probabilities become entry and exit
signals, predicted returns or scores become sizes, and predicted weights for several assets become
target allocations for from_orders. Models trained on a GPU or in another framework only need
their output converted to NumPy or pandas before it reaches the simulator, which runs on the CPU.
Tutorial
The members-only Cross-validation tutorial applies splitters to models and parameter searches without leaking the future into training.
Related pages
- Optimization and validation › Walk-forward and cross-validationTest optimized parameters out of sample with rolling, expanding, and purged splits
- Trading signalsGenerate, combine, clean, and analyze entry and exit signals before backtesting
- Backtesting › Realistic backtestsAvoid look-ahead bias, impossible fills, and survivorship bias in your backtests
- Backtesting › Stop loss and take profitBacktest stop losses, trailing stops, take profits, time stops, and exit ladders
Copyright © 2021–2026 Oleg Polakow. All rights reserved.
Site content and documentation are provided for using and evaluating VectorBT PRO and for educational purposes. Any other use, including building or supporting competing products or services, requires prior written consent.