# Machine learning labels (/features/indicators/machine-learning-labels)

How do you create targets for a trading model, and how do you backtest what it predicts? VBT
generates machine learning labels in Python, including fixed-horizon returns, triple-barrier
outcomes, trends, and pivots, and runs model predictions through the same backtester as any other
signal, with the look-ahead in the labels kept explicit.

```python title="Label barrier outcomes, train a classifier walk-forward, and backtest its predictions"
>>> from sklearn.ensemble import RandomForestClassifier

>>> data = vbt.YFData.pull("BTC-USD", start="2018-01-01", end="2025-01-01")
>>> labels = vbt.BOLB.run(  # (1)
...     data.high, data.low, window=10, up_th=0.05, down_th=0.05
... ).labels.rename("label")
>>> labels.value_counts(dropna=False).sort_index()
label
-1.0    1089
 0.0     151
 1.0    1317
Name: count, dtype: int64

>>> returns = data.close.pct_change()
>>> features = pd.DataFrame({  # (2)
...     "rsi": vbt.RSI.run(data.close, window=14).rsi,
...     "ret_5": data.close.pct_change(5),
...     "ret_20": data.close.pct_change(20),
...     "vol_20": returns.rolling(20).std(),
...     "dist_sma_50": data.close / data.close.rolling(50).mean() - 1,
... })

>>> proba = pd.Series(np.nan, index=data.index)
>>> for year in [2021, 2022, 2023, 2024]:  # (3)
...     test = data.index.year == year
...     train_end = data.index[test][0] - pd.Timedelta(days=10)
...     train = (
...         (data.index < train_end)
...         & features.notna().all(axis=1)
...         & labels.notna()
...         & (labels != 0)
...     )
...     model = RandomForestClassifier(n_estimators=200, max_depth=4, random_state=42)
...     model.fit(features[train], labels[train] == 1)
...     proba[test] = model.predict_proba(features[test].fillna(0))[:, 1]

>>> kwargs = dict(
...     price="nextopen",  # (4)
...     tp_stop=0.05,
...     sl_stop=0.05,
...     td_stop=10,
...     time_delta_format="rows",
...     fees=0.001,
... )
>>> test_data = data.loc["2021":]
>>> model_pf = vbt.PF.from_signals(test_data, (proba > 0.55).loc["2021":], **kwargs)
>>> always_pf = vbt.PF.from_signals(test_data, True, **kwargs)  # (5)
>>> metrics = ["total_trades", "win_rate", "total_return", "sharpe_ratio"]
>>> pd.DataFrame({
...     "model": model_pf.stats(metrics),
...     "every bar": always_pf.stats(metrics),
... }).round(2)
                      model  every bar
Total Trades            145        302
Win Rate [%]          51.03      47.84
Total Return [%]       9.14     -42.96
Sharpe Ratio           0.24       0.02
```

1.  Scan the next 10 bars for a 5% rise from the current low (1) or a 5% fall from the current high
    (-1). Zero means no unambiguous breakout was found. These price and time barriers form the
    labeling target.
2.  Features use only data up to each bar's close.
3.  Train on all earlier years, leaving out the last 10 days before each test year so no training
    label looks into it.
4.  A prediction made at a bar's close is traded at the next open. Stops use the same 5% distances
    and 10-bar limit, measured from the simulated trade rather than the label's high and low.
5.  The baseline takes every trade the barriers allow, without the model.

The model's entries won 51% of the time against 48% for entering on every bar, and turned a loss
into a small gain over four out-of-sample years. A modest edge, measured honestly: labels from the
future, features from the past, and predictions only on years the model never saw.

The classifier is trained on the upward and downward outcomes, leaving out zero labels. Its
probability threshold selects entries, and the portfolio simulation measures what those entries
actually earn after fees. You can change the target, model, or threshold and keep this same research
workflow.

## Choose what the model should predict \[#choose-what-the-model-should-predict]

A label defines the question you want the model to answer. Use a future return when the size of the
move matters, a barrier outcome when the order of price moves matters, or a trend label when you
want to study broader market moves. These targets can serve regression or classification models in
your preferred Python framework.

## Label generators \[#label-generators]

Every label generator is an indicator that looks forward instead of back:

| Indicator     | Label                                                                         | Useful for                                              |
| ------------- | ----------------------------------------------------------------------------- | ------------------------------------------------------- |
| `vbt.FIXLB`   | Return from each bar to the bar a fixed number of bars ahead                  | Comparing prediction horizons                           |
| `vbt.MEANLB`  | Return to the average of a future window                                      | Predicting an average future price level                |
| `vbt.BOLB`    | Which threshold, up or down, is reached first within a future window          | Classifying future breakout direction                   |
| `vbt.PIVOTLB` | Whether each bar is a high or a low pivot, judged with hindsight              | Learning turning-point targets                          |
| `vbt.TRENDLB` | Trend between pivots, as a direction, a position within the move, or a return | Classifying trends or predicting progress within a move |

Like other indicators, they take parameter lists, so a grid of horizons or thresholds produces a
grid of labels, and they plot over price for a quick check. `vbt.PIVOTLB` and `vbt.TRENDLB` read
highs and lows and take two positive thresholds: `up_th`, the rise from a trough that confirms it,
and `down_th`, the fall from a peak. Equal percentages are not symmetric, since undoing a 50% fall
takes a 100% rise.

For breakout targets, `BOLB` offers two reference choices. The default `"swing"` mode measures
upward moves from the current low and downward moves from the current high. `"range"` mode puts the
thresholds beyond the current high and low, which suits questions about breaking out of the bar's
range. Thresholds can also vary across observations, for example using a volatility-based series you
calculate from past data.

## Compare prediction horizons \[#compare-prediction-horizons]

The horizon is part of the research question. This fictional series produces different return
targets one and three bars ahead, with percentages shown below:

```python title="Build one-bar and three-bar return targets together"
>>> close = pd.Series([100.0, 110.0, 105.0, 120.0, 108.0],
...                   index=pd.date_range("2025-01-01", periods=5))
>>> labels = vbt.FIXLB.run(close, n=[1, 3]).labels
>>> (labels * 100).round(2)
fixlb_n         1      3
2025-01-01  10.00  20.00
2025-01-02  -4.55  -1.82
2025-01-03  14.29    NaN
2025-01-04 -10.00    NaN
2025-01-05    NaN    NaN
```

`FIXLB` leaves a target missing when the required future price is unavailable. Select the horizon
you want to model and train on rows where both its target and your features are available.

!!! warning "Labels use the future"
    Labels use future observations. Use them as training targets or for retrospective analysis, not as
    information available to a trading decision. Exclude training labels whose future observations
    overlap the test period. A gap covers the fixed horizon in the example above, while pivot and
    trend labels need the time at which their outcomes became known. See
    [Walk-forward and cross-validation](/features/optimization/time-series-cross-validation/).

## Forward statistics \[#forward-statistics]

`vbt.FMEAN`, `vbt.FMAX`, `vbt.FMIN`, and `vbt.FSTD` compute the mean, maximum, minimum, and standard
deviation of a future window, with a `wait` before the window starts. They answer questions such as
"how far did the price rise in the next 20 days" and serve as continuous targets.

## Meta-labeling \[#meta-labeling]

A model does not have to find trades on its own. In meta-labeling, a rule-based strategy proposes
trades, each trade is labeled by its outcome, for example profitable after fees or not, and a second
model learns which proposals to take or how much to put into them. The labels come from the
strategy's own trade records, such as `pf.trades.returns`, and the model's decision becomes a filter
on the entries or a size for each one. Keep the features aligned to the original entry opportunity,
then compare the filtered strategy with the original under the same fees and sizing rules.
[Trade analytics](/features/analytics/trade-analytics/) helps inspect which trades the filter
accepts and rejects.

## Backtesting predictions \[#backtesting-predictions]

Predictions are arrays like any other input. Thresholded probabilities become entry and exit
signals, predicted returns or scores become sizes, and predicted weights for several assets become
target allocations for `from_orders`. Models trained on a GPU or in another framework only need
their output converted to NumPy or pandas before it reaches the simulator, which runs on the CPU.

!!! info "Tutorial"
    The members-only [Cross-validation](https://members.vectorbt.pro/tutorials/cross-validation/applications/)
    tutorial applies splitters to models and parameter searches without leaking the future into
    training.


## Related pages

*   [Walk-forward and cross-validation](/features/optimization/time-series-cross-validation/): Test optimized parameters out of sample with rolling, expanding, and purged splits
*   [Trading signals](/features/indicators/trading-signals/): Generate, combine, clean, and analyze entry and exit signals before backtesting
*   [Realistic backtests](/features/backtesting/backtest-realism/): Avoid look-ahead bias, impossible fills, and survivorship bias in your backtests
*   [Stop loss and take profit](/features/backtesting/stop-loss-and-take-profit/): Backtest stop losses, trailing stops, take profits, time stops, and exit ladders