Tutorials

Cross-validation

Design robust cross-validation workflows for trading strategies

After developing a rule-based or machine learning-based strategy, it is time to backtest it. If our first backtest yields a low Sharpe ratio, we might tweak the strategy to try to improve it. After several rounds of parameter adjustments, we may arrive at a "flawless" set of parameters and a strategy showing an exceptional Sharpe ratio. However, in live trading, the strategy may perform poorly, resulting in losses. What went wrong?

Markets naturally contain noise—small and frequent inconsistencies in price data. When designing a strategy, we should avoid optimizing for a single period, because the model may fit the historical data so closely that it fails to predict the future effectively. This is similar to tuning a car for one specific racetrack and expecting it to perform equally well everywhere. Especially with VBT, which allows us to search large databases of historical market data for patterns, it can be easy to create complex rules that appear highly accurate at predicting price changes (see p-hacking) but make random guesses when applied to data outside the sample used to build the model.

Overfitting (also known as curve fitting) typically occurs for one or more of these reasons: mistaking noise for signal, or excessively tweaking too many parameters. To avoid overfitting, we should use cross-validation (CV), which involves splitting a data sample into complementary subsets, analyzing one subset called the training or in-sample (IS) set, and validating the analysis on the other subset called the validation or out-of-sample (OOS) set. This procedure is repeated until we have multiple OOS periods and can calculate statistics from all these results combined. The key questions we need to ask are: are our parameter choices robust in the IS periods? Is our performance robust in the OOS periods? If not, we are essentially guessing, and as quant investors we should not leave room for second-guessing when real money is at stake.

Let's consider a simple strategy based on a moving average crossover.

First, we will pull some data:

from vectorbtpro import *

data = vbt.BinanceData.pull("BTCUSDT", end="2022-11-01 UTC")
data.index
DatetimeIndex(['2017-08-17 00:00:00+00:00', '2017-08-18 00:00:00+00:00',
               '2017-08-19 00:00:00+00:00', '2017-08-20 00:00:00+00:00',
               ...
               '2022-10-28 00:00:00+00:00', '2022-10-29 00:00:00+00:00',
               '2022-10-30 00:00:00+00:00', '2022-10-31 00:00:00+00:00'],
    dtype='datetime64[ns, UTC]', name='Open time', length=1902, freq='D')

Next, let's create a parameterized mini-pipeline that takes data and parameters and returns the Sharpe ratio, which should reflect our strategy's performance on that test period:

@vbt.parameterized(merge_func="concat")  
def sma_crossover_perf(data, fast_window, slow_window):
    fast_sma = data.run("sma", fast_window, short_name="fast_sma")  
    slow_sma = data.run("sma", slow_window, short_name="slow_sma")
    entries = fast_sma.real_crossed_above(slow_sma)
    exits = fast_sma.real_crossed_below(slow_sma)
    pf = vbt.Portfolio.from_signals(
        data, entries, exits, direction="both")  
    return pf.sharpe_ratio  

Let's test a grid of fast_window and slow_window combinations over one year of that data:

perf = sma_crossover_perf(  
    data["2020":"2020"],  
    vbt.Param(np.arange(5, 50), condition="x < slow_window"),  
    vbt.Param(np.arange(5, 50)),  
    _execute_kwargs=dict(  
        clear_cache=50,  
        collect_garbage=50
    )
)
perf
fast_window  slow_window
5            6              0.625318
             7              0.333243
             8              1.171861
             9              1.062940
             10             0.635302
                                 ...
46           48             0.534582
             49             0.573196
47           48             0.445239
             49             0.357548
48           49            -0.826995
Length: 990, dtype: float64
Combination 990/990

When trading strategies that have shown strong performance in backtesting are applied to new data, they often prove to be fragile or may even fail catastrophically. This usually happens due to backtest overfitting, which can be detected and avoided with proper cross-validation.

✅ We will start with manual cross-validation, then cover different splitting schemes and techniques for analyzing split distributions. Finally, we will go over how to automate various cross-validation steps, from splitting regular arrays and complex VBT objects to testing entire parameter-based and ML-based pipelines 🔥

Topics

Copyright © 2021–2026 Oleg Polakow. All rights reserved.

Site content and documentation are provided for using and evaluating VectorBT PRO and for educational purposes. Any other use, including building or supporting competing products or services, requires prior written consent.