Features

Robustness and overfitting

Check whether a backtest is luck with baselines, permutation tests, and deflated Sharpe

Is the edge in a backtest real, or did the search just find the luckiest parameters? VBT tests backtest overfitting in Python with baselines, parameter sensitivity, permutation tests on shuffled data, and the probabilistic and deflated Sharpe ratios.

Compare the best of 72 strategies with the best on 1,000 shuffled histories
from itertools import product

data = vbt.YFData.pull("BTC-USD", start="2020-01-01", end="2025-01-01")
pairs = list(product(range(10, 55, 5), range(60, 220, 20)))

def best_sharpe(close):  
    sharpe = []
    for fast_window, slow_window in pairs:
        fast = close.rolling(fast_window).mean()
        slow = close.rolling(slow_window).mean()
        pf = vbt.PF.from_signals(
            close, fast > slow, fast < slow, fees=0.001, freq="1D"
        )
        sharpe.append(pf.sharpe_ratio)
    return pd.concat(sharpe, axis=1).max(axis=1)

real_best = best_sharpe(data.close.to_frame()).iloc[0]
returns = data.close.pct_change().iloc[1:]
shuffled = pd.DataFrame(np.tile(returns.values[:, None], 1000), index=returns.index)
shuffled = shuffled.vbt.shuffle(seed=42)  
perm_best = best_sharpe(data.close.iloc[0] * (1 + shuffled).cumprod())
p_value = ((perm_best >= real_best).sum() + 1) / (len(perm_best) + 1)
print(round(real_best, 3), round(perm_best.median(), 3), round(p_value, 3))
1.406 1.162 0.154
Histogram of the best Sharpe ratio on 1,000 shuffled Bitcoin histories, with the best Sharpe ratio on real prices marked Figure data (JSON)

The best crossover on real Bitcoin prices reached a Sharpe ratio of 1.41. Searching the same 72 pairs on prices reconstructed from shuffled returns found a better result 15% of the time. A p-value of 0.15 means the search cannot show that the crossover timing adds anything beyond Bitcoin's rise.

Deflate the best Sharpe ratio for 72 trials
fast = vbt.MA.run(data.close, window=[f for f, _ in pairs], short_name="fast")
slow = vbt.MA.run(data.close, window=[s for _, s in pairs], short_name="slow")
pf = vbt.PF.from_signals(data, fast.ma_above(slow), fast.ma_below(slow), fees=0.001)
best = pf.sharpe_ratio.idxmax()  
print(round(pf.deflated_sharpe_ratio[best], 3), round(pf.prob_sharpe_ratio[best], 3))
0.995 0.737

The deflated Sharpe ratio gives a 99.5% probability that the best pair's true Sharpe ratio is above what the luckiest of 72 skill-free trials would show, but it says nothing about holding. The probabilistic Sharpe ratio compares the same pair with buying and holding Bitcoin and puts the chance of beating it at 74%. Each test asks a different question, which is why one is rarely enough.

Choose a robustness test

Different tests help explain different weaknesses. You can use the same strategy function and compare the results in labeled pandas tables.

What you want to checkWhat to compare
Whether entry timing adds valueBuy and hold and random entries
Whether a precise setting drives the resultNearby parameters
Whether an extra rule helpsThe strategy with each filter on and off
Whether execution assumptions drive the resultSignal delays and higher costs
Whether the search could find a similar winner by chanceRepeated searches on shuffled data and adjusted Sharpe measures

For each comparison, keep the dates, sizing, and execution assumptions consistent unless those are what you are testing. Alongside returns, inspect drawdowns and trade counts and exposure to see what changed.

Baselines

Every result needs something to beat. vbt.PF.from_holding buys once and holds, and vbt.PF.from_random_signals enters and exits at random with a seed, using either a fixed number of signals or a probability per bar.

Compare with buy and hold and 1,000 random portfolios
hold = vbt.PF.from_holding(data, fees=0.001)
n_trades = pf.trades.count()[best]
rand = vbt.PF.from_random_signals(data, n=[n_trades] * 1000, seed=42, fees=0.001)
print(n_trades, round(hold.sharpe_ratio, 3), round(rand.sharpe_ratio.median(), 3))
8 1.126 1.015

The best pair beats holding and every one of the 1,000 random portfolios with the same 8 trades. That is still one strategy chosen after the fact from 72, so it does not settle the question. The permutation test above repeats the whole search on each shuffled history, which accounts for the search itself.

Parameter sensitivity

A robust strategy has a plateau: neighboring parameters perform about as well as the best one. A single peak surrounded by weak results usually means one lucky trade or one lucky period. Laying the Sharpe ratios out by parameter shows which kind of result you have.

Sharpe ratio for each pair of windows
pf.sharpe_ratio.unstack("slow_window").iloc[:5, :6].round(2)
slow_window   60    80    100   120   140   160
fast_window
10           1.23  1.17  1.37  1.32  1.17  1.29
15           1.25  1.12  1.37  1.16  1.17  1.17
20           1.30  1.14  1.22  1.15  1.22  1.25
25           1.21  1.24  1.41  1.20  1.32  1.27
30           1.02  1.19  1.32  1.20  1.20  1.27

The eight neighbors of 25 and 100 score between 1.14 and 1.32, so this result is not an isolated spike. The same layout works for ablation tests: pass each filter of a strategy as an on and off parameter and compare the strategy with and without it. The Parameter optimization page covers heatmaps and large parameter grids.

Test which filters help

An extra filter can improve the headline return simply by leaving fewer trades. An ablation test shows what each rule contributes. Build one signal column per variant and backtest them together, using the same prices, exits, and fees.

This small example uses made-up prices and volumes. It compares a core entry rule with a trend filter, a volume filter, and both filters together. All signals move forward one bar before trading.

Compare a strategy with and without entry filters
close = pd.Series([100, 102, 105, 103, 99, 101, 106, 110, 107, 102, 104, 109, 114, 111, 106, 108, 115, 118, 113, 110])
volume = pd.Series([90, 110, 80, 120, 90, 130, 140, 80, 110, 90, 120, 80, 140, 100, 90, 120, 80, 130, 110, 90])
core = close > close.shift(1)
trend_filter = close > close.rolling(4).mean()
volume_filter = volume >= 100
entries = pd.DataFrame({
    "Core": core,
    "Trend": core & trend_filter,
    "Volume": core & volume_filter,
    "Both": core & trend_filter & volume_filter,
}).vbt.signals.fshift(1)
exits = (close < close.shift(1)).vbt.signals.fshift(1)
variants = vbt.PF.from_signals(close, entries, exits, fees=0.001, freq="1D")
comparison = pd.DataFrame({
    "Return [%]": variants.total_return * 100,
    "Max drawdown [%]": variants.max_drawdown * 100,
    "Closed trades": variants.trades.status_closed.count(),
})
comparison.round(2)
        Return [%]  Max drawdown [%]  Closed trades
Core        -16.28            -16.28              4
Trend       -20.11            -20.11              3
Volume      -16.28            -16.28              4
Both        -14.32            -14.32              3

See exactly what each rule contributes: the volume filter alone leaves the result unchanged, while combining both filters produces the smallest loss on this sample. Use the same comparison to test entry rules across a larger strategy, with filter switches as parameters and repeated walk-forward test periods.

Execution sensitivity

Stress-test execution by delaying signals and raising costs. Compare how the strategy responds to a one-bar delay, a two-bar delay, and different fee levels. Each delay is one column and each fee level one parameter value, so the grid runs as a single backtest.

Delay the best pair's signals and raise its fees
entries = fast.ma_above(slow)[best]
exits = fast.ma_below(slow)[best]
delays = pd.Index([0, 1, 2], name="delay")
pf_exec = vbt.PF.from_signals(
    data,
    pd.concat([entries.vbt.signals.fshift(d) for d in delays], axis=1, keys=delays),
    pd.concat([exits.vbt.signals.fshift(d) for d in delays], axis=1, keys=delays),
    fees=vbt.Param([0.001, 0.002, 0.005], name="fees"),
)
pf_exec.sharpe_ratio.unstack("fees").round(2)
fees   0.001  0.002  0.005
delay
0       1.41   1.40   1.38
1       1.33   1.33   1.31
2       1.32   1.31   1.29

With only 8 trades, even 0.5% fees barely move the result, while a one-bar delay costs more than all of them. A strategy that trades more often pays more for each step in fees. The Backtest realism page covers slippage, order delays, and other execution assumptions.

Perturbed and permuted data

A strategy that only works on the exact price history it was tuned on is fragile. Each of these varies the data in a different way:

  • Shuffled returns. .vbt.shuffle(seed=...) shuffles each column separately, as in the first example, to break the original ordering while keeping the return distribution.
  • Noise. Adding small random noise to prices, for example with np.random.default_rng(seed) scaled to each bar's range, tests whether signals depend on exact price levels.
  • Mirrored prices. data.mirror_ohlc() turns rises into falls and keeps each bar's high and low consistent, which tests whether a long-only result depends on the market's direction.
  • Block bootstrap. vbt.Splitter.from_n_random draws random windows, and shuffle_splits with replace=True resamples existing windows, keeping the order of prices inside each block.
  • Synthetic prices. vbt.GBMOHLCData and vbt.RandomOHLCData generate seeded price paths with chosen drift and volatility, as the Synthetic data page shows.

Each variant is one more column, so hundreds of perturbed histories run as one backtest.

Statistical tests

The returns accessor and the portfolio both compute the Sharpe ratio tests by López de Prado and Bailey:

MethodAnswers
sharpe_ratio_stdHow uncertain the Sharpe ratio is, given the sample's length, skew, and kurtosis
prob_sharpe_ratioHow likely the true Sharpe ratio is above a benchmark's, such as buy and hold
deflated_sharpe_ratioHow likely the Sharpe ratio exceeds a benchmark adjusted for the search across columns

prob_sharpe_ratio on a portfolio compares with its benchmark, which is buying and holding the traded asset unless you pass bm_returns. To compare with a Sharpe ratio of zero instead, call it on pf.get_returns_acc(bm_returns=False).

deflated_sharpe_ratio treats each column as one trial, so call it on the returns of every combination you tested, not only on the winner. A single column returns NaN, and a grouped portfolio counts each group as one trial.

Probability of backtest overfitting

VBT provides splitters and labeled results for calculating the probability of backtest overfitting (PBO) in a research workflow. Compare the in-sample and out-of-sample results of every combination over combinatorial splits, using half the folds for training and half for testing. Build those splits with vbt.Splitter.from_purged_kfold, covered on the Walk-forward and cross-validation page. PBO is the share of splits in which the best combination on the training half ranks below the median on the test half. Values near 0.5 or higher mean the selection does no better than picking a combination at random.

Correlated trials

The deflated Sharpe ratio assumes the trials are independent. Neighboring parameters of one strategy are strongly correlated, so a grid of 72 pairs behaves like far fewer independent trials. The permutation example complements this by rerunning the same correlated parameter grid on each shuffled history.

Tutorial

The members-only Cross-validation tutorial carries each period's winning parameters into the next period and compares them with every alternative out of sample.

Copyright © 2021–2026 Oleg Polakow. All rights reserved.

Site content and documentation are provided for using and evaluating VectorBT PRO and for educational purposes. Any other use, including building or supporting competing products or services, requires prior written consent.