Features
Robustness and overfitting
Check whether a backtest is luck with baselines, permutation tests, and deflated Sharpe
Is the edge in a backtest real, or did the search just find the luckiest parameters? VBT tests backtest overfitting in Python with baselines, parameter sensitivity, permutation tests on shuffled data, and the probabilistic and deflated Sharpe ratios.
from itertools import product
data = vbt.YFData.pull("BTC-USD", start="2020-01-01", end="2025-01-01")
pairs = list(product(range(10, 55, 5), range(60, 220, 20)))
def best_sharpe(close):
sharpe = []
for fast_window, slow_window in pairs:
fast = close.rolling(fast_window).mean()
slow = close.rolling(slow_window).mean()
pf = vbt.PF.from_signals(
close, fast > slow, fast < slow, fees=0.001, freq="1D"
)
sharpe.append(pf.sharpe_ratio)
return pd.concat(sharpe, axis=1).max(axis=1)
real_best = best_sharpe(data.close.to_frame()).iloc[0]
returns = data.close.pct_change().iloc[1:]
shuffled = pd.DataFrame(np.tile(returns.values[:, None], 1000), index=returns.index)
shuffled = shuffled.vbt.shuffle(seed=42)
perm_best = best_sharpe(data.close.iloc[0] * (1 + shuffled).cumprod())
p_value = ((perm_best >= real_best).sum() + 1) / (len(perm_best) + 1)
print(round(real_best, 3), round(perm_best.median(), 3), round(p_value, 3))1.406 1.162 0.154The best crossover on real Bitcoin prices reached a Sharpe ratio of 1.41. Searching the same 72 pairs on prices reconstructed from shuffled returns found a better result 15% of the time. A p-value of 0.15 means the search cannot show that the crossover timing adds anything beyond Bitcoin's rise.
fast = vbt.MA.run(data.close, window=[f for f, _ in pairs], short_name="fast")
slow = vbt.MA.run(data.close, window=[s for _, s in pairs], short_name="slow")
pf = vbt.PF.from_signals(data, fast.ma_above(slow), fast.ma_below(slow), fees=0.001)
best = pf.sharpe_ratio.idxmax()
print(round(pf.deflated_sharpe_ratio[best], 3), round(pf.prob_sharpe_ratio[best], 3))0.995 0.737The deflated Sharpe ratio gives a 99.5% probability that the best pair's true Sharpe ratio is above what the luckiest of 72 skill-free trials would show, but it says nothing about holding. The probabilistic Sharpe ratio compares the same pair with buying and holding Bitcoin and puts the chance of beating it at 74%. Each test asks a different question, which is why one is rarely enough.
Choose a robustness test
Different tests help explain different weaknesses. You can use the same strategy function and compare the results in labeled pandas tables.
| What you want to check | What to compare |
|---|---|
| Whether entry timing adds value | Buy and hold and random entries |
| Whether a precise setting drives the result | Nearby parameters |
| Whether an extra rule helps | The strategy with each filter on and off |
| Whether execution assumptions drive the result | Signal delays and higher costs |
| Whether the search could find a similar winner by chance | Repeated searches on shuffled data and adjusted Sharpe measures |
For each comparison, keep the dates, sizing, and execution assumptions consistent unless those are what you are testing. Alongside returns, inspect drawdowns and trade counts and exposure to see what changed.
Baselines
Every result needs something to beat. vbt.PF.from_holding buys once and holds, and
vbt.PF.from_random_signals enters and exits at random with a seed, using either a fixed number of
signals or a probability per bar.
hold = vbt.PF.from_holding(data, fees=0.001)
n_trades = pf.trades.count()[best]
rand = vbt.PF.from_random_signals(data, n=[n_trades] * 1000, seed=42, fees=0.001)
print(n_trades, round(hold.sharpe_ratio, 3), round(rand.sharpe_ratio.median(), 3))8 1.126 1.015The best pair beats holding and every one of the 1,000 random portfolios with the same 8 trades. That is still one strategy chosen after the fact from 72, so it does not settle the question. The permutation test above repeats the whole search on each shuffled history, which accounts for the search itself.
Parameter sensitivity
A robust strategy has a plateau: neighboring parameters perform about as well as the best one. A single peak surrounded by weak results usually means one lucky trade or one lucky period. Laying the Sharpe ratios out by parameter shows which kind of result you have.
pf.sharpe_ratio.unstack("slow_window").iloc[:5, :6].round(2)slow_window 60 80 100 120 140 160
fast_window
10 1.23 1.17 1.37 1.32 1.17 1.29
15 1.25 1.12 1.37 1.16 1.17 1.17
20 1.30 1.14 1.22 1.15 1.22 1.25
25 1.21 1.24 1.41 1.20 1.32 1.27
30 1.02 1.19 1.32 1.20 1.20 1.27The eight neighbors of 25 and 100 score between 1.14 and 1.32, so this result is not an isolated spike. The same layout works for ablation tests: pass each filter of a strategy as an on and off parameter and compare the strategy with and without it. The Parameter optimization page covers heatmaps and large parameter grids.
Test which filters help
An extra filter can improve the headline return simply by leaving fewer trades. An ablation test shows what each rule contributes. Build one signal column per variant and backtest them together, using the same prices, exits, and fees.
This small example uses made-up prices and volumes. It compares a core entry rule with a trend filter, a volume filter, and both filters together. All signals move forward one bar before trading.
close = pd.Series([100, 102, 105, 103, 99, 101, 106, 110, 107, 102, 104, 109, 114, 111, 106, 108, 115, 118, 113, 110])
volume = pd.Series([90, 110, 80, 120, 90, 130, 140, 80, 110, 90, 120, 80, 140, 100, 90, 120, 80, 130, 110, 90])
core = close > close.shift(1)
trend_filter = close > close.rolling(4).mean()
volume_filter = volume >= 100
entries = pd.DataFrame({
"Core": core,
"Trend": core & trend_filter,
"Volume": core & volume_filter,
"Both": core & trend_filter & volume_filter,
}).vbt.signals.fshift(1)
exits = (close < close.shift(1)).vbt.signals.fshift(1)
variants = vbt.PF.from_signals(close, entries, exits, fees=0.001, freq="1D")
comparison = pd.DataFrame({
"Return [%]": variants.total_return * 100,
"Max drawdown [%]": variants.max_drawdown * 100,
"Closed trades": variants.trades.status_closed.count(),
})
comparison.round(2) Return [%] Max drawdown [%] Closed trades
Core -16.28 -16.28 4
Trend -20.11 -20.11 3
Volume -16.28 -16.28 4
Both -14.32 -14.32 3See exactly what each rule contributes: the volume filter alone leaves the result unchanged, while combining both filters produces the smallest loss on this sample. Use the same comparison to test entry rules across a larger strategy, with filter switches as parameters and repeated walk-forward test periods.
Execution sensitivity
Stress-test execution by delaying signals and raising costs. Compare how the strategy responds to a one-bar delay, a two-bar delay, and different fee levels. Each delay is one column and each fee level one parameter value, so the grid runs as a single backtest.
entries = fast.ma_above(slow)[best]
exits = fast.ma_below(slow)[best]
delays = pd.Index([0, 1, 2], name="delay")
pf_exec = vbt.PF.from_signals(
data,
pd.concat([entries.vbt.signals.fshift(d) for d in delays], axis=1, keys=delays),
pd.concat([exits.vbt.signals.fshift(d) for d in delays], axis=1, keys=delays),
fees=vbt.Param([0.001, 0.002, 0.005], name="fees"),
)
pf_exec.sharpe_ratio.unstack("fees").round(2)fees 0.001 0.002 0.005
delay
0 1.41 1.40 1.38
1 1.33 1.33 1.31
2 1.32 1.31 1.29With only 8 trades, even 0.5% fees barely move the result, while a one-bar delay costs more than all of them. A strategy that trades more often pays more for each step in fees. The Backtest realism page covers slippage, order delays, and other execution assumptions.
Perturbed and permuted data
A strategy that only works on the exact price history it was tuned on is fragile. Each of these varies the data in a different way:
- Shuffled returns.
.vbt.shuffle(seed=...)shuffles each column separately, as in the first example, to break the original ordering while keeping the return distribution. - Noise. Adding small random noise to prices, for example with
np.random.default_rng(seed)scaled to each bar's range, tests whether signals depend on exact price levels. - Mirrored prices.
data.mirror_ohlc()turns rises into falls and keeps each bar's high and low consistent, which tests whether a long-only result depends on the market's direction. - Block bootstrap.
vbt.Splitter.from_n_randomdraws random windows, andshuffle_splitswithreplace=Trueresamples existing windows, keeping the order of prices inside each block. - Synthetic prices.
vbt.GBMOHLCDataandvbt.RandomOHLCDatagenerate seeded price paths with chosen drift and volatility, as the Synthetic data page shows.
Each variant is one more column, so hundreds of perturbed histories run as one backtest.
Statistical tests
The returns accessor and the portfolio both compute the Sharpe ratio tests by López de Prado and Bailey:
| Method | Answers |
|---|---|
sharpe_ratio_std | How uncertain the Sharpe ratio is, given the sample's length, skew, and kurtosis |
prob_sharpe_ratio | How likely the true Sharpe ratio is above a benchmark's, such as buy and hold |
deflated_sharpe_ratio | How likely the Sharpe ratio exceeds a benchmark adjusted for the search across columns |
prob_sharpe_ratio on a portfolio compares with its benchmark, which is buying and holding the
traded asset unless you pass bm_returns. To compare with a Sharpe ratio of zero instead, call it
on pf.get_returns_acc(bm_returns=False).
deflated_sharpe_ratio treats each column as one trial, so call it on the returns of every
combination you tested, not only on the winner. A single column returns NaN, and a grouped portfolio
counts each group as one trial.
Probability of backtest overfitting
VBT provides splitters and labeled results for calculating the probability of backtest overfitting
(PBO) in a research workflow. Compare the in-sample and out-of-sample results of every combination
over combinatorial splits, using half the folds for training and half for testing. Build those
splits with vbt.Splitter.from_purged_kfold, covered on the
Walk-forward and cross-validation page. PBO
is the share of splits in which the best combination on the training half ranks below the median on
the test half. Values near 0.5 or higher mean the selection does no better than picking a
combination at random.
Correlated trials
The deflated Sharpe ratio assumes the trials are independent. Neighboring parameters of one strategy are strongly correlated, so a grid of 72 pairs behaves like far fewer independent trials. The permutation example complements this by rerunning the same correlated parameter grid on each shuffled history.
Tutorial
The members-only Cross-validation tutorial carries each period's winning parameters into the next period and compares them with every alternative out of sample.
Related pages
- Walk-forward and cross-validationTest optimized parameters out of sample with rolling, expanding, and purged splits
- Analysis › Performance and risk metricsCompute returns, Sharpe, Sortino, benchmark, and rolling metrics on any backtest
- Parameter optimizationSweep millions of parameter combinations with grids, conditions, and random search
- Backtesting › Backtesting engineSimulate orders, signals, and callbacks across many assets and parameters at once
Copyright © 2021–2026 Oleg Polakow. All rights reserved.
Site content and documentation are provided for using and evaluating VectorBT PRO and for educational purposes. Any other use, including building or supporting competing products or services, requires prior written consent.