# Robustness and overfitting (/features/optimization/robustness-and-overfitting)

Is the edge in a backtest real, or did the search just find the luckiest parameters? VBT tests
backtest overfitting in Python with baselines, parameter sensitivity, permutation tests on shuffled
data, and the probabilistic and deflated Sharpe ratios.

```python title="Compare the best of 72 strategies with the best on 1,000 shuffled histories"
>>> from itertools import product

>>> data = vbt.YFData.pull("BTC-USD", start="2020-01-01", end="2025-01-01")
>>> pairs = list(product(range(10, 55, 5), range(60, 220, 20)))

>>> def best_sharpe(close):  # (1)
...     sharpe = []
...     for fast_window, slow_window in pairs:
...         fast = close.rolling(fast_window).mean()
...         slow = close.rolling(slow_window).mean()
...         pf = vbt.PF.from_signals(
...             close, fast > slow, fast < slow, fees=0.001, freq="1D"
...         )
...         sharpe.append(pf.sharpe_ratio)
...     return pd.concat(sharpe, axis=1).max(axis=1)

>>> real_best = best_sharpe(data.close.to_frame()).iloc[0]
>>> returns = data.close.pct_change().iloc[1:]
>>> shuffled = pd.DataFrame(np.tile(returns.values[:, None], 1000), index=returns.index)
>>> shuffled = shuffled.vbt.shuffle(seed=42)  # (2)
>>> perm_best = best_sharpe(data.close.iloc[0] * (1 + shuffled).cumprod())
>>> p_value = ((perm_best >= real_best).sum() + 1) / (len(perm_best) + 1)
>>> print(round(real_best, 3), round(perm_best.median(), 3), round(p_value, 3))
1.406 1.162 0.154
```

1.  Runs all 72 moving average pairs on every column and keeps the best Sharpe ratio per column. With
    1,000 columns, that is 72,000 backtests, which took 12 seconds on an Apple M3.
2.  Shuffles the daily returns separately in each of the 1,000 columns. Every shuffled history has
    the same returns and the same final price as Bitcoin, in a random order.

Histogram of the best Sharpe ratio on 1,000 shuffled Bitcoin histories, with the best Sharpe ratio on real prices marked. [Figure data (JSON)](/assets/figures/features/optimization/permutation-test.56a0e3418d25.json)

The best crossover on real Bitcoin prices reached a Sharpe ratio of 1.41. Searching the same 72
pairs on prices reconstructed from shuffled returns found a better result 15% of the time. A p-value
of 0.15 means the search cannot show that the crossover timing adds anything beyond Bitcoin's rise.

```python title="Deflate the best Sharpe ratio for 72 trials"
>>> fast = vbt.MA.run(data.close, window=[f for f, _ in pairs], short_name="fast")
>>> slow = vbt.MA.run(data.close, window=[s for _, s in pairs], short_name="slow")
>>> pf = vbt.PF.from_signals(data, fast.ma_above(slow), fast.ma_below(slow), fees=0.001)
>>> best = pf.sharpe_ratio.idxmax()  # (1)
>>> print(round(pf.deflated_sharpe_ratio[best], 3), round(pf.prob_sharpe_ratio[best], 3))
0.995 0.737
```

1.  The same winner as above: a 25-day fast and a 100-day slow average.

The deflated Sharpe ratio gives a 99.5% probability that the best pair's true Sharpe ratio is above
what the luckiest of 72 skill-free trials would show, but it says nothing about holding. The
probabilistic Sharpe ratio compares the same pair with buying and holding Bitcoin and puts the
chance of beating it at 74%. Each test asks a different question, which is why one is rarely enough.

## Choose a robustness test \[#choose-a-robustness-test]

Different tests help explain different weaknesses. You can use the same strategy function and
compare the results in labeled pandas tables.

| What you want to check                                   | What to compare                                                                                                       |
| -------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- |
| Whether entry timing adds value                          | [Buy and hold and random entries](#baselines)                                                                         |
| Whether a precise setting drives the result              | [Nearby parameters](#parameter-sensitivity)                                                                           |
| Whether an extra rule helps                              | [The strategy with each filter on and off](#test-which-filters-help)                                                  |
| Whether execution assumptions drive the result           | [Signal delays and higher costs](#execution-sensitivity)                                                              |
| Whether the search could find a similar winner by chance | [Repeated searches on shuffled data](#perturbed-and-permuted-data) and [adjusted Sharpe measures](#statistical-tests) |

For each comparison, keep the dates, sizing, and execution assumptions consistent unless those are
what you are testing. Alongside returns, inspect [drawdowns](/features/analytics/drawdown-analysis/)
and [trade counts and exposure](/features/analytics/trade-analytics/) to see what changed.

## Baselines \[#baselines]

Every result needs something to beat. `vbt.PF.from_holding` buys once and holds, and
`vbt.PF.from_random_signals` enters and exits at random with a seed, using either a fixed number of
signals or a probability per bar.

```python title="Compare with buy and hold and 1,000 random portfolios"
>>> hold = vbt.PF.from_holding(data, fees=0.001)
>>> n_trades = pf.trades.count()[best]
>>> rand = vbt.PF.from_random_signals(data, n=[n_trades] * 1000, seed=42, fees=0.001)
>>> print(n_trades, round(hold.sharpe_ratio, 3), round(rand.sharpe_ratio.median(), 3))
8 1.126 1.015
```

The best pair beats holding and every one of the 1,000 random portfolios with the same 8 trades.
That is still one strategy chosen after the fact from 72, so it does not settle the question. The
permutation test above repeats the whole search on each shuffled history, which accounts for the
search itself.

## Parameter sensitivity \[#parameter-sensitivity]

A robust strategy has a plateau: neighboring parameters perform about as well as the best one. A
single peak surrounded by weak results usually means one lucky trade or one lucky period. Laying the
Sharpe ratios out by parameter shows which kind of result you have.

```python title="Sharpe ratio for each pair of windows"
>>> pf.sharpe_ratio.unstack("slow_window").iloc[:5, :6].round(2)
slow_window   60    80    100   120   140   160
fast_window
10           1.23  1.17  1.37  1.32  1.17  1.29
15           1.25  1.12  1.37  1.16  1.17  1.17
20           1.30  1.14  1.22  1.15  1.22  1.25
25           1.21  1.24  1.41  1.20  1.32  1.27
30           1.02  1.19  1.32  1.20  1.20  1.27
```

The eight neighbors of 25 and 100 score between 1.14 and 1.32, so this result is not an isolated
spike. The same layout works for ablation tests: pass each filter of a strategy as an on and off
parameter and compare the strategy with and without it. The
[Parameter optimization](/features/optimization/strategy-optimization/) page covers heatmaps and
large parameter grids.

### Test which filters help \[#test-which-filters-help]

An extra filter can improve the headline return simply by leaving fewer trades. An ablation test
shows what each rule contributes. Build one signal column per variant and backtest them together,
using the same prices, exits, and fees.

This small example uses made-up prices and volumes. It compares a core entry rule with a trend
filter, a volume filter, and both filters together. All signals move forward one bar before trading.

```python title="Compare a strategy with and without entry filters"
>>> close = pd.Series([100, 102, 105, 103, 99, 101, 106, 110, 107, 102, 104, 109, 114, 111, 106, 108, 115, 118, 113, 110])
>>> volume = pd.Series([90, 110, 80, 120, 90, 130, 140, 80, 110, 90, 120, 80, 140, 100, 90, 120, 80, 130, 110, 90])
>>> core = close > close.shift(1)
>>> trend_filter = close > close.rolling(4).mean()
>>> volume_filter = volume >= 100
>>> entries = pd.DataFrame({
...     "Core": core,
...     "Trend": core & trend_filter,
...     "Volume": core & volume_filter,
...     "Both": core & trend_filter & volume_filter,
... }).vbt.signals.fshift(1)
>>> exits = (close < close.shift(1)).vbt.signals.fshift(1)
>>> variants = vbt.PF.from_signals(close, entries, exits, fees=0.001, freq="1D")
>>> comparison = pd.DataFrame({
...     "Return [%]": variants.total_return * 100,
...     "Max drawdown [%]": variants.max_drawdown * 100,
...     "Closed trades": variants.trades.status_closed.count(),
... })
>>> comparison.round(2)
        Return [%]  Max drawdown [%]  Closed trades
Core        -16.28            -16.28              4
Trend       -20.11            -20.11              3
Volume      -16.28            -16.28              4
Both        -14.32            -14.32              3
```

See exactly what each rule contributes: the volume filter alone leaves the result unchanged, while
combining both filters produces the smallest loss on this sample. Use the same comparison to test
entry rules across a larger strategy, with filter switches as parameters and repeated
[walk-forward test periods](/features/optimization/time-series-cross-validation/).

## Execution sensitivity \[#execution-sensitivity]

Stress-test execution by delaying signals and raising costs. Compare how the strategy responds to a
one-bar delay, a two-bar delay, and different fee levels. Each delay is one column and each fee
level one parameter value, so the grid runs as a single backtest.

```python title="Delay the best pair's signals and raise its fees"
>>> entries = fast.ma_above(slow)[best]
>>> exits = fast.ma_below(slow)[best]
>>> delays = pd.Index([0, 1, 2], name="delay")
>>> pf_exec = vbt.PF.from_signals(
...     data,
...     pd.concat([entries.vbt.signals.fshift(d) for d in delays], axis=1, keys=delays),
...     pd.concat([exits.vbt.signals.fshift(d) for d in delays], axis=1, keys=delays),
...     fees=vbt.Param([0.001, 0.002, 0.005], name="fees"),
... )
>>> pf_exec.sharpe_ratio.unstack("fees").round(2)
fees   0.001  0.002  0.005
delay
0       1.41   1.40   1.38
1       1.33   1.33   1.31
2       1.32   1.31   1.29
```

With only 8 trades, even 0.5% fees barely move the result, while a one-bar delay costs more than all
of them. A strategy that trades more often pays more for each step in fees. The
[Backtest realism](/features/backtesting/backtest-realism/) page covers slippage, order delays, and
other execution assumptions.

## Perturbed and permuted data \[#perturbed-and-permuted-data]

A strategy that only works on the exact price history it was tuned on is fragile. Each of these
varies the data in a different way:

*   **Shuffled returns.** `.vbt.shuffle(seed=...)` shuffles each column separately, as in the first
    example, to break the original ordering while keeping the return distribution.
*   **Noise.** Adding small random noise to prices, for example with `np.random.default_rng(seed)`
    scaled to each bar's range, tests whether signals depend on exact price levels.
*   **Mirrored prices.** `data.mirror_ohlc()` turns rises into falls and keeps each bar's high and low
    consistent, which tests whether a long-only result depends on the market's direction.
*   **Block bootstrap.** `vbt.Splitter.from_n_random` draws random windows, and `shuffle_splits` with
    `replace=True` resamples existing windows, keeping the order of prices inside each block.
*   **Synthetic prices.** `vbt.GBMOHLCData` and `vbt.RandomOHLCData` generate seeded price paths with
    chosen drift and volatility, as the [Synthetic data](/features/data/synthetic-data/) page shows.

Each variant is one more column, so hundreds of perturbed histories run as one backtest.

## Statistical tests \[#statistical-tests]

The returns accessor and the portfolio both compute the Sharpe ratio tests by López de Prado and
Bailey:

| Method                  | Answers                                                                                |
| ----------------------- | -------------------------------------------------------------------------------------- |
| `sharpe_ratio_std`      | How uncertain the Sharpe ratio is, given the sample's length, skew, and kurtosis       |
| `prob_sharpe_ratio`     | How likely the true Sharpe ratio is above a benchmark's, such as buy and hold          |
| `deflated_sharpe_ratio` | How likely the Sharpe ratio exceeds a benchmark adjusted for the search across columns |

`prob_sharpe_ratio` on a portfolio compares with its benchmark, which is buying and holding the
traded asset unless you pass `bm_returns`. To compare with a Sharpe ratio of zero instead, call it
on `pf.get_returns_acc(bm_returns=False)`.

`deflated_sharpe_ratio` treats each column as one trial, so call it on the returns of every
combination you tested, not only on the winner. A single column returns NaN, and a grouped portfolio
counts each group as one trial.

### Probability of backtest overfitting \[#probability-of-backtest-overfitting]

VBT provides splitters and labeled results for calculating the probability of backtest overfitting
(PBO) in a research workflow. Compare the in-sample and out-of-sample results of every combination
over combinatorial splits, using half the folds for training and half for testing. Build those
splits with `vbt.Splitter.from_purged_kfold`, covered on the
[Walk-forward and cross-validation](/features/optimization/time-series-cross-validation/) page. PBO
is the share of splits in which the best combination on the training half ranks below the median on
the test half. Values near 0.5 or higher mean the selection does no better than picking a
combination at random.

!!! warning "Correlated trials"
    The deflated Sharpe ratio assumes the trials are independent. Neighboring parameters of one
    strategy are strongly correlated, so a grid of 72 pairs behaves like far fewer independent trials.
    The permutation example complements this by rerunning the same correlated parameter grid on each
    shuffled history.

!!! info "Tutorial"
    The members-only [Cross-validation](https://members.vectorbt.pro/tutorials/cross-validation/workflow/) tutorial
    carries each period's winning parameters into the next period and compares them with every
    alternative out of sample.


## Related pages

*   [Walk-forward and cross-validation](/features/optimization/time-series-cross-validation/): Test optimized parameters out of sample with rolling, expanding, and purged splits
*   [Performance and risk metrics](/features/analytics/performance-metrics/): Compute returns, Sharpe, Sortino, benchmark, and rolling metrics on any backtest
*   [Parameter optimization](/features/optimization/strategy-optimization/): Sweep millions of parameter combinations with grids, conditions, and random search
*   [Backtesting engine](/features/backtesting/backtesting-engine/): Simulate orders, signals, and callbacks across many assets and parameters at once