# Pairs trading (/features/strategies/pairs-trading)

How do you find, trade, and validate pairs? This pairs trading backtest in Python screens 105 crypto
pairs for cointegration, trades the ten strongest on the z-score of a rolling regression spread, and
then checks whether the relationship held while it was being traded.

VBT gives you the parts to build your own statistical arbitrage workflow: rolling regression,
threshold signals, long and short positions, shared cash, and parameter comparisons. You choose the
pair-selection rule, spread model, sizing, and exits.

```python title="Screen 105 pairs for cointegration on two years of data"
>>> from itertools import combinations
>>> import statsmodels.tsa.stattools as ts

>>> symbols = [
...     "BTCUSDT", "ETHUSDT", "BNBUSDT", "XRPUSDT", "ADAUSDT", "SOLUSDT", "DOGEUSDT",
...     "LTCUSDT", "LINKUSDT", "DOTUSDT", "TRXUSDT", "BCHUSDT", "XLMUSDT", "ETCUSDT",
...     "ATOMUSDT",
... ]
>>> data = vbt.BinanceData.pull(symbols, start="2021-01-01", end="2025-01-01")
>>> log_close = np.log(data.close)
>>> select = log_close.loc["2021":"2022"]  # (1)
>>> pairs = list(combinations(symbols, 2))  # (2)
>>> pvalues = pd.Series(
...     {(s1, s2): ts.coint(select[s1], select[s2])[1] for s1, s2 in pairs},
...     name="pvalue",
... ).rename_axis(["s1", "s2"])
>>> top = pvalues.nsmallest(10)
>>> top.round(4)
s1        s2
LINKUSDT  XLMUSDT     0.0003
BCHUSDT   XLMUSDT     0.0005
BNBUSDT   BCHUSDT     0.0005
          XLMUSDT     0.0006
          DOTUSDT     0.0007
          LTCUSDT     0.0007
          ATOMUSDT    0.0007
ADAUSDT   BCHUSDT     0.0009
BNBUSDT   LINKUSDT    0.0012
          ETCUSDT     0.0012
Name: pvalue, dtype: float64
```

1.  Pairs are chosen on 2021 and 2022 only. Trading starts in 2023, on data the screen never saw.
2.  Each unordered pair once: 15 symbols give 105 pairs.

All ten pairs pass the Engle-Granger test with p-values near 0.001. Each pair is then traded as a
group of two columns that share one cash balance:

```python title="Trade every pair on z-score thresholds with shared cash"
>>> y = pd.concat({p: data.close[p[0]] for p in top.index}, axis=1, names=["s1", "s2"])
>>> x = pd.concat({p: data.close[p[1]] for p in top.index}, axis=1, names=["s1", "s2"])
>>> zscore = vbt.OLS.run(np.log(x), np.log(y), window=60, hide_params=True).zscore  # (1)

>>> def legs(y_signal, x_signal):  # (2)
...     both = pd.concat({"y": y_signal, "x": x_signal}, axis=1, names=["leg"])
...     return both.reorder_levels(["s1", "s2", "leg"], axis=1).sort_index(axis=1)

>>> def trade_pairs(period, threshold):
...     z = zscore.loc[period]
...     wide = z.vbt.crossed_above(threshold)  # (3)
...     narrow = z.vbt.crossed_below(-threshold)
...     revert = z.vbt.crossed_above(0) | z.vbt.crossed_below(0)
...     return vbt.PF.from_signals(
...         legs(y.loc[period], x.loc[period]),
...         long_entries=legs(narrow, wide),
...         short_entries=legs(wide, narrow),
...         long_exits=legs(revert, revert),
...         short_exits=legs(revert, revert),
...         size=0.5,
...         size_type="valuepercent",  # (4)
...         fees=0.001,
...         group_by=["s1", "s2"],
...         cash_sharing=True,
...         call_seq="auto",
...     )

>>> pf = trade_pairs("2023", 2.0)
>>> metrics = ["total_return", "sharpe_ratio", "max_dd", "total_trades"]
>>> pf.stats(metrics, agg_func=None).round(2)
                   Total Return [%]  Sharpe Ratio  Max Drawdown [%]  Total Trades
s1       s2
ADAUSDT  BCHUSDT             -90.43         -1.73             94.77            12
BCHUSDT  XLMUSDT             -35.72          0.28             74.99            10
BNBUSDT  ATOMUSDT             -2.12         -0.05             13.81            12
         BCHUSDT             -55.64         -1.56             61.99            12
         DOTUSDT             -11.63         -0.57             23.42            16
         ETCUSDT               9.38          0.58             10.82            16
         LINKUSDT              7.50          0.40             20.72            12
         LTCUSDT             -16.03         -0.53             30.92            12
         XLMUSDT               3.72          0.30             15.77            16
LINKUSDT XLMUSDT             -18.92         -0.76             33.01            18
```

1.  A 60-day rolling regression of one log price on the other gives the hedge ratio and the z-score
    of the spread, using current and earlier observations.
2.  Builds one column per leg, ordered so that both legs of a pair sit next to each other.
3.  When the spread is unusually wide, short the first asset and buy the second. When it is unusually
    narrow, do the opposite. Exit both legs when the spread returns to its mean.
4.  Each leg takes half of the pair's current value, and selling runs before buying within a bar.

The side-by-side results separate the three profitable pairs from the rest, including two that lost
more than half their capital. Now inspect how the selected relationships changed:

```python title="Test the same pairs for cointegration while they were traded"
>>> later = log_close.loc["2023":"2024"]
>>> still = pd.Series(
...     {(s1, s2): ts.coint(later[s1], later[s2])[1] for s1, s2 in top.index}
... )
>>> pd.DataFrame({"2021-2022": top, "2023-2024": still}).round(3)
                   2021-2022  2023-2024
LINKUSDT XLMUSDT       0.000      0.712
BCHUSDT  XLMUSDT       0.001      0.594
BNBUSDT  BCHUSDT       0.001      0.864
         XLMUSDT       0.001      0.925
         DOTUSDT       0.001      0.964
         LTCUSDT       0.001      0.963
         ATOMUSDT      0.001      0.826
ADAUSDT  BCHUSDT       0.001      0.175
BNBUSDT  LINKUSDT      0.001      0.902
         ETCUSDT       0.001      0.825
```

The later test finds that none of the ten pairs retained significance at the 5% level. VBT brings
the screen, rolling model, and trade results into one workflow, so you can investigate changing
relationships and compare new selection rules or exits.

## Finding pairs \[#finding-pairs]

Cointegration matters more than correlation: two correlated assets can drift apart for good, while a
cointegrated pair has a spread that returns to its mean. The
[Engle-Granger test from statsmodels](https://www.statsmodels.org/stable/generated/statsmodels.tsa.stattools.coint.html)
regresses one price on the other and tests the residual for stationarity. Testing every pair of a
large universe increases the chance of false positives. At a 5% significance level, some pairs with
no cointegrating relationship will pass by chance. Choose pairs on one period and trade them on a
later one, as above.

The test is not symmetric. It regresses the first price on the second, so swapping the two gives a
somewhat different p-value and hedge ratio. Testing each unordered pair once halves the work, at the
cost of ignoring the other direction.

For a larger universe, `@vbt.parameterized` can run the test over all pairs in parallel threads, and
`vbt.Param` with a condition such as `"s1 < s2"` builds each unordered pair once. The cointegration
test itself comes from `statsmodels` and runs in plain Python, so a rolling cointegration test over
many pairs and windows is slow. Only the rolling regression is compiled. See the
[Parameter optimization](/features/optimization/strategy-optimization/) page.

## Spread signals \[#spread-signals]

You can inspect the rolling model as well as its trades. `vbt.OLS` exposes the slope, intercept,
predicted value, residual, z-score, and R-squared. Its regression and z-score normalization windows
can be set separately, so you can compare how quickly the model and its entry thresholds adapt. See
[Rolling OLS](/features/indicators/technical-indicators/#rolling-ols).

The strategy can also use a spread you calculate yourself in pandas or another model. Convert that
spread into boolean entry and exit arrays and pass them to the same simulator. For example, compare
exiting at zero with exiting inside a narrower band, using the crossing helpers on the
[Trading signals](/features/indicators/trading-signals/) page.

The rolling regression from `vbt.OLS` gives the hedge ratio, the spread, and its z-score, and
`zscore_crossed_above` and `zscore_crossed_below` turn thresholds into signals. Both legs of a pair
go into one group with shared cash, so an entry sells one asset and buys the other with the same
capital, and `call_seq="auto"` places the sell first. Because every pair is two columns of one
portfolio, ten or a thousand pairs run in one simulation, and a symbol can appear in several pairs.

The legs above are dollar-neutral: each takes half of the pair's value, and the hedge ratio only
shapes the spread and its z-score. To size the legs by the hedge ratio instead, pass a `size` array
per leg built from the regression's `slope` output. On these pairs, the 60-day slope ranged from
-1.5 to 4.1 during 2023, so hedge-ratio sizing would also swing between legs from one trade to the
next.

## Inspect each pair and each leg \[#inspect-each-pair-and-each-leg]

The example gives each pair its own cash pool, shared by its two legs. That lets you compare pairs
independently, even when the same asset appears in several pairs. Pair labels keep those positions
separate in the results. For strategies that compete for one account's capital, configure the shared
group around that account, as explained in
[Portfolio accounting](/features/backtesting/portfolio-accounting/).

A pair's combined return can hide very different results on its long and short legs. Inspect
[orders and fills](/features/backtesting/orders-and-execution/),
[trade records](/features/analytics/trade-analytics/), and `pf.get_allocations()` to see how the
positions were entered, exited, and sized. Fees and slippage can be set per asset, so you can test
what happens when one leg costs more to trade than the other.

## Validate pair selection and trading rules \[#validate-pair-selection-and-trading-rules]

The thresholds, the regression window, and the exit rule are parameters, which the next step tests
out of sample:

```python title="Choose the threshold on 2023, then test it on 2024"
>>> thresholds = [1.5, 2.0, 2.5]
>>> in_sample = {t: trade_pairs("2023", t).sharpe_ratio.mean() for t in thresholds}
>>> best = max(in_sample, key=in_sample.get)
>>> print(best, round(in_sample[best], 2))
2.5 -0.34
>>> out_pf = trade_pairs("2024", best)
>>> print(round(out_pf.sharpe_ratio.mean(), 2), (out_pf.sharpe_ratio > 0).mean())
0.05 0.3
```

The least bad threshold in 2023 also did best in 2024, but three pairs in ten were profitable and
the average Sharpe ratio was close to zero. For more windows and a rolling selection of pairs, use
the splitters on the
[Walk-forward and cross-validation](/features/optimization/time-series-cross-validation/) page.

!!! info "Tutorial"
    The members-only [Pairs trading](https://members.vectorbt.pro/tutorials/pairs-trading/) tutorial screens a
    larger universe and builds the same strategy with indicators, Numba, and a custom simulator.


## Related pages

*   [Multidimensional research](/features/tooling/multidimensional-research/): Run assets, parameters, and strategies as labeled columns, then index and stack them
*   [Backtesting engine](/features/backtesting/backtesting-engine/): Simulate orders, signals, and callbacks across many assets and parameters at once
*   [Robustness and overfitting](/features/optimization/robustness-and-overfitting/): Check whether a backtest is luck with baselines, permutation tests, and deflated Sharpe
*   [Walk-forward and cross-validation](/features/optimization/time-series-cross-validation/): Test optimized parameters out of sample with rolling, expanding, and purged splits