Features
Synthetic data
Generate reproducible random walks, GBM paths, and synthetic OHLC bars
How does a strategy behave on prices that never happened? VBT generates synthetic price data in Python: random walks, geometric Brownian motion, and full OHLC bars, reproducible with a seed and returned as the same data object as real market data. Running a strategy on hundreds of synthetic paths shows the range of results a strategy can produce under a chosen model.
paths = vbt.GBMData.pull(
[f"path_{i}" for i in range(500)],
start="2020-01-01",
end="2024-01-01",
timeframe="1d",
seed=42,
)
close = paths.get()
close.shape(1461, 500)fast = close.rolling(20).mean()
slow = close.rolling(50).mean()
pf = vbt.PF.from_signals(
close,
fast.vbt.crossed_above(slow),
fast.vbt.crossed_below(slow),
fees=0.001,
)
pf.sharpe_ratio.describe().round(3)count 500.000
mean -0.096
std 0.494
min -1.598
25% -0.437
50% -0.129
75% 0.224
max 1.754
Name: sharpe_ratio, dtype: float64print(round((pf.sharpe_ratio > 0).mean(), 3))0.406pf.sharpe_ratio.vbt.histplot(trace_kwargs=dict(nbinsx=40)).show()Even with no predictable trend built into the model, the crossover had a positive Sharpe ratio on 41% of the paths, and the best path reached 1.75. This gives you a comparison for the same strategy on historical data. Look at the distribution across paths, rather than treating the best simulated result as a universal pass mark.
What can you test with synthetic data?
Use generated paths when you want control over the conditions:
- Compare strategy changes on the same seeded prices.
- Try calmer or more volatile markets without changing the strategy.
- Exercise stop, limit, and signal logic on generated OHLC bars.
- Check whether an unexpectedly strong result also appears on random data.
The generated data works with the same indicators, backtests, and plots as downloaded data, and needs no data-provider account. For broader validation on historical data, see robustness and overfitting.
Generators
vbt.RandomData compounds normally distributed returns, and vbt.GBMData simulates geometric
Brownian motion. Both take a start value, a mean, a standard deviation, and a seed, which can differ
per symbol, and both produce any number of symbols on any index or timeframe. With a seed, every run
gives the same paths, so experiments on synthetic data are as reproducible as those on stored data.
For random returns, symmetric=True transforms losses so that a gain and a loss of the same size
cancel out multiplicatively, which removes the downward drift of plain returns.
You can use an asset's returns to estimate the generator parameters. RandomData takes the mean and
standard deviation of sampled returns. For GBM, choose drift, volatility, and dt consistently with
the time step you are modeling. These describe a simple model: volatility stays constant, and there
is no clustering, no autocorrelation, and no fat tails. These generators test whether a rule finds
structure in noise, not how it would trade a real market. For more realistic dynamics, write your
own model as shown below.
Compare market scenarios
Parameters can differ by symbol, so each symbol can represent a scenario. Here both paths use the same random draws, with different volatility settings:
scenarios = vbt.GBMData.pull(
["calm", "volatile"],
start="2024-01-01",
periods=252,
timeframe="1d",
mean=0.0,
std=vbt.symbol_dict({"calm": 0.005, "volatile": 0.02}),
seed=42,
split_seed=False,
)
scenarios.get().shape(252, 2)(scenarios.get().pct_change().std() * 100).round(2)symbol
calm 0.48
volatile 1.94
dtype: float64The output is the sample standard deviation of daily returns, in percent. split_seed=False keeps
the random draws the same across these scenarios so you can isolate the volatility change. By
default, VBT derives a separate seed for each symbol, as in the 500-path example. Pass
scenarios.get() to your strategy to compare how it behaves in the two settings.
Synthetic OHLC bars
vbt.RandomOHLCData and vbt.GBMOHLCData simulate ticks within each bar and aggregate them into
open, high, low, and close, so stops, limits, and intrabar logic can be tested as well. The number
of simulated ticks is controlled by n_ticks, which can also vary by bar. The
Synthetic OHLC highlight below plots the resulting candles.
Real data can be turned into synthetic data too. data.mirror_ohlc() reverses log returns and
transforms the other OHLC prices around the mirrored path. Compare a strategy on the original and
its mirror to explore how much its behavior changes when price moves reverse direction.
Custom generators
Any model becomes a data class. Subclass vbt.SyntheticData and return a series for each symbol,
and the class handles symbols, seeds, and indexes like any other data source. You can also provide
an update method that continues your model from its previous state. Here returns follow a Student's
t distribution, which produces the fat tails that normal returns lack:
class StudentTData(vbt.SyntheticData):
@classmethod
def generate_key(
cls, key, index, df=3, scale=0.01, start_value=100.0, seed=None, **kwargs
):
if seed is not None:
np.random.seed(seed)
returns = np.random.standard_t(df, size=len(index)) * scale
return pd.Series(start_value * np.cumprod(1 + returns), index=index)
fat = StudentTData.pull(
["A", "B"],
start="2024-01-01",
end="2025-01-01",
timeframe="1d",
seed=42,
)
fat.get().pct_change().kurt().round(2) symbol
A 5.60
B 10.26
dtype: float64✅ New basic models are available for generating synthetic OHLC data. These are especially useful for leakage detection.
data = vbt.GBMOHLCData.pull("R", start="2022-01", end="2022-04", seed=42)
data.plot().show()Related pages
- Backtesting › Backtesting engineSimulate orders, signals, and callbacks across many assets and parameters at once
- Optimization and validation › Robustness and overfittingCheck whether a backtest is luck with baselines, permutation tests, and deflated Sharpe
- Research toolkit › Multidimensional researchRun assets, parameters, and strategies as labeled columns, then index and stack them
- Indicators and signals › Machine learning labelsCreate triple-barrier, trend, and pivot labels, and backtest model predictions
Copyright © 2021–2026 Oleg Polakow. All rights reserved.
Site content and documentation are provided for using and evaluating VectorBT PRO and for educational purposes. Any other use, including building or supporting competing products or services, requires prior written consent.