# Synthetic data (/features/data/synthetic-data)

How does a strategy behave on prices that never happened? VBT generates synthetic price data in
Python: random walks, geometric Brownian motion, and full OHLC bars, reproducible with a seed and
returned as the same data object as real market data. Running a strategy on hundreds of synthetic
paths shows the range of results a strategy can produce under a chosen model.

```python title="Run one strategy on 500 random price paths"
>>> paths = vbt.GBMData.pull(  # (1)
...     [f"path_{i}" for i in range(500)],
...     start="2020-01-01",
...     end="2024-01-01",
...     timeframe="1d",
...     seed=42,
... )
>>> close = paths.get()
>>> close.shape
(1461, 500)

>>> fast = close.rolling(20).mean()
>>> slow = close.rolling(50).mean()
>>> pf = vbt.PF.from_signals(
...     close,
...     fast.vbt.crossed_above(slow),
...     fast.vbt.crossed_below(slow),
...     fees=0.001,
... )
>>> pf.sharpe_ratio.describe().round(3)
count    500.000
mean      -0.096
std        0.494
min       -1.598
25%       -0.437
50%       -0.129
75%        0.224
max        1.754
Name: sharpe_ratio, dtype: float64
>>> print(round((pf.sharpe_ratio > 0).mean(), 3))
0.406

>>> pf.sharpe_ratio.vbt.histplot(trace_kwargs=dict(nbinsx=40)).show()
```

1.  Geometric Brownian motion with zero drift and 1% daily volatility, one independent path per
    symbol. No predictable trend is built into the model.

Histogram of Sharpe ratios of a moving average crossover strategy on 500 synthetic random price paths. [Figure data (JSON)](/assets/figures/features/data/synthetic-sharpe-distribution.de0439809b7d.json)

Even with no predictable trend built into the model, the crossover had a positive Sharpe ratio on
41% of the paths, and the best path reached 1.75. This gives you a comparison for the same strategy
on historical data. Look at the distribution across paths, rather than treating the best simulated
result as a universal pass mark.

## What can you test with synthetic data? \[#what-can-you-test-with-synthetic-data]

Use generated paths when you want control over the conditions:

*   Compare strategy changes on the same seeded prices.
*   Try calmer or more volatile markets without changing the strategy.
*   Exercise stop, limit, and signal logic on generated OHLC bars.
*   Check whether an unexpectedly strong result also appears on random data.

The generated data works with the same indicators, backtests, and plots as downloaded data, and
needs no data-provider account. For broader validation on historical data, see
[robustness and overfitting](/features/optimization/robustness-and-overfitting/).

## Generators \[#generators]

`vbt.RandomData` compounds normally distributed returns, and `vbt.GBMData` simulates geometric
Brownian motion. Both take a start value, a mean, a standard deviation, and a seed, which can differ
per symbol, and both produce any number of symbols on any index or timeframe. With a seed, every run
gives the same paths, so experiments on synthetic data are as reproducible as those on stored data.
For random returns, `symmetric=True` transforms losses so that a gain and a loss of the same size
cancel out multiplicatively, which removes the downward drift of plain returns.

You can use an asset's returns to estimate the generator parameters. `RandomData` takes the mean and
standard deviation of sampled returns. For GBM, choose drift, volatility, and `dt` consistently with
the time step you are modeling. These describe a simple model: volatility stays constant, and there
is no clustering, no autocorrelation, and no fat tails. These generators test whether a rule finds
structure in noise, not how it would trade a real market. For more realistic dynamics, write your
own model as shown below.

## Compare market scenarios \[#compare-market-scenarios]

Parameters can differ by symbol, so each symbol can represent a scenario. Here both paths use the
same random draws, with different volatility settings:

```python title="Compare volatility scenarios with the same random draws"
>>> scenarios = vbt.GBMData.pull(
...     ["calm", "volatile"],
...     start="2024-01-01",
...     periods=252,
...     timeframe="1d",
...     mean=0.0,
...     std=vbt.symbol_dict({"calm": 0.005, "volatile": 0.02}),
...     seed=42,
...     split_seed=False,
... )
>>> scenarios.get().shape
(252, 2)

>>> (scenarios.get().pct_change().std() * 100).round(2)
symbol
calm        0.48
volatile    1.94
dtype: float64
```

The output is the sample standard deviation of daily returns, in percent. `split_seed=False` keeps
the random draws the same across these scenarios so you can isolate the volatility change. By
default, VBT derives a separate seed for each symbol, as in the 500-path example. Pass
`scenarios.get()` to your strategy to compare how it behaves in the two settings.

## Synthetic OHLC bars \[#synthetic-ohlc-bars]

`vbt.RandomOHLCData` and `vbt.GBMOHLCData` simulate ticks within each bar and aggregate them into
open, high, low, and close, so stops, limits, and intrabar logic can be tested as well. The number
of simulated ticks is controlled by `n_ticks`, which can also vary by bar. The
[Synthetic OHLC](#synthetic-ohlc) highlight below plots the resulting candles.

Real data can be turned into synthetic data too. `data.mirror_ohlc()` reverses log returns and
transforms the other OHLC prices around the mirrored path. Compare a strategy on the original and
its mirror to explore how much its behavior changes when price moves reverse direction.

## Custom generators \[#custom-generators]

Any model becomes a data class. Subclass `vbt.SyntheticData` and return a series for each symbol,
and the class handles symbols, seeds, and indexes like any other data source. You can also provide
an update method that continues your model from its previous state. Here returns follow a Student's
t distribution, which produces the fat tails that normal returns lack:

```python title="Generate fat-tailed price paths with a custom class"
>>> class StudentTData(vbt.SyntheticData):
...     @classmethod
...     def generate_key(
...         cls, key, index, df=3, scale=0.01, start_value=100.0, seed=None, **kwargs
...     ):
...         if seed is not None:
...             np.random.seed(seed)
...         returns = np.random.standard_t(df, size=len(index)) * scale
...         return pd.Series(start_value * np.cumprod(1 + returns), index=index)

>>> fat = StudentTData.pull(
...     ["A", "B"],
...     start="2024-01-01",
...     end="2025-01-01",
...     timeframe="1d",
...     seed=42,
... )
>>> fat.get().pct_change().kurt().round(2)  # (1)
symbol
A     5.60
B    10.26
dtype: float64
```

1.  Excess kurtosis well above zero, the value for normally distributed returns.

## Synthetic OHLC \[#synthetic-ohlc]

New in 1.2.1

✅ New basic models are available for generating synthetic OHLC data. These are especially useful for
leakage detection.

```python title="Generate 3 months of synthetic data using Geometric Brownian Motion"
>>> data = vbt.GBMOHLCData.pull("R", start="2022-01", end="2022-04", seed=42)
>>> data.plot().show()
```

Three months of synthetic OHLC data generated with geometric Brownian motion. [Figure data (JSON)](/assets/figures/features/data/synthetic-ohlc.2cf9918db1d2.json)


## Related pages

*   [Backtesting engine](/features/backtesting/backtesting-engine/): Simulate orders, signals, and callbacks across many assets and parameters at once
*   [Robustness and overfitting](/features/optimization/robustness-and-overfitting/): Check whether a backtest is luck with baselines, permutation tests, and deflated Sharpe
*   [Multidimensional research](/features/tooling/multidimensional-research/): Run assets, parameters, and strategies as labeled columns, then index and stack them
*   [Machine learning labels](/features/indicators/machine-learning-labels/): Create triple-barrier, trend, and pivot labels, and backtest model predictions