# Rust engine (/features/performance/native-rust)

Bring Rust speed to your Python research, or write the entire strategy in Rust. VBT PRO ships a
native Rust backtesting engine for both workflows. Python users get precompiled calculations through
the optional extension. Rust users can also write stateful strategy callbacks and run independent
simulations across CPU cores.

## Choose your route to Rust \[#choose-your-route-to-rust]

| What you want to do                                                      | Where to start                                                                                 |
| ------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------- |
| Speed up supported calculations in an existing Python notebook or script | Install the optional extension and keep using the Python API.                                  |
| Keep custom trading rules written in Python                              | Use Numba callbacks, with Rust available for supported calculations elsewhere in the pipeline. |
| Write a Rust strategy that reacts to fills, cash, or positions           | Use the native simulators and give the strategy its own state.                                 |
| Process new bars as they arrive in a Rust program                        | Use a streaming stepper to carry the portfolio and strategy state forward.                     |

The Python extension and native Rust library share calculation code. You can adopt the extension
without learning Rust, and use the native library to build standalone Rust strategies. Native Rust
programs can use the library without a Python runtime.

## A stateful strategy in Python and Rust \[#a-stateful-strategy-in-python-and-rust]

The same strategy below runs twice: once as Numba callbacks through the Python API, once as a Rust
strategy. It trades 100 synthetic assets over 50,000 hourly bars, enters on a moving average
crossover, exits on the opposite cross or a 3% stop loss or 6% take profit, and waits 24 bars after
any stop before entering that asset again. Whether that wait applies depends on fills the simulation
itself produces, so it cannot be precomputed as an array.

```python title="Write the cooldown as Numba callbacks and run it through the Python API"
>>> from collections import namedtuple

>>> COOLDOWN_BARS = 24
>>> np.random.seed(42)
>>> returns = np.random.normal(0, 0.01, size=(50_000, 100))
>>> close = pd.DataFrame(100 * np.exp(returns.cumsum(axis=0)))
>>> fast = close.rolling(20).mean()
>>> slow = close.rolling(100).mean()
>>> entries = fast.vbt.crossed_above(slow)
>>> exits = fast.vbt.crossed_below(slow)
>>> Memory = namedtuple("Memory", ["blocked_until"])

>>> @njit
... def cooldown_signal_func_nb(c, entries, exits, memory):
...     cooling = c.i < memory.blocked_until[c.col]
...     entry = vbt.nb.flex_select_nb(entries, c.i, c.col) and not cooling
...     exit = vbt.nb.flex_select_nb(exits, c.i, c.col)
...     return entry, exit, False, False

>>> @njit
... def cooldown_post_order_func_nb(c, memory):
...     if vbt.pf_nb.order_closed_position_nb(c):  # (1)
...         if vbt.pf_nb.get_last_order_nb(c)["stop_type"] >= 0:
...             memory.blocked_until[c.col] = c.i + 1 + COOLDOWN_BARS

>>> memory = Memory(blocked_until=np.zeros(close.shape[1], dtype=int_))
>>> pf = vbt.PF.from_signals(
...     close,
...     signal_func_nb=cooldown_signal_func_nb,
...     signal_args=(vbt.Rep("entries"), vbt.Rep("exits"), memory),
...     broadcast_named_args=dict(entries=entries, exits=exits),
...     post_order_func_nb=cooldown_post_order_func_nb,
...     post_order_args=(memory,),
...     size=1.0,
...     fees=0.001,
...     sl_stop=0.03,
...     tp_stop=0.06,
... )
```

1.  Runs after every order. When an order closed the position and its record carries a stop type, the
    asset is blocked for the next 24 bars.

```rust title="Write the same cooldown as a Rust strategy"
const COOLDOWN_BARS: usize = 24;

struct Cooldown<'a> {
    entries: ArrayView2<'a, bool>,
    exits: ArrayView2<'a, bool>,
    blocked_until: Vec<usize>, // (1)
}

impl SignalStrategy for Cooldown<'_> {
    fn signals(&mut self, ctx: &mut SignalContext<'_, '_>) -> VbtResult<Signals> {
        let cell = [ctx.i(), ctx.col()];
        let cooling = ctx.i() < self.blocked_until[ctx.local_col()];
        Ok(Signals {
            long_entry: self.entries[cell] && !cooling,
            long_exit: self.exits[cell],
            ..Signals::default()
        })
    }

    fn on_order_end(
        &mut self,
        ctx: &mut OrderEndContext<'_, '_, SignalOrderRecord>,
        report: &ExecutionReport,
    ) -> VbtResult<()> {
        if report.before.position != 0.0 && report.after.position == 0.0 {
            let record = ctx.order_record_now().expect("a close emits a record");
            if record.stop_type >= 0 {
                self.blocked_until[ctx.local_col()] = ctx.i() + 1 + COOLDOWN_BARS;
            }
        }
        Ok(())
    }
}
```

1.  A Rust strategy owns its state, so the memory tuple becomes a field.

```rust title="Run it on one core and on all cores"
let simulator = SignalSimulator::builder()
    .config(
        SimulationConfig::builder()
            .target_shape(close.dim())
            .group_lens(group_lens.view())
            .close(close.view())
            .build()?,
    )
    .signal_config(
        SignalConfig::builder().size(1.0).fees(0.001).sl_stop(0.03).tp_stop(0.06).build(),
    )
    .build()?;
let new_cooldown = |info: GroupInfo| {
    Ok(Cooldown {
        entries: entries.view(),
        exits: exits.view(),
        blocked_until: vec![0; info.to_col - info.from_col],
    })
};
let serial = simulator.run(new_cooldown)?; // (1)
let parallel = simulator.run_parallel(new_cooldown)?;
```

1.  `run` creates one strategy per group and runs the groups in order. `run_parallel` distributes
    independent groups across CPU cores, and the combined output keeps the same order.

Both versions produced the same 65,278 orders, matching on every column, bar, size, and price.
Measured on an Apple M3 with 8 cores, best of five runs:

| Run                                           | Milliseconds |
| --------------------------------------------- | -----------: |
| Numba callbacks through `vbt.PF.from_signals` |        1,129 |
| Rust strategy, one core                       |          538 |
| Rust strategy, all cores                      |          201 |

The Python time includes argument preparation and portfolio construction, and the Rust times cover
the simulation alone. The table therefore measures the two workflows at different levels. The
research script that produced it builds the Rust program, runs both versions, and compares their
orders.

## Drop-in Rust kernels \[#drop-in-rust-kernels]

Python users get the Rust engine without writing Rust. With the optional `vectorbtpro-rust` package
installed, supported functions such as rolling statistics, signal helpers, return metrics, and the
`from_signals` and `from_orders` simulators run as precompiled Rust kernels, and fall back to Numba
for anything Rust does not cover. Install it from the platform wheels of the matching release before
installing `vectorbtpro`, as described in the
[installation guide](https://members.vectorbt.pro/installation/#rust-backend). VBT uses it only when its version
matches exactly. `jitted=dict(jitter="rs", parallel=True)` spreads supported calculations across
cores, over columns or independent portfolio groups.

Because the kernels are compiled ahead of time, the first call in a session does not wait for Numba.
On the machine that built this page, the first call of a rolling standard deviation in a fresh
process took about 0.08 seconds with Rust. Numba took 1.3 seconds the first time it saw those
inputs, since it compiled the function, and about 0.17 seconds in later processes, which load the
compiled version from disk. Later calls in the same process took under a millisecond either way, so
short scripts, scheduled jobs, and new worker processes gain the most. The
[Compute backends](/features/performance/compute-backends/) page shows when Rust wins and how to
choose a backend per call, and the [Rust backend](#rust-backend) highlight below shows the controls.

## Native simulators \[#native-simulators]

The same simulators are a Rust library. A Rust program can run a whole history at once with
`run_single`, `run`, or `run_parallel`, or feed bars one at a time to a stepper such as
`SignalStepper` and receive each bar's fills as they happen, with checkpoints to save and restore
its state. You can replay a history and then continue as new bars arrive, keeping cash, positions,
and strategy memory such as the cooldown above. See
[live simulation](/features/backtesting/live-simulation/) for this workflow.

Batch and streaming runs can be compared with `ensure_simulation_eq`, which checks records and
terminal portfolio state and reports the first mismatch. This gives you a way to check that moving a
strategy from a full-history run to a bar-by-bar feed preserves its simulated trades. The crate's
test suite compares its simulators with the Numba ones, as the example above did with orders.

## Rust callbacks \[#rust-callbacks]

Strategies in Rust are types that implement a trait, such as `SignalStrategy` for signals,
`OrderStrategy` for one order per element, and `FlexOrderStrategy` for any number of orders per bar.
Their methods receive a context with the current bar, column, cash, and positions, and an execution
report after each order. Other native engines take callbacks for signal generation, apply and reduce
operations, and portfolio allocation.

Each parallel simulation group gets its own strategy instance. A group can contain several assets
that share cash, so you can keep portfolio-level decisions together while running independent
portfolios in parallel. Rows and callbacks within each group still run in order.

Rust callbacks run inside Rust programs. From Python, strategy logic is written as Numba callbacks
like the ones above, which run on the Numba simulators. Errors in Rust carry a category, such as an
invalid value or a rejected order, and surface in Python as the matching exception type.

## Bring Rust results back to Python \[#bring-rust-results-back-to-python]

A native Rust simulation can save its output to a NumPy-compatible archive with `save_npz`. Load it
with `vbt.load_simulation_npz` and build a Portfolio using the same prices, labels, initial cash,
and grouping as the simulation. You can then use VBT's
[performance analysis](/features/analytics/performance-metrics/) and
[trade analytics](/features/analytics/trade-analytics/) on the Rust results. This lets you run the
simulation in Rust and keep your Python analysis and reporting workflow.

!!! info "Tutorial"
    The members-only [From Python to Rust](https://members.vectorbt.pro/tutorials/from-python-to-rust/) tutorial
    builds one strategy with the high-level API, Numba kernels, Python bindings, and native Rust, then
    moves it to streaming.

## Native Rust simulators \[#native-rust-simulators]

New in v2026.9.5

✅ Take your strategy directly to Rust. VBT's native simulators let the same strategy process a full
price array or step through individual bars as they arrive. Strategy callbacks can inspect cash,
positions, and orders during execution, keeping portfolio state within reach of your trading logic.

!!! note "Note"
    The examples below share the strategy definition and require `vectorbtpro-rust` and
    `ndarray = "0.16"`. See the [Rust setup guide](/documentation/rust/) for installation.

```rust title="Define a strategy with price and position conditions"
use vectorbtpro_rust::error::VbtResult;
use vectorbtpro_rust::portfolio::enums::Order;
use vectorbtpro_rust::portfolio::simulator::{
    FnOrderStrategy, OrderContext, OrderStrategy,
};

fn buy_the_dip() -> impl OrderStrategy {
    FnOrderStrategy::new(|ctx: &OrderContext<'_, '_>| {
        let price = ctx.close(ctx.col());
        let position = ctx.position(ctx.col());
        let size = if price <= 100.0 && position == 0.0 { // (1)
            1.0
        } else if price >= 110.0 && position > 0.0 {
            -position
        } else {
            return Ok(None);
        };
        Ok(Some(Order::builder().size(size).build()))
    })
}
```

1.  Buy one share only when flat. The position check prevents another buy at $98.

=== "Batch"
    ```rust title="Run a batch simulation"
    use ndarray::array;
    use vectorbtpro_rust::portfolio::simulator::{OrderSimulator, SimulationConfig};

    fn main() -> VbtResult<()> {
        let close = array![[100.0], [98.0], [105.0], [112.0]];
        let groups = array![1];
        let config = SimulationConfig::builder()
            .target_shape(close.dim())
            .group_lens(groups.view())
            .close(close.view())
            .init_cash(1000.0)
            .build()?;
        let simulator = OrderSimulator::builder().config(config).build()?;
        let output = simulator.run_single(&mut buy_the_dip())?;
        for order in &output.order_records {
            println!("row {}: {:.0} share at ${:.0}", order.idx, order.size, order.price);
        }
        Ok(())
    }
    ```

    ```text title="Output"
    row 0: 1 share at $100
    row 3: 1 share at $112
    ```

=== "Streaming"
    ```rust title="Run a streaming simulation"
    use ndarray::array;
    use vectorbtpro_rust::portfolio::simulator::{OrderStepper, StreamConfig, StreamRow};

    fn main() -> VbtResult<()> {
        let close = array![[100.0], [98.0], [105.0], [112.0]];
        let groups = array![1];
        let config = StreamConfig::builder()
            .group_lens(groups.view())
            .init_cash(1000.0)
            .build()?;
        let mut live = OrderStepper::new_single(config, None, buy_the_dip())?;
        for row in close.rows() {
            let step = live.step(StreamRow::builder().close(row).build())?; // (1)
            for order in step.order_records() {
                println!("row {}: {:.0} share at ${:.0}", order.idx, order.size, order.price);
            }
        }
        Ok(())
    }
    ```

    1.  Preserve portfolio state between calls and return the current bar's fills. The stepper does not
        retain the complete record history by default.

    ```text title="Output"
    row 0: 1 share at $100
    row 3: 1 share at $112
    ```

!!! info "Tutorial"
    Learn more in the [From Python to Rust](/tutorials/from-python-to-rust/) tutorial.

## Rust backend \[#rust-backend]

New in v2026.6.27

✅ Install the optional `vectorbtpro-rust` extension and compatible jitted calls can take the Rust
fast lane automatically. VBT exposes the extension as `vbt.rs`, registers Rust kernels under
`jitted="rs"`, and falls back to the normal implementation, usually Numba, when Rust is unavailable
or unsupported.

=== "Out of the box"
    ```python title="Keep your workflow, get the Rust lane"
    >>> data = vbt.YFData.pull("BTC-USD", start="2024")

    >>> fast_ma = data.close.vbt.rolling_mean(20)  # (1)
    >>> slow_ma = data.close.vbt.rolling_mean(50)
    >>> entries = fast_ma.vbt.crossed_above(slow_ma)  # (2)
    >>> exits = fast_ma.vbt.crossed_below(slow_ma)

    >>> pf = vbt.Portfolio.from_signals(  # (3)
    ...     data,
    ...     entries=entries,
    ...     exits=exits,
    ...     sl_stop=0.05,
    ...     tp_stop=0.15,
    ...     fees=0.001,
    ... )
    ```

    1.  `rolling_mean_nb` routes to `vbt.rs.generic.rolling.rolling_mean_rs`.
    2.  `crossed_above_nb` routes to `vbt.rs.generic.base.crossed_above_rs`.
    3.  `from_signals_nb` routes to `vbt.rs.portfolio.from_signals.from_signals_rs`.

=== "Speed check"
    ```python title="Compare Numba, parallel Numba, Rust, and parallel Rust"
    >>> np.random.seed(42)
    >>> returns = pd.DataFrame(np.random.normal(0, 0.01, size=(200_000, 64)))  # (1)

    >>> nb = lambda: returns.vbt.returns.rolling_profit_factor(
    ...     window=50,
    ...     jitted="nb"
    ... )
    >>> nb_parallel = lambda: returns.vbt.returns.rolling_profit_factor(
    ...     window=50,
    ...     jitted=dict(jitter="nb", parallel=True)
    ... )
    >>> rs = lambda: returns.vbt.returns.rolling_profit_factor(
    ...     window=50,
    ...     jitted="rs"
    ... )
    >>> rs_parallel = lambda: returns.vbt.returns.rolling_profit_factor(
    ...     window=50,
    ...     jitted=dict(jitter="rs", parallel=True)
    ... )

    >>> nb(); nb_parallel(); rs(); rs_parallel()  # (2)

    >>> %timeit nb()
    1.21 s ± 655 μs per loop (mean ± std. dev. of 7 runs, 1 loop each)

    >>> %timeit nb_parallel()
    214 ms ± 15.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

    >>> %timeit rs()
    272 ms ± 736 μs per loop (mean ± std. dev. of 7 runs, 1 loop each)

    >>> %timeit rs_parallel()
    54.6 ms ± 1.23 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)
    ```

    1.  Use a wide enough array to give parallel kernels room to stretch.
    2.  Warm up first so Numba compilation stays out of the race. These timings were measured on an
        Apple M3 and will vary by machine.

=== "AutoBench"
    ```bash title="Build the benchmark cache first"
    python -m vectorbtpro.benchmarks.bench_matrix_cli --full
    ```

    !!! info "Note"
        The full benchmark matrix can take several hours to finish, depending on your machine.

    ```python title="Let local timings pick the lane"
    >>> vbt.settings.jitting["backends"]["auto_bench_path"] = "benchmarks/cache.json"  # (1)
    >>> vbt.settings.jitting["backends"]["auto_mode"] = "bench"  # (2)

    >>> data = vbt.YFData.pull("BTC-USD", start="2024")
    >>> sharpe_ratio = data.returns.vbt.returns.sharpe_ratio()  # (3)
    ```

    1.  Run benchmarks before enabling AutoBench. The matrix command writes this cache by default.
        Change the path if you pass a custom `--output-dir` or `--cache`.
    2.  Use `"bench"` for AutoBench or `"bench_mixed"` for AutoBenchMixed. AutoBenchMixed can pick
        across serial and parallel candidates automatically.
    3.  Leave `jitted` empty. Compatible calls consult your machine's benchmark cache automatically.

=== "Control panel"
    ```python title="Tune Rust preference"
    >>> vbt.settings.jitting["backends"]["auto_mode"] = False  # (1)
    >>> vbt.settings.jitting["resolve_overrides"]["returns_module"] = dict(  # (2)
    ...     match="vectorbtpro.returns",
    ...     resolve_kwargs=dict(auto_mode=True),
    ... )

    >>> data = vbt.YFData.pull("BTC-USD", start="2024")
    >>> sharpe_ratio = data.returns.vbt.returns.sharpe_ratio()
    ```

    1.  Disable Rust-first automatic dispatch globally.
    2.  Re-enable automatic backend dispatch only for returns-module calls.

=== "Raw Rust, fastest lane"
    ```python title="Resolve or call the Rust function directly"
    >>> total_return_rs = vbt.resolve_jitted(  # (1)
    ...     vbt.ret_nb.total_return_nb,
    ...     jitted="rs",
    ...     use_backend_wrapper=False,
    ... )

    >>> data = vbt.YFData.pull("BTC-USD", start="2024")
    >>> returns = data.returns.vbt.to_2d_array()  # (2)

    >>> total_return_rs(returns)  # (3)
    array([0.652341])

    >>> vbt.rs.returns.total_return_rs(returns)  # (4)
    array([0.652341])
    ```

    1.  Let VBT find the Rust backend registered for `total_return_nb`. With
        `use_backend_wrapper=False`, this returns the raw Rust function itself.
    2.  Prepare arguments. Raw kernels expect raw arrays, not wrapped Series/DataFrames.
    3.  Call the resolved Rust function with raw-array arguments.
    4.  Or import the same Rust function from its module path. The API page for each Numba function
        lists its Rust counterpart path when one exists.

    !!! info
        Raw Rust functions bypass the registry's compatibility checks, argument preparation, and
        fallbacks. Use them only when you have prepared arguments in the exact format expected by the
        Rust function.


## Related pages

*   [Compute backends](/features/performance/compute-backends/): Run kernels with Numba, parallel Numba, Rust, or NumPy and choose the backend per call
*   [Benchmarks](/features/performance/benchmarks/): Compare simulators on the same orders and measure each backend on your own machine
*   [Event-driven backtesting](/features/backtesting/event-driven-backtesting/): Write compiled callbacks for cooldowns, position limits, and custom simulators
*   [Backtesting engine](/features/backtesting/backtesting-engine/): Simulate orders, signals, and callbacks across many assets and parameters at once