Features

Benchmarks

Compare simulators on the same orders and measure each backend on your own machine

See VBT's backtesting speed across millions of bars and orders. Explore the interactive comparisons below, from high-level Python calls to low-level Rust simulation. Each result uses matching data and orders, passes a correctness check, and comes with code you can run on your own machine.

Compare backtesting speed

Choose a workload and the number of rows per column below. Sparse orders show the cost of scanning a long history with occasional trades. Dense orders show the cost when every bar produces an order. The high-level results are the starting point if you want to use VBT's Python portfolio API.

Benchmark packageDownload ZIP

1,000,000 one-minute bars and 1,388 filled orders per column.

High-level
320M bars/s
25 ms for this workload
Mid-level
579M bars/s
14 ms for this workload, 1.8x high-level
Low-level
4.7B bars/s
1.7 ms for this workload, 14.6x high-level
Sparse simulator benchmarks at 1,000,000 rows per column
SimulatorVersionModeStatusColumnsWarm throughput (M bars/s)First-call throughput (M bars/s)Warm runtime (s)First-call runtime (s)Compilation (s)
PRO low-level Rust2026.6.27ParallelCompleted84665.743238317416—0.001714625——
PRO low-level Numba2026.6.27ParallelCompleted84163.32302204300052.1160832578238130.001921541988849643.78056958317756653.778648041188717
PRO low-level Rust2026.6.27SerialCompleted11424.9236240937485—0.000701792——
PRO low-level Numba2026.6.27SerialCompleted11226.99398639811830.3510514305229050.00081499991938471792.848585458006712.847770458087325
PRO mid-level Rust2026.6.27ParallelCompleted8579.383003451964—0.013807792——
PRO high-level Rust2026.6.27ParallelCompleted8319.6106135270101—0.025030457880347967——
Manifold-BT0.19.0ParallelCompleted8190.96269066431148—0.041893——
PRO mid-level Rust2026.6.27SerialCompleted1112.24867661616487—0.008908791——
Manifold-BT0.19.0SerialCompleted1109.29210028648235—0.0091497921384871——
PRO mid-level Numba2026.6.27ParallelCompleted8103.393794030940540.51608457685265530.0773740829899907115.5013351663947115.42396108340472
PRO high-level Numba2026.6.27ParallelCompleted895.374864984275660.50732255504796010.0838795420713722715.76906037470325815.685180832631886
PRO high-level Rust2026.6.27SerialCompleted182.13186841345684—0.012175541836768389——
PRO mid-level Numba2026.6.27SerialCompleted171.250019799195840.085270530565565190.01403508381918072711.7273809998296211.71334591601044
PRO high-level Numba2026.6.27SerialCompleted154.482551433632680.083487449425977150.0183545001782476911.9778482499532411.959493749774992
OSS Numba1.1.0SerialCompleted121.4367055113375960.336708461622389930.0466489591635763652.9699283326044682.9232793734408915
RaptorBT0.9.0SerialCompleted120.865011249168635—0.04792712489143014——
Nautilus Python v22.0.0rc3SerialCompleted10.7077891993928654—1.412850041873753——
Nautilus Rust v20.62.0SerialCompleted10.5931788730150103—1.6858321250000001——
Backtrader1.9.78.123ParallelCompleted80.12284478022451349—65.12283212505281——
PyBroker Numba2.0.0SerialCompleted10.081236121315575770.0739374233395184612.30979500012472313.5249506249092521.2151556247845292
Backtrader1.9.78.123SerialCompleted10.020621870727887733—48.49220583308488——
Horizontal ranking of sparse simulator benchmarks at 1,000,000 rows per column by warm bar throughput, with first-call throughput shown as a translucent marker

Environment: Apple M3 (8 logical cores), arm64, macOS 26.5.2

Each simulator receives the same one-minute bars and the same orders: a sparse workload with a few orders per column, or a dense one with an order on every bar. A result counts only after its filled orders, final positions, and portfolio value match. VBT runs at three API levels: high-level portfolio construction, mid-level simulation with prepared inputs, and low-level order processing, each on Numba and on Rust.

With 1,000,000 bars in each of 8 columns run in parallel, the sparse workload took 1.7 ms on the low-level Rust simulator and 25 ms through high-level Rust, against 42 ms for Manifold-BT and 65 seconds for Backtrader. In the dense workload, with 8,000,000 orders, high-level Rust took 178 ms, Manifold-BT 213 ms, and Backtrader did not finish within the 300-second limit. The translucent markers show the first call: Numba compiles on first use, which took about 16 seconds at the high level, while Rust is compiled ahead of time. All numbers were measured on an Apple M3 with 8 cores, with each library's version shown in the chart's tooltips. The ZIP contains the full harness to rerun them.

Which benchmark matches your research?

Your workflowWhat to compare
Build portfolios from Python and inspect their resultsHigh-level calls, including input preparation and portfolio construction.
Reuse prepared arrays in repeated simulationsMid-level calls that time the simulator with those inputs ready.
Write a custom numerical simulation loopLow-level order processing, where you handle the surrounding workflow.
Spend most of your time on indicators or performance metricsThe function benchmarks below, using shapes close to your own data.

The warm timings help estimate repeated research runs. First-call timings matter when you start a new script or worker process. Both are shown so you can judge the cost that applies to your workflow. The Rust engine page also demonstrates a strategy with callbacks that react to fills, while parallel execution and caching shows a large parameter sweep with measured memory usage.

Benchmarking your functions

The package includes a benchmark engine for its own functions. vbt.run_benchmarks generates seeded inputs of a chosen shape, checks backend outputs for agreement, and then times each one. The checks include returned values and registered in-place outputs, so timings are tied to the calculation being tested:

Compare Numba and Rust on six functions with 100,000 × 10 inputs
results = vbt.run_benchmarks(
    input_model="2d",
    backend_ids=("nb", "rs"),
    patterns=[
        "generic.rolling.rolling_mean",
        "generic.rolling.ewm_mean",
        "generic.base.crossed_above",
        "returns.sharpe_ratio",
        "returns.max_drawdown",
        "portfolio.from_signals.from_signals",
    ],
    rows=100_000,
    cols=10,
)
rows = {}
for r in results:
    ms = {backend: s * 1e3 for backend, s in r.runtime_by_backend.items()}
    rows[r.name] = ms | {"speedup": r.speedup}
pd.DataFrame(rows).T.round(2)  
                                                  nb      rs  speedup
generic.base.crossed_above                      4.45    5.67     0.78
generic.rolling.ewm_mean                        5.46    2.35     2.32
generic.rolling.rolling_mean                    3.81    3.96     0.96
portfolio.from_signals.from_signals[dense]    167.33  180.72     0.93
portfolio.from_signals.from_signals[grouped]  134.67  129.49     1.04
portfolio.from_signals.from_signals[sparse]    14.15   11.91     1.19
returns.max_drawdown                            2.73    0.94     2.92
returns.sharpe_ratio                            4.55    1.09     4.19

Rust was more than four times faster on the Sharpe ratio and nearly three times on maximum drawdown, and slightly slower on the rolling mean, crossovers, and the dense simulation. Results like these are why VBT does not assume one backend is always faster. The command-line version runs the same engine over whole families of functions, and the full matrix covers almost every function in the package, as the Benchmarks highlight below shows.

Generated matrix reports record the CPU, operating system, Python and package versions, Rust tools, and selected thread settings alongside the results. You can save the reports with your research and compare a new environment against the same workloads. Start with the functions you use most, then expand to the full matrix when you need broader coverage.

Measured backend selection

AutoBench uses those measurements. After you build a benchmark cache on your machine, setting vbt.settings.jitting["backends"]["auto_mode"] = "bench" uses stored timings and the input shape to choose an eligible backend, and "bench_mixed" also chooses between serial and parallel versions. Policies keep the choice conservative: a backend must beat the default by at least 5% by default, and the comparison can use the minimum or another statistic of the measured runs. The Compute backends page covers the other ways to choose a backend.

Measuring your code

vbt.Timer and vbt.MemTracer measure elapsed time and peak traced memory for a block of code, as the Resource management highlight below shows. vbt.timeit times a function and returns a readable duration, and the @vbt.with_timer and @vbt.with_memtracer decorators print the time or memory of every call to a function.

Measure the stages you actually run: preparing data, calculating signals, simulating orders, and building the statistics you want to keep. This shows where a faster backend, reused inputs, or smaller chunks would help. MemTracer uses Python's tracemalloc to track allocations during the block. Use whole-process memory measurements as well when planning how many workers fit in RAM.

Small inputs mostly measure fixed costs. On the machine that built this page, one vbt.PF.from_signals call took 5.3 ms on 365 bars and 6.4 ms on 100,000 bars: nearly all of it went to preparing inputs and building the portfolio, not to the simulation. A comparison on one year of daily data therefore says little about speed at scale. Many small backtests run faster as columns of one call, or through the lower-level functions the showdown above includes.

Benchmark what you run

Speed depends on the function, the input shape and memory layout, the number of cores, and whether code is already compiled. Time the calls that dominate your workload, on your data and hardware, after a warmup call.

Benchmarks

✅ VBT now ships with correctness-aware backend benchmarks for almost the entire codebase. The latest benchmark overview covers 748 functions, 768 variants, 1,976 kernels, 15,709 cases, and 59,600 benchmark runs across Numba, Rust, raw backend calls, AutoBench, and AutoBenchMixed. This makes performance inspectable: users can see which backend wins, at which input size, and whether parallel execution is worth it for their workload.

Benchmark one family from the terminal
python -m vectorbtpro.benchmarks.bench_engine_cli \
  --backend nb \
  --backend rs \
  --input-model 1d \
  --pattern returns
Run the same benchmark from Python
results = vbt.run_benchmarks(
    input_model="1d",
    backend_ids=("nb", "rs"),
    patterns=["returns"],
)
print(results)  
function,task_id,ndim,shape,nb_s,rs_s,speedup,auto_bench_backend,nb_elapsed_s,rs_elapsed_s
...
Benchmark median runtime ranks by backend and total element count Figure data (JSON)

✅ New profiling tools help you measure the execution time and memory usage of any code block 🧰

Profile getting the Sharpe ratio of a random portfolio
data = vbt.YFData.pull("BTC-USD")

with (
    vbt.Timer() as timer,
    vbt.MemTracer() as mem_tracer
):
    print(vbt.PF.from_random_signals(data.close, n=100, seed=42).sharpe_ratio)
1.0410760501518814
print(timer.elapsed())
74.15 milliseconds
print(mem_tracer.peak_usage())
459.7 kB

Copyright © 2021–2026 Oleg Polakow. All rights reserved.

Site content and documentation are provided for using and evaluating VectorBT PRO and for educational purposes. Any other use, including building or supporting competing products or services, requires prior written consent.