# Records and mapped arrays (/features/analytics/records-and-mapped-arrays)

Orders, trades, drawdowns, and pattern matches are events, not time series: most bars have none, and
some have several. VBT stores such sparse events and intervals as records with a column map, so you
can filter, reduce, regroup, and plot them across thousands of columns without building a dense
array for every field.

```python title="Find the order with the worst fill among 20,000 orders in 1,000 columns"
>>> close = vbt.GBMOHLCData.pull("S", start="2024-01-01", end="2025-01-01", seed=42).close
>>> np.random.seed(42)
>>> slippage = pd.DataFrame(
...     np.random.uniform(0, 0.005, size=(len(close), 1000)),
...     index=close.index,
... )
>>> pf = vbt.PF.from_random_signals(close, n=10, seed=42, slippage=slippage)  # (1)
>>> orders = pf.orders
>>> print(orders.count().sum())
20000

>>> fill_vs_close = orders.price.values / close.values[orders.idx.values] - 1
>>> cost = orders.map_array(  # (2)
...     np.where(orders.side.values == 0, fill_vs_close, -fill_vs_close)
... )
>>> cost.max().describe().round(4)  # (3)
count    1000.0000
mean        0.0048
std         0.0002
min         0.0035
25%         0.0047
50%         0.0048
75%         0.0049
max         0.0050
Name: max, dtype: float64

>>> readable = cost.to_readable()
>>> readable.loc[readable["Value"].idxmax()]
Id                                1
Column                          165
Index     2024-03-12 00:00:00+00:00
Value                      0.004999
Name: 3301, dtype: object
```

1.  One synthetic price, 1,000 columns with random slippage, and 10 random entries and exits per
    column.
2.  One value per order: how much worse than the close it filled.
3.  Reduce the 20,000 values to the worst fill of each column, then summarize the 1,000 results.

Twenty thousand orders take twenty thousand rows, not 366 bars times 1,000 columns times every
field. The worst fill was the second order in column 165, on March 12, 2024.

## Records \[#records]

A record array is a structured NumPy array, one row per event, with fields such as the column, the
bar, the price, and the size. Records classes wrap it with the time index and column labels of the
data, so the same object knows both the raw positions and their dates. Orders, logs, trades,
positions, drawdowns, and pattern matches are all records.

Ranges are records with a start and an end. Any boolean mask becomes ranges, which turns periods
into objects you can measure:

```python title="Turn overbought periods into ranges"
>>> btc = vbt.YFData.pull("BTC-USD", start="2020-01-01", end="2025-01-01")
>>> rsi = btc.run("rsi", window=14).rsi.rename("RSI")
>>> overbought = vbt.Ranges.from_array(rsi > 70)
>>> print(overbought.count(), overbought.coverage.round(3))  # (1)
42 0.128

>>> longest = overbought.apply_mask(overbought.duration.top_n_mask(3))
>>> longest.readable[["Start Index", "End Index", "Status"]]
                Start Index                 End Index  Status
0 2023-01-11 00:00:00+00:00 2023-01-30 00:00:00+00:00  Closed
1 2023-10-20 00:00:00+00:00 2023-11-14 00:00:00+00:00  Closed
2 2024-11-06 00:00:00+00:00 2024-11-25 00:00:00+00:00  Closed
>>> longest.duration.values
array([19, 25, 19])
```

1.  42 separate overbought periods covering 12.8% of the bars.

Records can also be built from your own data. A DataFrame or a list of dictionaries with dates and
symbols becomes typed records, which is how
[index records](/features/backtesting/backtesting-engine/#index-records) pass sparse orders into a
simulation.

## Mapped arrays \[#mapped-arrays]

A mapped array holds one value per record together with the record's column and bar. Every field of
a records object is available as one, for example `pf.trades.pnl` or `pf.drawdowns.duration`, and
`map_array` attaches any values you compute. Mapped arrays reduce per column or group with `mean`,
`max`, `idxmax`, `count`, or any compiled function, filter with masks such as `top_n_mask`, and
convert to pandas with `to_pd` or `to_readable` when you need a table. `to_pd` places each value at
its record's bar. When several records of one column share a bar, it keeps the latest and warns.
Pass `reduce_func_nb`, such as `"sum"`, to combine them, `repeat_index=True` to keep all of them, or
`ignore_index=True` to stack the values without dates.

This lets you ask more specific questions of a backtest: how long did losing trades last when they
barely moved into profit, or which orders had the largest costs? Filter the records or their mapped
values, then run the same reductions on the selected events. The results keep their asset and
parameter labels. [Trade analytics](/features/analytics/trade-analytics/) shows this workflow for
long and short trades and their price excursions.

## Indexing and regrouping \[#indexing-and-regrouping]

Records stay sorted by column, and a column map remembers where each column's records start.
Selecting a column or a group, as in `pf.trades["BTC-USD"]`, therefore does not scan all records.
Grouping can change after the fact: the same trade records report per asset with `group_by=False` or
per strategy with a different grouping. For a chronological view across columns, sort the readable
table by its index column.

For example, report three assets separately, then combine A and B into one research group:

```python title="Regroup trade profits without rerunning the backtest"
>>> group_close = pd.DataFrame({"A": [10.0, 11.0, 12.0], "B": [20.0, 18.0, 19.0], "C": [30.0, 33.0, 36.0]}, index=pd.date_range("2026-01-01", periods=3))
>>> group_pf = vbt.PF.from_orders(group_close, size=np.array([1, 0, -1])[:, None])
>>> trade_pnl = group_pf.trades.status_closed.pnl
>>> trade_pnl.sum()
A    2.0
B   -1.0
C    6.0
Name: sum, dtype: float64

>>> trade_pnl.sum(group_by=["Group 1", "Group 1", "Group 2"])
group
Group 1    1.0
Group 2    6.0
Name: sum, dtype: float64
```

The same records now answer a different reporting question. You can group by sectors, strategy
families, or another label for each column. To have assets share capital while the simulation runs,
use [shared cash](/features/backtesting/portfolio-accounting/#shared-cash).

## Custom record classes \[#custom-record-classes]

Your own events can have their own class. Define a NumPy dtype with the fields you need, attach
titles and value mappings so readable tables show names instead of codes, and subclass `vbt.Records`
or `vbt.Ranges` to get filters, mapped fields, statistics, and plots for free. Records save to disk
and load back like any other VBT object. Inside Numba code, read record fields with brackets, such
as `record["price"]`, which works whether compilation is on or off.

## OHLC-native classes \[#ohlc-native-classes]

New in 1.3.0

✅ Previously, OHLC data was used for simulation, but only the close price was analyzed. Now, most
classes let you track all OHLC data for more accurate quantitative and qualitative analysis.

```python title="Plot trades of a random portfolio"
>>> data = vbt.YFData.pull("BTC-USD", start="2020-01", end="2020-03")
>>> pf = vbt.PF.from_random_signals(
...     open=data.open,
...     high=data.high,
...     low=data.low,
...     close=data.close,
...     n=10,
...     seed=42
... )
>>> pf.trades.plot().show()
```

BTC-USD OHLC with random portfolio trade entries and exits. [Figure data (JSON)](/assets/figures/features/analysis/ohlc-native-classes.4ffa9efd631b.json)


## Related pages

*   [Multidimensional research](/features/tooling/multidimensional-research/): Run assets, parameters, and strategies as labeled columns, then index and stack them
*   [Backtesting engine](/features/backtesting/backtesting-engine/): Simulate orders, signals, and callbacks across many assets and parameters at once
*   [Trade analytics](/features/analytics/trade-analytics/): Analyze trades and positions with MAE, MFE, edge ratio, profit factor, and plots
*   [Time-series operations](/features/tooling/time-series-operations/): Run compiled rolling windows, reductions, and date logic on labeled pandas data