Examples#

The examples are executable Jupyter notebooks, committed with their outputs, so every page below shows real results. They live in the repository’s examples/ directory; run them yourself with the examples pixi environment (pixi run -e examples jupyter lab).

Each notebook is self-contained. Most write their files to the system temporary directory; the taxi notebook builds its file from a small committed sample of the public NYC TLC trip data, so it too runs offline. The Arrow notebook needs the optional pyarrow dependency, which the examples environment already has.

Notebook

Covers

Quickstart

Create → append → read → validate, and reopen from disk.

Column types

Fixed-length strings (no silent truncation), categoricals, booleans, valid ranges, and missing values.

List columns

list<float>, list<str>, nested lists, null versus empty, and the offsets layout.

Filters and storage

Per-column filter pipelines, hdf5plugin codecs, compression ratios, and the raw HDF5 layout.

JSON logs

Modeling semi-structured log records with columns and list columns.

NYC taxi trips

A real dataset end to end: categoricals, both missing-value styles, a datetime codec, filters, and indexed queries with explain().

Exporting to Arrow

to_arrow(): real nulls instead of fill values, categoricals as dictionaries, list columns intact, column attributes carried as field metadata, and the hop to pandas and Parquet.

Reading part of a table

Subscript and read_rows on a 200,000-row table, measuring what a range actually reads: chunk-coalesced reads for scalar columns, spanning ranges for list columns, and the column[...] versus column.dataset[...] trap.