Filters and storage#

Column-oriented layouts and compression are natural partners. A column is a long run of same-typed, often similar values, which is exactly what compressors like. In H5Col each column dataset carries its own HDF5 filter pipeline and its own chunk shape, so every column can be stored the way its data deserves.

Filter pipelines#

A FilterPipeline is an ordered list of filter plugins, representing an HDF5 dataset’s own chunk filter pipeline. On write, each chunk passes through the plugins in order; on read, HDF5 reverses them transparently.

from h5col import ColumnSpec, Deflate, FilterPipeline, Fletcher32, Shuffle

ColumnSpec(
    name="fare_amount",
    dtype="float64",
    filters=FilterPipeline([Shuffle(), Deflate(5)]),
)

Shuffle() reorders bytes so same-significance bytes group together which typically improves compression ratio of numeric data. Deflate() is zlib/gzip compression at the given level. Fletcher32() appends a checksum to each chunk, so bit rot is detected at read time. If using it, place it in the pipeline alongside the others (recommendation is to be the last).

The pattern Shuffle() then Deflate(...), or shuffle then any general-purpose compressor, is the sensible default for numeric columns.

Ordering is yours to decide. Every column, whatever its datatype, is created through HDF5’s dataset-creation property list, so a pipeline is stored exactly as written: the filter plugins stay in the order you gave them, each keeps its optional flag, and a pipeline may hold up to 32 plugins.

H5col does not rearrange a pipeline or check that its order makes sense. Some filters do have a natural place — a bit-rounding filter belongs before a shuffle, and a shuffle before a compressor, because each prepares the bytes the next one works on. Which order is right depends on the filter and on the data, so the decision is yours.

A filter can only be applied by the plugin that implements it, and a plugin has to be installed to run. If a column declares a filter whose plugin is missing here, creating it fails with a message naming the plugin. Marking that filter optional instead tells HDF5 to skip it and write the column without it.

Naming a filter#

Every filter plugin has an identifier, and The HDF Group’s plugin registry gives each one a canonical name. h5col knows those names, so a plugin can be asked for by name rather than by number:

from h5col import from_name, plugin_name

FilterPipeline([Shuffle(), from_name("zstd", 5)])

FilterPipeline(["shuffle", "zstd"])   # names alone, plugin defaults

plugin_name(32015)                    # 'zstd'

Names are lower case and matched exactly, as the registry specifies. One it does not know answers plugin_<id>, which is still something to print and to search for.

Reading a pipeline back#

filters() gives the pipeline a column was created with:

>>> table["temp_k"].filters
<h5col.FilterPipeline shuffle(4) | deflate(6)>

A list column is several datasets rather than one, so its filters answers with a pipeline per dataset, keyed by path within the column group — OFFSETS, VALUES, MASK, and for strings VALUES/OFFSETS and VALUES/CHARS.

from_dataset() does the same for any h5py dataset, whether or not it belongs to a H5Col table.

One thing to expect: reported plugin configuration is not always what was its input. A plugin may add client data of its own while processing the dataset’s creation properties. Shuffle records the size of the element it was applied to, so a Shuffle() declared with no parameters at all reads back as shuffle(4) on a 4-byte column. Blosc2 records the element size and the chunk’s size in bytes. What is read back describes that one dataset, it is not a template for creating another.

The wider filter ecosystem#

HDF5 filters are available via registered filter plugins, and the hdf5plugin package (a regular dependency of h5col) provides the widely used modern ones. Its filter objects drop straight into a pipeline:

import hdf5plugin

ColumnSpec(
    name="total_amount",
    dtype="float64",
    filters=FilterPipeline([hdf5plugin.Zstd(clevel=5)]),
)

(from_hdf5plugin() is the explicit adapter behind that coercion.) Beyond hdf5plugin’s catalog, any registered HDF5 filter, such as domain-specific or lossy compressors, can be named directly by its plugin id with Filter:

Filter(plugin_id=32015, cd_values=(5,), name="zstd")

Chunks written with a plugin filter need that plugin present at read time too. Within Python this is handled by importing h5col which imports hdf5plugin. Readers outside Python need the corresponding HDF5 plugin binaries. Choose filter plugins with your consumers in mind.

Chunking#

Chunks are the unit of I/O, of filtering, and of chunk-level index pruning. Every column accepts a chunks= row count in its spec.

When chunks is omitted, this implementation picks a chunk size aimed at large tables: a few mebibytes per chunk (2–8 MiB, scaled to the dataset chunk cache), so scans stream efficiently. Two situations call for an explicit value:

  • Small tables. A small table with an auto-sized, million-element chunk allocates the whole chunk on disk regardless — a 265-row lookup table can cost megabytes. Set chunks near the expected row count. In the taxi example, doing so on the zones table shrank the file from about 5.6 MB to 1.4 MB by releasing the over-allocated chunk space.

  • Selective queries over big columns. A CHUNK_MINMAX index prunes whole chunks, so its resolution is the chunk size; the taxi trips table chunks its columns at 4,096 rows partly to make that pruning effective.

Seeing what storage decisions cost#

Storage questions deserve measurements. storage gives the bytes a column’s values would occupy unfiltered, the bytes actually used in the file, and what the filter pipeline bought:

>>> table["fare_amount"].storage
<h5col.Storage 200.0 kB -> 41.0 kB (4.9x)>

>>> for name in table.column_names:
...     print(f"{name:24s} {table[name].storage!r}")

A list column is several datasets rather than one, and its storage adds them together. dataset_storage() and group_storage() do the same for any h5py dataset or group.

ratio is None, and prints as —, when there is nothing to divide. An empty column and a column left entirely at its fill value both leave HDF5 with no chunks to allocate, and neither is infinite compression.

Measuring the stored size means walking the chunk index. That is quick on a local file and a round trip or several on a remote one, so it is not something to ask for in a loop over a wide table held in object storage unless reading from a cloud optimized file.

On the 25,000-row taxi sample, shuffle-plus-DEFLATE pipelines compress the trips table about 4.9× overall. Numbers like that are dataset-dependent — measure your own.