Filters and storage#
Column-oriented layouts and compression are natural partners. A column is a long run of same-typed, often similar values, which is exactly what compressors like. In H5Col each column dataset carries its own HDF5 filter pipeline and its own chunk shape, so every column can be stored the way its data deserves.
Filter pipelines#
A FilterPipeline is an ordered list of filter plugins,
representing an HDF5 dataset’s own chunk filter pipeline. On write, each chunk
passes through the plugins in order; on read, HDF5 reverses them transparently.
from h5col import ColumnSpec, Deflate, FilterPipeline, Fletcher32, Shuffle
ColumnSpec(
name="fare_amount",
dtype="float64",
filters=FilterPipeline([Shuffle(), Deflate(5)]),
)
Shuffle() reorders bytes so same-significance bytes group together
which typically improves compression ratio of numeric data.
Deflate() is zlib/gzip compression at the given level.
Fletcher32() appends a checksum to each chunk, so bit rot is
detected at read time. If using it, place it in the pipeline alongside the
others (recommendation is to be the last).
The pattern Shuffle() then Deflate(...), or shuffle then any
general-purpose compressor, is the sensible default for numeric columns.
Ordering is yours to decide. Every column, whatever its datatype, is created
through HDF5’s dataset-creation property list, so a pipeline is stored exactly
as written: the filter plugins stay in the order you gave them, each keeps its
optional flag, and a pipeline may hold up to 32 plugins.
H5col does not rearrange a pipeline or check that its order makes sense. Some filters do have a natural place — a bit-rounding filter belongs before a shuffle, and a shuffle before a compressor, because each prepares the bytes the next one works on. Which order is right depends on the filter and on the data, so the decision is yours.
A filter can only be applied by the plugin that implements it, and a plugin has
to be installed to run. If a column declares a filter whose plugin is missing
here, creating it fails with a message naming the plugin. Marking that filter
optional instead tells HDF5 to skip it and write the column without it.
Naming a filter#
Every filter plugin has an identifier, and The HDF Group’s plugin registry
gives each one a canonical name. h5col knows those names, so a plugin can be
asked for by name rather than by number:
from h5col import from_name, plugin_name
FilterPipeline([Shuffle(), from_name("zstd", 5)])
FilterPipeline(["shuffle", "zstd"]) # names alone, plugin defaults
plugin_name(32015) # 'zstd'
Names are lower case and matched exactly, as the registry specifies. One it
does not know answers plugin_<id>, which is still something to print and to
search for.
Reading a pipeline back#
filters() gives the pipeline a column was created with:
>>> table["temp_k"].filters
<h5col.FilterPipeline shuffle(4) | deflate(6)>
A list column is several datasets rather than one, so its filters answers
with a pipeline per dataset, keyed by path within the column group — OFFSETS,
VALUES, MASK, and for strings VALUES/OFFSETS and VALUES/CHARS.
from_dataset() does the same for any h5py dataset,
whether or not it belongs to a H5Col table.
One thing to expect: reported plugin configuration is not always what was its
input. A plugin may add client data of its own while processing the dataset’s
creation properties. Shuffle records the size of the element it was applied to,
so a Shuffle() declared with no parameters at all reads back as shuffle(4)
on a 4-byte column. Blosc2 records the element size and the chunk’s size in
bytes. What is read back describes that one dataset, it is not a template for
creating another.
The wider filter ecosystem#
HDF5 filters are available via registered filter plugins, and the
hdf5plugin package (a regular
dependency of h5col) provides the widely used modern ones. Its filter
objects drop straight into a pipeline:
import hdf5plugin
ColumnSpec(
name="total_amount",
dtype="float64",
filters=FilterPipeline([hdf5plugin.Zstd(clevel=5)]),
)
(from_hdf5plugin() is the explicit adapter behind that coercion.)
Beyond hdf5plugin’s catalog, any registered HDF5 filter, such as
domain-specific or lossy compressors, can be named directly by its plugin id
with Filter:
Filter(plugin_id=32015, cd_values=(5,), name="zstd")
Chunks written with a plugin filter need that plugin present at read time too.
Within Python this is handled by importing h5col which imports hdf5plugin.
Readers outside Python need the corresponding HDF5 plugin
binaries. Choose filter plugins with
your consumers in mind.
Chunking#
Chunks are the unit of I/O, of filtering, and of
chunk-level index pruning. Every column accepts a chunks=
row count in its spec.
When chunks is omitted, this implementation picks a chunk size aimed at large
tables: a few mebibytes per chunk (2–8 MiB, scaled to the dataset chunk cache),
so scans stream efficiently. Two situations call for an explicit value:
Small tables. A small table with an auto-sized, million-element chunk allocates the whole chunk on disk regardless — a 265-row lookup table can cost megabytes. Set
chunksnear the expected row count. In the taxi example, doing so on the zones table shrank the file from about 5.6 MB to 1.4 MB by releasing the over-allocated chunk space.Selective queries over big columns. A
CHUNK_MINMAXindex prunes whole chunks, so its resolution is the chunk size; the taxi trips table chunks its columns at 4,096 rows partly to make that pruning effective.
Seeing what storage decisions cost#
Storage questions deserve measurements. storage gives the
bytes a column’s values would occupy unfiltered, the bytes actually used in the
file, and what the filter pipeline bought:
>>> table["fare_amount"].storage
<h5col.Storage 200.0 kB -> 41.0 kB (4.9x)>
>>> for name in table.column_names:
... print(f"{name:24s} {table[name].storage!r}")
A list column is several datasets rather than one, and its storage adds them
together. dataset_storage() and group_storage() do the
same for any h5py dataset or group.
ratio is None, and prints as —, when there is nothing to divide. An empty
column and a column left entirely at its fill value both leave HDF5 with no
chunks to allocate, and neither is infinite compression.
Measuring the stored size means walking the chunk index. That is quick on a local file and a round trip or several on a remote one, so it is not something to ask for in a loop over a wide table held in object storage unless reading from a cloud optimized file.
On the 25,000-row taxi sample, shuffle-plus-DEFLATE pipelines compress the trips table about 4.9× overall. Numbers like that are dataset-dependent — measure your own.