Filters#

A column’s storage pipeline, mirroring HDF5’s chunk filter pipeline. The filters and storage chapter covers usage and the hdf5plugin ecosystem.

class h5col.FilterPipeline(filters: Iterable[Any] = ())[source]#

An ordered, immutable sequence of Filter entries.

Accepts Filter instances and hdf5plugin filter objects, adapting the latter automatically.

Raises:

FilterError – If an entry is neither a Filter nor an hdf5plugin filter.

classmethod from_dcpl(dcpl: Any) → FilterPipeline[source]#

The pipeline recorded in a dataset-creation property list.

The counterpart to apply(), though not its inverse. What comes back is what HDF5 stored, which is not always what was handed over: a plugin may add client data of its own while the dataset is being created. Shuffle records the element size it was applied to, and Blosc2 records that plus the chunk’s size in bytes. For the same reason the result describes one particular dataset and is not a template for creating another.

Added in version 0.5.0.

Parameters:

dcpl – A dataset-creation property list, left unmodified. One with no filters gives an empty pipeline.

classmethod from_dataset(dataset: Any) → FilterPipeline[source]#

The pipeline a dataset was created with.

Added in version 0.5.0.

Parameters:

dataset – Any h5py dataset, not only a H5Col column. An unfiltered or contiguous one gives an empty pipeline.

apply(dcpl: Any) → None[source]#

Add every filter, in order, to an HDF5 dataset-creation property list.

Changed in version 0.5.0: A mandatory filter that cannot write in this process is refused here, rather than failing later inside HDF5 with a message that does not say which filter or what to do about it.

Parameters:

dcpl – A dataset-creation property list, modified in place. Declaration order is pipeline order, so filters are added exactly as listed.

Raises:

FilterError – If a filter this process cannot run is mandatory. An optional one is allowed through: HDF5 skips a filter it cannot apply, and the data is written without it.

to_h5py_kwargs() → dict[str, Any][source]#

Map the pipeline to h5py high-level create_dataset keyword arguments.

Deprecated since version 0.5.0: H5Col builds every column from a dataset-creation property list, so this is no longer part of writing a table. It stays for the moment as a convenience for code calling h5py’s create_dataset directly, and will be removed in a later release. Use apply() instead.

The builtin shuffle and fletcher32 filters map to their boolean keywords, and the single remaining compressor maps to compression and compression_opts.

What this mapping cannot express is a fair summary of why H5Col stopped using it: h5py’s keywords hold one compressor, in an order h5py chooses, with flags h5py chooses. apply() has none of those limits.

Raises FilterError if the pipeline needs more than one compressor filter, which the high-level API cannot express (combine them into a Blosc/Blosc2 meta-compressor, or drive the low-level DCPL via apply()).

class h5col.Filter(plugin_id: int, cd_values: tuple[int, ...] = (), optional: bool = False, name: str = '')[source]#

One entry of a filter pipeline.

Parameters:
  • plugin_id (int) – The registered HDF5 filter plugin identifier. A filter plugin is the piece of software that connects the HDF5 library to the code that actually filters a chunk’s bytes.

  • cd_values (tuple[int, ...]) – Client data (the filter’s unsigned-int parameters), in order.

  • optional (bool) – If True, HDF5 may skip the filter for a chunk it cannot process instead of failing the write (the H5Z_FLAG_OPTIONAL flag).

  • name (str) – Human-readable label (informational only).

property flags: int#

The HDF5 filter flags for this entry.

property available: bool#

Whether this plugin is registered with the HDF5 library right now.

A plugin has to be installed and loadable for its filter to run. One that is not stays absent for this process only, so a column written with it elsewhere is still perfectly good — it just cannot be filtered or unfiltered here.

Added in version 0.5.0.

property can_encode: bool#

Whether this filter can transform data on its way into a file.

A plugin may ship able to read what it wrote elsewhere but not to write anything itself; szip is the one usually built that way.

Added in version 0.5.0.

property can_decode: bool#

Whether this filter can restore data on its way out of a file.

Added in version 0.5.0.

h5col.Deflate(level: int = 4) → Filter[source]#

The built-in deflate compression filter (HDF5’s H5Z_FILTER_DEFLATE).

Parameters:

level – Compression level, 0–9 (default 4).

Raises:

FilterError – If level is outside 0–9.

h5col.Shuffle() → Filter[source]#

The built-in byte-shuffle filter.

h5col.Fletcher32() → Filter[source]#

The built-in Fletcher-32 checksum filter.

h5col.from_hdf5plugin(obj: Any) → Filter[source]#

Adapt an hdf5plugin filter instance to a Filter.

hdf5plugin filter objects behave like mappings of h5py create_dataset keyword arguments (compression = filter id, compression_opts = client data) and expose a filter_id attribute.

Parameters:

obj – An hdf5plugin filter instance such as hdf5plugin.Zstd(). Anything else that converts to such a mapping and carries a filter id also works; the class itself is not required.

Raises:

FilterError – If obj cannot be interpreted as an hdf5plugin filter or carries no filter id.