Filters#
A column’s storage pipeline, mirroring HDF5’s chunk filter pipeline. The
filters and storage chapter covers usage and the
hdf5plugin ecosystem.
- class h5col.FilterPipeline(filters: Iterable[Any] = ())[source]#
An ordered, immutable sequence of
Filterentries.Accepts
Filterinstances andhdf5pluginfilter objects, adapting the latter automatically.- Raises:
FilterError – If an entry is neither a
Filternor anhdf5pluginfilter.
- classmethod from_dcpl(dcpl: Any) FilterPipeline[source]#
The pipeline recorded in a dataset-creation property list.
The counterpart to
apply(), though not its inverse. What comes back is what HDF5 stored, which is not always what was handed over: a plugin may add client data of its own while the dataset is being created. Shuffle records the element size it was applied to, and Blosc2 records that plus the chunk’s size in bytes. For the same reason the result describes one particular dataset and is not a template for creating another.Added in version 0.5.0.
- Parameters:
dcpl – A dataset-creation property list, left unmodified. One with no filters gives an empty pipeline.
- classmethod from_dataset(dataset: Any) FilterPipeline[source]#
The pipeline a dataset was created with.
Added in version 0.5.0.
- Parameters:
dataset – Any h5py dataset, not only a H5Col column. An unfiltered or contiguous one gives an empty pipeline.
- apply(dcpl: Any) None[source]#
Add every filter, in order, to an HDF5 dataset-creation property list.
Changed in version 0.5.0: A mandatory filter that cannot write in this process is refused here, rather than failing later inside HDF5 with a message that does not say which filter or what to do about it.
- Parameters:
dcpl – A dataset-creation property list, modified in place. Declaration order is pipeline order, so filters are added exactly as listed.
- Raises:
FilterError – If a filter this process cannot run is mandatory. An optional one is allowed through: HDF5 skips a filter it cannot apply, and the data is written without it.
- to_h5py_kwargs() dict[str, Any][source]#
Map the pipeline to h5py high-level
create_datasetkeyword arguments.Deprecated since version 0.5.0: H5Col builds every column from a dataset-creation property list, so this is no longer part of writing a table. It stays for the moment as a convenience for code calling h5py’s
create_datasetdirectly, and will be removed in a later release. Useapply()instead.The builtin shuffle and fletcher32 filters map to their boolean keywords, and the single remaining compressor maps to
compressionandcompression_opts.What this mapping cannot express is a fair summary of why H5Col stopped using it: h5py’s keywords hold one compressor, in an order h5py chooses, with flags h5py chooses.
apply()has none of those limits.Raises
FilterErrorif the pipeline needs more than one compressor filter, which the high-level API cannot express (combine them into a Blosc/Blosc2 meta-compressor, or drive the low-level DCPL viaapply()).
- class h5col.Filter(plugin_id: int, cd_values: tuple[int, ...] = (), optional: bool = False, name: str = '')[source]#
One entry of a filter pipeline.
- Parameters:
plugin_id (int) – The registered HDF5 filter plugin identifier. A filter plugin is the piece of software that connects the HDF5 library to the code that actually filters a chunk’s bytes.
cd_values (tuple[int, ...]) – Client data (the filter’s unsigned-int parameters), in order.
optional (bool) – If True, HDF5 may skip the filter for a chunk it cannot process instead of failing the write (the
H5Z_FLAG_OPTIONALflag).name (str) – Human-readable label (informational only).
- property available: bool#
Whether this plugin is registered with the HDF5 library right now.
A plugin has to be installed and loadable for its filter to run. One that is not stays absent for this process only, so a column written with it elsewhere is still perfectly good — it just cannot be filtered or unfiltered here.
Added in version 0.5.0.
- h5col.Deflate(level: int = 4) Filter[source]#
The built-in deflate compression filter (HDF5’s
H5Z_FILTER_DEFLATE).- Parameters:
level – Compression level,
0–9(default4).- Raises:
FilterError – If level is outside
0–9.
- h5col.from_hdf5plugin(obj: Any) Filter[source]#
Adapt an
hdf5pluginfilter instance to aFilter.hdf5pluginfilter objects behave like mappings of h5py create_dataset keyword arguments (compression= filter id,compression_opts= client data) and expose afilter_idattribute.- Parameters:
obj – An
hdf5pluginfilter instance such ashdf5plugin.Zstd(). Anything else that converts to such a mapping and carries a filter id also works; the class itself is not required.- Raises:
FilterError – If obj cannot be interpreted as an
hdf5pluginfilter or carries no filter id.