Column datatypes#

Helpers for the datatypes HDF5 does not hand to NumPy directly: the fixed-length UTF-8 string, the H5Col boolean enumeration, and the opaque column of raw bytes. The column datatypes chapter explains when to reach for each.

Fixed-length strings#

class h5col.FixedString(nbytes: int, encoding: str = 'utf-8')[source]#

A fixed-length HDF5 string datatype of nbytes bytes.

Parameters:
  • nbytes (int) – Storage width in bytes (> 0).

  • encoding (str) – "utf-8" (default) or "ascii".

Raises:

SchemaError – If nbytes is not a positive integer, or encoding is unsupported.

property dtype: dtype#

The NumPy/h5py dtype for creating a dataset or attribute of this type.

encode_scalar(value: object, *, index: int | None = None) → bytes[source]#

Encode a single value to bytes, enforcing the byte budget.

Parameters:
  • value – bytes is stored as given; anything else is passed through str() and encoded with this type’s encoding.

  • index – Position of the value in the array being written, carried into the error message so an oversized row can be located. None when the value did not come from an array.

Raises:

OversizedStringError – If the encoded form exceeds nbytes. Nothing is truncated.

encode(values: Any) → NDArray[bytes_][source]#

Encode an array-like of strings to a |S{nbytes} array.

Parameters:

values – Any array-like of values encode_scalar() accepts. The shape is preserved.

Raises:

OversizedStringError – On the first value whose encoding exceeds nbytes, carrying its position. No value is ever truncated.

decode_scalar(raw: bytes | bytearray | bytes_) → str[source]#

Decode a single stored value to str (trailing NULs stripped).

Parameters:

raw – One stored value. HDF5 pads a fixed-length string to its full width with NUL bytes, which are stripped before decoding.

decode(values: Any) → NDArray[Any][source]#

Decode an array-like of stored bytes to a NumPy string array.

The result carries decoded_string_dtype(), which keeps the text in one compact arena instead of allocating a Python str object per element the way a dtype=object array does.

NumPy’s own S → StringDType cast replaces the per-element Python loop this used to run: it strips the trailing NULs HDF5 pads with and takes the bytes as UTF-8, which also covers ascii columns since ASCII is a subset.

The cast copies eagerly but validates lazily. Bytes that are not valid UTF-8 — only reachable from a non-conformant producer, since encode() validates on write — therefore raise UnicodeDecodeError when the offending value is read out of the array, not when the column is read.

Parameters:

values – An array-like of stored bytes, as read from a fixed-length string dataset. The shape is preserved.

classmethod from_dtype(dtype: Any) → FixedString[source]#

Build a FixedString from an existing fixed-length string dtype.

Parameters:

dtype – An h5py string dtype, normally taken from an existing dataset. Its length and encoding become the new instance’s nbytes and encoding; an absent encoding defaults to UTF-8.

Raises:

SchemaError – If dtype is not an HDF5 string dtype, or is a variable-length (not fixed-length) string dtype.

static is_fixed_string(dtype: Any) → bool[source]#

Return True if dtype is a fixed-length HDF5 string dtype.

Parameters:

dtype – Any dtype. One that is not a string dtype at all, and a variable-length string dtype, both answer False.

h5col.ascii_token_dtype(value: str) → dtype[source]#

Return a fixed-length ASCII string dtype sized to hold value plus a NUL.

Used for H5Col reserved-token attributes (CLASS, VERSION, KIND), whose values are ASCII and are stored null-terminated/null-padded.

Parameters:

value – The token the dtype has to hold. Only its encoded length is used, so any token of the same length gives the same dtype.

Booleans#

h5col.bool_dtype() → dtype[source]#

Return the H5Col boolean dtype (enum over little-endian signed int8).

On little-endian platforms the <i1 base is byte-identical to H5T_STD_I8LE, which is what H5Col mandates.

h5col.is_bool_dtype(dtype: Any) → bool[source]#

Return True if dtype is an acceptable H5Col boolean datatype.

Applies the consumer-lenient rule: any one-byte integer enumeration, of either signedness, whose members are exactly FALSE = 0 and TRUE = 1. Also accepts NumPy bool: h5py normalizes its FALSE/TRUE-over-int8 enum (which is the H5Col boolean datatype on disk) back to bool on read.

Parameters:

dtype – Anything numpy.dtype() accepts, including a dtype carrying h5py enumeration metadata. A dtype that is not an enumeration at all is not an error; it simply answers False.

h5col.encode_bool(values: Any) → NDArray[int8][source]#

Encode a boolean/0-1 array-like to an int8 array of codes.

Parameters:

values – A sequence or array of Python bools, NumPy bools, or integers that are all 0 or 1. Any other integer is rejected rather than coerced, and floats such as 1.5 are rejected rather than truncated.

Raises:

SchemaError – If an input value is neither 0 nor 1 (an H5Col boolean column may hold only those two codes).

h5col.decode_bool(values: Any) → NDArray[bool][source]#

Decode stored integer codes to a NumPy boolean array.

Every code must be 0 (FALSE) or 1 (TRUE). A code outside that domain is a non-conformant boolean value; H5Col forbids interpreting it as either FALSE or TRUE, so this raises instead of silently coercing (NumPy maps every nonzero code to True).

Parameters:

values – The stored codes, as read from a boolean column’s dataset.

Raises:

ConformanceError – If any code is neither 0 nor 1.

Opaque bytes#

h5col.is_opaque_dtype(dtype: Any) → bool[source]#

True for a plain opaque datatype — raw bytes of a fixed width.

V is NumPy’s void kind, which covers plain raw bytes, structured datatypes and fixed-shape sub-arrays alike. h5py reads three different HDF5 datatypes through it — H5T_OPAQUE, H5T_COMPOUND and H5T_ARRAY — so a column’s kind alone does not say which arrived. Only the last two carry a structure, and those two structural fields are what separates them.

Added in version 0.4.0.

Parameters:

dtype – Anything numpy.dtype() accepts.

h5col.opaque_fill_bytes(nbytes: int) → bytes[source]#

H5Col’s recommended fill value for an opaque column nbytes wide.

Since no byte string can be reserved by being out of range, the next best thing is one that is vanishingly unlikely to occur: the ASCII FILL marker, followed by rising byte values that wrap through zero.

The rising tail is the part that earns its keep. A fill of all zeros or all 0xFF would collide constantly — that is what zero padding, erased flash and uninitialized memory leave behind — while a counting sequence is something real data almost never is. An eight-byte column gets 46 49 4c 4c 01 02 03 04, which reads as FILL.... in a hex dump.

A column narrower than the marker gets as much of it as fits. Note that a one-byte opaque column has 256 possible values and this claims one of them, which is a real risk rather than a negligible one; such a column is better given a fill value chosen against its own data.

Added in version 0.4.0.

Parameters:

nbytes – The column’s fixed width. The result is exactly this long.