Low-level modules#

Seven submodules are exported as public, stable entry points below the class-based API. They operate directly on h5py objects — most take a table group or a dataset rather than a Table — and exist for tools that need the convention’s mechanics without the wrappers: validators, migration scripts, or readers in constrained environments. Most users never need them.

h5col.arrow#

Builds the Arrow export behind Table.to_arrow and Column.to_arrow. Needs the optional pyarrow dependency; nothing else in the package imports it.

Export H5Col columns and tables as Apache Arrow.

Arrow is the one representation that carries the whole H5Col data model without loss. NumPy has no type for three of the things H5Col stores: a null distinct from the value that marks it, a dictionary of codes and labels, and (from the list-column work) ragged values with nulls at every level. Arrow has all three natively, and its memory layout is close enough to H5Col’s that most of the conversion is handing buffers over rather than rewriting them.

pyarrow is an optional dependency — install h5col[arrow]. Nothing else in the package imports it.

Type mapping#

H5Col column

Arrow type

numeric

the matching primitive, data buffer shared

boolean

bool_

fixed-length string

large_string

opaque

fixed_size_binary[n], data buffer shared

categorical

dictionary<indices=<code dtype>, values=...>

list column

large_list<...>, buffers shared

Missing rows become real Arrow nulls in every case, so an Arrow consumer never sees the fill value.

h5col.arrow.METADATA_PREFIX = 'h5col.'

Prefix for the Arrow field-metadata keys carrying HDF5 column attributes, so they cannot collide with a consumer’s own metadata.

h5col.arrow.require_pyarrow() → Any[source]

Import and return pyarrow, or explain how to get it.

Raises:

ModuleNotFoundError – If pyarrow is not installed.

h5col.arrow.string_array(raw: NDArray[Any], nbytes: int, mask: NDArray[bool]) → Any[source]

Build an Arrow large_string from a fixed-width string column block.

A fixed-width column has no offsets to lend, so unlike a list column this is a real conversion. It is still done with array arithmetic rather than a Python loop: the trailing NULs HDF5 pads with are counted per row, the lengths accumulate into the offsets buffer, and the payload bytes are taken in one masked gather.

Parameters:
  • raw – The stored fixed-width bytes, one row per element.

  • nbytes – The column’s fixed width, i.e. the stored size of every row.

  • mask – True where the row is missing. Masked rows become Arrow nulls and contribute no payload bytes.

h5col.arrow.column_array(col: Column, rows: Any = None) → Any[source]

Convert one scalar column (or rows of it) to an Arrow array.

Parameters:
  • col – The scalar column to convert. List columns go through list_array() instead.

  • rows – Which rows to convert, in any form read_rows() accepts. None converts the whole column.

h5col.arrow.column_metadata(col: Column) → dict[str, str][source]

The column’s HDF5 attributes, as Arrow field metadata.

Arrow metadata is a flat string-to-string map, so numeric attributes are rendered with str. The keys are prefixed with h5col. and survive a Parquet round trip, which is what makes them worth carrying at all.

Parameters:

col – The scalar column whose attributes are collected. Attributes that are unset are simply left out.

h5col.arrow.list_column_metadata(col: Any) → dict[str, str][source]

A list column’s HDF5 attributes, as Arrow field metadata.

Parameters:

col – The list column whose attributes are collected. Attributes that are unset are simply left out.

h5col.arrow.table_arrow(table: Any, columns: Any = None, rows: Any = None) → Any[source]

Convert a table (or rows of it) to an Arrow table.

Parameters:
  • table – The table to convert.

  • columns – Names to convert, in the order given. None converts every column.

  • rows – Which rows to convert, as a slice, a sequence of positions, or a boolean mask. None converts every row.

Raises:

KeyError – If a requested column name is not a column of the table.

h5col.arrow.list_array(group: Any, count: int) → Any[source]

One level of a list column as an Arrow large_list.

Recurses through nesting, so an inner null — which no top-level mask can express — survives at whatever depth it was written.

Parameters:
  • group – A list column group, or an inner nesting level of one.

  • count – How many entries of this level to wrap, counted from row 0.

h5col.arrow.specs_from_arrow(table: Any) → list[Any][source]

The column specs this package would import an Arrow table with.

Returned so they can be inspected and adjusted before anything is written. Chunking and filters have no Arrow equivalent and so cannot be inferred at all. The string widths and category sets are inferred from the data, which is worth checking on a table you did not write:

specs = h5col.specs_from_arrow(tbl)
specs[2].chunks = 8192
specs[2].filters = FilterPipeline([Shuffle(), Deflate(4)])

Nothing is read from or written to a file here.

Arrow’s model is wider than H5Col’s, so a type with no exact H5Col equivalent is refused by name rather than approximated: timestamps, dates, times, durations, decimals, structs, maps, unions, opaque binary and fixed-size lists.

Field metadata under h5col. becomes the column’s own annotations, the same keys column_metadata() writes. Any other metadata is carried across as a producer attribute, unless its name is one H5Col reserves.

Arrow marks a missing value with a null; H5Col marks one with a value drawn from the column’s own domain. The specs come back with fill_value unset, which means the recommended value for the datatype, as it does anywhere else. Whether that value is safe for this data is not decided here — a fill that already occurs in a column is refused when the table is written, since setting one is the remedy and it has to be possible to get the specs in order to set it:

specs = h5col.specs_from_arrow(tbl)
specs[0].fill_value = 254
h5col.Table.from_arrow(group, tbl, specs=specs)

Added in version 0.4.0.

Parameters:

table – A pyarrow.Table.

Raises:
  • SchemaError – If two fields share a name (HDF5 links are unique), if a column’s type has no H5Col equivalent, or if a h5col. metadata key is not one this importer understands.

  • ReservedNameError – If a column name, or a producer metadata key, is a name H5Col reserves.

h5col.arrow.prepared_specs(table: Any, specs: Any = None) → list[Any][source]

Specs for importing table, with every fill value checked against data.

The fill checks are not skippable by supplying specs: choosing a value that already occurs in the column is the one importing mistake that produces a conformant file with unreadable rows, so it is verified whatever the specs came from. Supplied specs are copied rather than mutated.

Parameters:
  • table – The pyarrow.Table being imported.

  • specs – A complete list of column specs, as from specs_from_arrow(), or None to infer them.

Raises:

SchemaError – If specs does not name exactly the table’s columns.

h5col.arrow.append_values(spec: Any, column: Any) → Any[source]

One record-batch column in a form append() accepts.

Numeric columns have their nulls replaced by the column’s fill value before NumPy sees them: to_numpy on a nullable integer column upcasts to float64 and turns the nulls into NaN, which would change the datatype of every value in the column, not only the missing ones.

Parameters:
  • spec – The column’s spec, which carries the fill value and the datatype.

  • column – The batch’s column, as a pyarrow.Array.

h5col.missing#

Fill values and the canonical H5Col missing-value test.

A column marks a missing row with its HDF5 fill value. H5Col recommends a per-datatype sentinel chosen to lie outside the column’s logical value range, and defines a single canonical missing-value test that both the fill-equality and NaN cases reduce to.

h5col.missing.recommended_fill(dtype: Any) → Any[source]

Return H5Col’s recommended fill value for dtype.

  • Fixed- or variable-length string dtypes → b"".

  • Opaque dtypes → opaque_fill_bytes() for that width.

  • Enumerations with a MISSING member → the integer code of that member (the spec’s enum fill convention).

  • Enumerations without a MISSING member, including the H5Col boolean datatype (which MUST NOT declare a fill value at all) → raises.

  • Integer and float families → the tabulated value for that width.

  • Anything else (e.g. float16) → raises FillValueError.

Parameters:

dtype – Anything numpy.dtype() accepts, including h5py string and enumeration dtypes, whose metadata decides which rule above applies.

h5col.missing.masked_to_none(values: Any) → Any[source]

Rewrite a 1-D masked array to a list carrying None where it was masked.

A masked element and a None say the same thing on append — this row is missing — so folding one into the other keeps a single missing-value path through every encoder. The payload under a mask is discarded deliberately: numpy.ma promises nothing about it, and in practice it holds whatever stale arithmetic left behind rather than the column’s fill value.

Anything that is not a 1-D masked array is returned unchanged, and a masked array with nothing masked yields its plain data, so the typed fast paths downstream are left undisturbed.

Parameters:

values – Usually a 1-D numpy.ma.MaskedArray. Anything else, including a plain array or a list, is returned unchanged.

h5col.missing.is_missing(values: Any, fill_value: Any) → NDArray[bool][source]

Apply the canonical missing-value test element-wise.

missing(v, f) = isnan(f) ? isnan(v) : v == f — i.e. when the fill value is a NaN bit pattern the test is isnan(v); otherwise it is bit/value equality.

Parameters:
  • values – The stored values to test, as read from a column.

  • fill_value – The column’s declared fill value. A NaN may be given as a Python float, a NumPy scalar or a 0-d array; all three take the isnan branch.

h5col.missing.validate_fill_outside_range(fill: Any, valid_min: Any | None = None, valid_max: Any | None = None) → None[source]

Check that fill lies strictly outside [valid_min, valid_max].

Parameters:
  • fill – The column’s fill value.

  • valid_min – Lower bound of the column’s declared valid range, or None for unbounded below.

  • valid_max – Upper bound, or None for unbounded above. With both bounds None there is nothing to check and the call succeeds.

Raises:

FillValueError – If fill falls inside the declared range, where a genuine value could collide with it.

h5col.opaque#

Opaque columns: raw bytes of a fixed width per row.

HDF5’s opaque datatype (H5T_OPAQUE) stores a fixed number of bytes per value with no interpretation attached — digests, ciphertext, packed records, anything whose meaning lives outside the file. H5Col permits it as a column datatype and gives it a sorting order and a search-index hash, but the recommended fill-value table has no row for it, because no byte string can be reserved on the grounds of being out of range.

This module supplies the one that H5Col writes by default, and the predicate that tells a plain opaque datatype apart from the compound and array datatypes h5py also reads through NumPy’s V kind.

h5col.opaque.OPAQUE_FILL_MAGIC = b'FILL'

ASCII FILL, so the value announces itself in a hex dump rather than looking like data.

Type:

Opening bytes of an opaque column’s recommended fill value

h5col.opaque.is_opaque_dtype(dtype: Any) → bool[source]

True for a plain opaque datatype — raw bytes of a fixed width.

V is NumPy’s void kind, which covers plain raw bytes, structured datatypes and fixed-shape sub-arrays alike. h5py reads three different HDF5 datatypes through it — H5T_OPAQUE, H5T_COMPOUND and H5T_ARRAY — so a column’s kind alone does not say which arrived. Only the last two carry a structure, and those two structural fields are what separates them.

Added in version 0.4.0.

Parameters:

dtype – Anything numpy.dtype() accepts.

h5col.opaque.opaque_fill_bytes(nbytes: int) → bytes[source]

H5Col’s recommended fill value for an opaque column nbytes wide.

Since no byte string can be reserved by being out of range, the next best thing is one that is vanishingly unlikely to occur: the ASCII FILL marker, followed by rising byte values that wrap through zero.

The rising tail is the part that earns its keep. A fill of all zeros or all 0xFF would collide constantly — that is what zero padding, erased flash and uninitialized memory leave behind — while a counting sequence is something real data almost never is. An eight-byte column gets 46 49 4c 4c 01 02 03 04, which reads as FILL.... in a hex dump.

A column narrower than the marker gets as much of it as fits. Note that a one-byte opaque column has 256 possible values and this claims one of them, which is a real risk rather than a negligible one; such a column is better given a fill value chosen against its own data.

Added in version 0.4.0.

Parameters:

nbytes – The column’s fixed width. The result is exactly this long.

h5col.ordering#

The H5Col canonical ordering: orderability and min/max under the defined order.

H5Col defines a total order for a fixed set of datatypes (spec, “Sorted-row permutation index” / Ordering): integers arithmetically, floats by IEEE 754 over finite values and infinities (NaN unordered), booleans and enums by code, and strings byte-wise over UTF-8 with trailing storage padding stripped. CHUNK_MINMAX and SORTED_ROWS indexes may only be built over orderable datatypes; is_orderable() is that predicate.

NumPy’s S-dtype comparison is full-width memcmp over NUL-padded values, which is order-equivalent to the spec’s trailing-NUL-stripped byte-wise rule (NUL sorts below every byte), so NULLTERM/NULLPAD fixed strings compare natively. SPACEPAD strings need their trailing spaces stripped first (normalize_strings()).

h5col.ordering.is_orderable(dtype: Any) → bool[source]#

True if dtype has an H5Col-defined order.

Excluded per the spec: object/region references, compound datatypes, array datatypes, and variable-length-array datatypes. Everything else the spec enumerates — integers, floats, booleans, strings (fixed and variable length), opaque, and enumerations — is orderable.

Parameters:

dtype – Anything numpy.dtype() accepts, including h5py reference, string, variable-length and enumeration dtypes, whose metadata decides the answer.

h5col.ordering.is_spacepad(dataset: Any) → bool[source]#

True if dataset stores fixed-length strings with space padding.

Parameters:

dataset – An h5py dataset. Anything whose datatype cannot be inspected answers False, so a non-string dataset is safe to pass.

h5col.ordering.normalize_strings(values: ndarray, *, spacepad: bool) → ndarray[source]#

Strip trailing storage padding so byte-wise comparison matches the spec.

NUL padding needs no work (memcmp order-equivalence above); space padding is stripped explicitly.

Parameters:
  • values – Stored fixed-length string values.

  • spacepad – Whether the dataset pads with spaces rather than NULs, as reported by is_spacepad(). False returns values unchanged.

h5col.ordering.min_max(values: ndarray) → tuple[Any, Any][source]#

Min and max of values under the H5Col order.

Flexible dtypes (fixed strings) cannot use NumPy reductions, so they sort instead.

Parameters:

values – Orderable elements with the missing and NaN ones already removed, and at least one left. Both conditions are the caller’s to meet.

Raises:

SchemaError – If values is empty.

h5col.references#

HDF5 object-reference backend for H5Col.

All object-reference creation and resolution in H5Col goes through this module, so the on-disk reference representation can be changed in one place.

Warning

H5Col mandates the unified H5T_STD_REF datatype (HDF5 1.12+) and forbids the deprecated H5T_STD_REF_OBJ. h5py (as of 3.16) cannot create H5T_STD_REF, so this backend currently writes H5T_STD_REF_OBJ. This is a documented deviation (see docs/DEVIATIONS.md D1). The read side accepts either representation. A conformant backend can replace this module without changing any caller.

h5col.references.ref_dtype() → dtype[source]#

Return the NumPy dtype used to store object references.

h5col.references.is_reference_dtype(dtype: Any) → bool[source]#

Return True if dtype is an HDF5 object/region reference dtype.

Parameters:

dtype – Any dtype. One carrying no h5py reference metadata answers False.

h5col.references.make_ref(obj: Any) → Reference[source]#

Create an object reference to an open HDF5 object.

Parameters:

obj – An open h5py dataset or group. It must still be open — a reference cannot be made from a closed object.

Raises:

ObjectReferenceError – If obj exposes no ref (it is not a referable HDF5 object).

h5col.references.is_null_ref(ref: Any) → bool[source]#

Return True if ref is a null object reference.

Parameters:

ref – An object reference, typically read from an attribute. A null one points at nothing and cannot be dereferenced.

h5col.references.write_ref_attr(target: Any, name: str, obj: Any) → None[source]#

Write a scalar object-reference attribute onto target.

Parameters:
  • target – The object to write the attribute on.

  • name – Attribute name. An existing attribute of that name is replaced.

  • obj – The object to point at.

h5col.references.write_ref_array_attr(target: Any, name: str, objs: Iterable[Any]) → None[source]#

Write a 1-D object-reference array attribute onto target.

Parameters:
  • target – The object to write the attribute on.

  • name – Attribute name. An existing attribute of that name is replaced.

  • objs – The objects to point at, in order. An empty iterable writes an empty array.

h5col.references.append_ref_to_array_attr(target: Any, name: str, obj: Any) → None[source]#

Append a reference to obj to a 1-D reference-array attribute on target.

Creates the attribute when absent. HDF5 attributes cannot be resized in place, so an existing attribute is rewritten with the extended array. The new reference is created before the old attribute is touched, and the old array is restored if the rewrite fails, so a failed append cannot silently drop the existing references. (A hard crash between the delete and the create can still lose the attribute — HDF5 offers no atomic attribute rewrite.)

Parameters:
  • target – The object carrying the attribute.

  • name – Attribute name. It is created when absent.

  • obj – The object to append a reference to.

h5col.references.resolve(where: Any, ref: Any) → Any[source]#

Dereference ref relative to file/group where.

Parameters:
  • where – Any h5py file or group in the same file; references are file-scoped, so which one does not matter.

  • ref – The object reference to resolve.

Raises:

ObjectReferenceError – If ref is null, or does not resolve to an object in the file.

h5col.reserved#

H5Col reserved names, tokens, and name-validation helpers.

H5Col writes its reserved attribute and group names in fixed-length uppercase ASCII (with a few lowercase exceptions borrowed from broader community practice or from AnnData). This module centralizes those tokens so the rest of H5Col never hard-codes a spelling, and provides validators for producer-chosen names.

h5col.reserved.RESERVED_ATTRIBUTE_NAMES = frozenset({'CATEGORIES', 'CLASS', 'GENERATION', 'INDEX_COLUMNS', 'KIND', 'MASK', 'NROWS', 'SEARCH_INDEX_LIST', 'SOURCE_GENERATION', 'SOURCE_NROWS', 'TITLE', 'VALUES', 'VERSION', 'valid_max', 'valid_min'})#

Attribute names H5Col treats as reserved (producers must not repurpose them).

Return True if name is usable as an HDF5 link name (UTF-8, no //NUL).

Parameters:

name – Any object. A non-string, the empty string, and a string containing / or a NUL all answer False rather than raising.

h5col.reserved.validate_column_name(name: str) → str[source]#

Validate a producer-chosen column name and return it unchanged.

Parameters:

name – The proposed column name.

Raises:
  • SchemaError – If name is not a valid HDF5 link name.

  • ReservedNameError – If name collides with a H5Col reserved group, member, or attribute name.

h5col.reserved.validate_index_dataset_name(name: str) → str[source]#

Validate a search-index dataset name and return it unchanged.

Datasets in SEARCH_INDEXES may have any name — the spec assigns names no meaning — but the name must still be a single HDF5 link name, and reserved-names rule 2 forbids reusing any reserved name for a search-index dataset.

Parameters:

name – The proposed index dataset name.

Raises:
h5col.reserved.is_discouraged_column_name(name: str) → bool[source]#

Return True for names H5Col says producers SHOULD avoid (leading _).

Parameters:

name – A column name. This is advisory only — such a name is discouraged by the convention, not rejected.

h5col.reserved.COLUMN_ANNOTATION_NAMES = frozenset({'description', 'ordered', 'units', 'units_vocabulary'})#

Attribute names H5Col itself writes on a column, beyond the reserved set. These are not reserved names — a column may legitimately be called units — but a producer attribute of the same name would land on the very attribute the convention writes there, so free-form metadata may not use one.

h5col.reserved.validate_attribute_names(attributes: dict[str, object] | None, column: str) → None[source]#

Check that producer attribute names do not shadow H5Col’s own.

Parameters:
  • attributes – Extra attributes destined for a column, or None.

  • column – The column’s name, used only to make the error message specific.

Raises:
  • SchemaError – If a name is not a valid HDF5 attribute name.

  • ReservedNameError – If a name is one H5Col reserves. The convention gives those names a meaning, so a producer’s own metadata must not occupy them — a CLASS or NROWS attribute written as free-form metadata would make the column unreadable.

h5col.lists#

List columns: the H5Col offsets encoding for variable-length row values.

A list column is an HDF5 group (a direct child of the table group) with CLASS="LIST_COLUMN" and KIND="OFFSETS". It stores a variable-length list per row using the Apache Arrow offsets layout: all elements are flattened back-to-back into a VALUES member, and a monotonic OFFSETS dataset records each entry’s slice. VALUES is one of three things — a rank-1 leaf dataset, a nested LIST_COLUMN group (lists of lists), or a STRING_VALUES group (variable-length UTF-8 via a second OFFSETS/CHARS level). An optional MASK at any level distinguishes a null entry from an empty one.

The create/append/read/validate functions here are driven by the file structure (not the Python spec), so a table can be reopened and appended without its original ListColumnSpec. Writing follows the H5Col leaf-first order (deepest elements first, each enclosing OFFSETS last) so that committed rows stay fully described at every moment.

h5col.lists.reject_vlen(dtype: Any) → None[source]#

Raise if dtype — or any datatype nested inside it — is variable-length.

H5Col rule 11 forbids any HDF5 variable-length datatype anywhere below a list column, including one hidden inside a compound field or an array subtype. h5py’s check_string_dtype / check_vlen_dtype only inspect the top level, so this descends into compound fields and array bases first.

Parameters:

dtype – Any dtype, including a compound or array dtype whose members are inspected recursively.

h5col.lists.validate_list_column_spec(spec: ListColumnSpec) → None[source]#

Validate a list column spec without touching the file.

Parameters:

spec – The spec to check. Its name and its whole values tree, however deeply nested, are validated.

h5col.lists.create_list_column(table_group: Any, spec: ListColumnSpec, *, default_chunk_bytes: int | None = None) → Any[source]#

Create an empty list column group under table_group from spec.

Parameters:
  • table_group – The table group to create the column group in.

  • spec – The list column’s spec, including its nesting and its leaf values.

  • default_chunk_bytes – Target chunk size for the datasets created here, when the spec sets no explicit chunks shape.

h5col.lists.append_list_column(level_group: Any, rows: list[Any], n_old: int) → None[source]#

Append rows (a list of per-row list values) to a list column group.

Parameters:
  • level_group – The list column group, or an inner nesting level of one.

  • rows – One entry per new row: a list of values, or None for a null row.

  • n_old – The row count before this append, which is where the new rows are written.

h5col.lists.read_list_column(level_group: Any, count: int, *, start: int = 0) → list[Any][source]#

Read count entries from row start as a list of (list | None).

Only the rows asked for are read, at every level of nesting: the range’s own OFFSETS say where its values begin and end in the child, so the child read is narrowed to that span, and so on down.

The one thing to be careful about is that offsets are absolute positions in the child, while the child block just read begins at the range’s first offset. Every offset therefore has to be rebased against that first one before it can index the block.

Parameters:
  • level_group – The list column group, or an inner nesting level of one.

  • count – How many entries to read. Zero or fewer reads nothing and touches no dataset.

  • start – The first row to read. Defaults to 0, so the older two-argument call still reads [0, count).

h5col.lists.validate_list_column(level_group: Any, count: int) → None[source]#

Validate a list column subtree at count entries (recursively).

Parameters:
  • level_group – The list column group, or an inner nesting level of one.

  • count – How many entries this level is expected to hold, which its OFFSETS and MASK extents are checked against.

h5col.indexes#

Search-index engine: validity tokens, SEARCH_INDEXES, and index families.

Like h5col.lists, this engine is file-driven: every operation reads the structure it needs from the file, so indexes built in one session are maintainable and queryable after reopening.

The validity-token protocol (spec, “Index validity tokens”) is the backbone:

  • GENERATION (table group) identifies the current state of the column data; it is created the first time an index is added and incremented by every mutation thereafter.

  • SOURCE_GENERATION / SOURCE_NROWS (each index dataset) name the table state the index content was built against. index_is_valid() is the consumer check; a failed check means “treat the index as absent”, never an error.

  • Write ordering keeps every crash state detectable. Mutations gated by the NROWS commit (append) write future-valued tokens before the index content; building or refreshing an index over an already-committed state writes content first and current-valued tokens last.

h5col.indexes.INDEX_CHUNK_BYTES = 65536#

Byte target for one chunk of a search-index dataset. Index datasets scale with the source column’s chunk count, not its row count, so the column chunk policy would waste a mostly-empty multi-MiB chunk here; 64 KiB keeps index datasets compact while still amortizing appends.

h5col.indexes.MINMAX_FIELDS = ('min', 'max', 'nan_count', 'fill_count', 'n')#

CHUNK_MINMAX compound fields, in required declaration order.

h5col.indexes.SUPPORTED_KINDS = frozenset({'BITMAP', 'CHUNK_MINMAX', 'SORTED_ROWS'})#

Index kinds this implementation can build and maintain (grows in 4c).

h5col.indexes.append_refresh_indexes(table_group: Any, g_old: int, n_new: int) → bool[source]#

Mutation-protocol step 4: maintain every supported index for the new state.

Used by any NROWS-gated mutation — append and truncation share the same steps 4-6, with n_new the post-mutation row count. For each maintained index, the future-valued tokens (SOURCE_GENERATION = g_old + 1, SOURCE_NROWS = n_new) are written before the content — the index fails the validity check throughout its own rebuild, because the new generation does not yet exist on the table group. The caller commits GENERATION and NROWS afterwards.

Indexes this producer cannot rebuild are left untouched, with their tokens intact — including any index claimed by more than one column, where there is no correct column to rebuild against (the spec forbids the state, and validate reports it). Returns True when any index was rewritten, so the caller knows to flush.

Parameters:
  • table_group – The table’s HDF5 group.

  • g_old – The pre-mutation GENERATION, from mutation_generation(). The tokens are written for g_old + 1, the generation the caller is about to commit.

  • n_new – The post-mutation row count the caller is about to commit.

h5col.indexes.bitmap_bytes(nrows: int) → int[source]#

Bytes per bitmap row for a table of nrows (ceil(nrows / 8)).

Parameters:

nrows – The table’s row count, one bit apiece.

h5col.indexes.bitmap_values_dataset(table_group: Any, index_ds: Any) → Any | None[source]#

Resolve a BITMAP index’s accompanying values dataset, or None.

Consumer-lenient: a missing, non-scalar, null, dangling, or unlinked VALUES reference — or a target that is not a KIND-less rank-1 dataset sitting next to the index under SEARCH_INDEXES — yields None, making the bitmap unusable rather than an error; validate reports the violation separately.

Parameters:
  • table_group – The table’s HDF5 group, against which the reference is resolved.

  • index_ds – The BITMAP index dataset carrying the VALUES reference.

h5col.indexes.column_datasets(table_group: Any) → dict[str, Any][source]#

Direct-child rank-1 datasets of the table group (the scalar columns).

Parameters:

table_group – The table’s HDF5 group. List columns are groups and so are not returned.

h5col.indexes.column_index_datasets(table_group: Any, column_ds: Any) → list[Any][source]#

Resolve the column’s SEARCH_INDEX_LIST references, in order.

Consumer-lenient: null, dangling, and unlinked references are skipped — their indexes are treated as absent, per the spec’s tolerance rules — while validate reports them as rule-4 violations. A malformed (non-1-D) attribute raises, because no reference can be read from it.

Parameters:
  • table_group – The table’s HDF5 group, against which the references are resolved.

  • column_ds – The column dataset whose SEARCH_INDEX_LIST is read.

h5col.indexes.compute_bitmap(column_ds: Any, nrows: int) → tuple[ndarray, ndarray, bool][source]#

Recompute a BITMAP enumeration for rows [0, nrows).

Returns (values, bits, exhaustive): the distinct non-missing values in H5Col order, the (K, ceil(nrows / 8)) uint8 bit matrix with bit r % 8 of byte r // 8 set where row r equals the k-th value (pad bits zero), and the exhaustive claim. NaN cannot be enumerated — IEEE 754 equality never matches it — so non-missing NaN elements are left out of the enumeration and make the claim exhaustive = False.

Parameters:
  • column_ds – The column dataset to derive the index from.

  • nrows – The table’s row count; rows at or above it are reserved storage and are not indexed.

h5col.indexes.compute_chunk_minmax(column_ds: Any, nrows: int) → ndarray[source]#

Recompute the CHUNK_MINMAX entries for rows [0, nrows).

Creation, append maintenance, refresh, and deep validation all derive the index content from this one function, so they cannot disagree.

Parameters:
  • column_ds – The column dataset to derive the index from.

  • nrows – The table’s row count; rows at or above it are reserved storage and are not indexed.

h5col.indexes.compute_sorted_rows(column_ds: Any, nrows: int) → tuple[ndarray, int, int][source]#

Recompute the SORTED_ROWS permutation for rows [0, nrows).

Returns (permutation, fill_tail_length, nan_tail_length). The permutation is total and deterministic: the body is sorted under the H5Col order with ties broken by increasing row position (a stable argsort over rows already in increasing order), followed by the fill tail and then the NaN tail, each in increasing row order. A row goes to the NaN tail if its value is NaN, and otherwise to the fill tail if it matches a non-NaN fill; with a NaN fill every missing row is a NaN row and the fill tail is empty.

Parameters:
  • column_ds – The column dataset to derive the index from.

  • nrows – The table’s row count; rows at or above it are reserved storage and are not indexed.

h5col.indexes.create_bitmap(table_group: Any, column_ds: Any, *, name: str | None = None, description: str | None = None) → Any[source]#

Build a BITMAP index over column_ds, with its values dataset.

The accompanying values dataset is created as <name>_values next to the bitmap (its name carries no meaning; the linkage is the bitmap’s scalar VALUES object reference). The enumeration is the distinct non-missing values in H5Col order, so ordered is true; exhaustive is true unless the column holds non-missing NaN elements, which IEEE 754 equality makes impossible to enumerate.

Parameters:
  • table_group – The table’s HDF5 group. The index is created under its SEARCH_INDEXES group, which is created when absent.

  • column_ds – The column dataset to index.

  • name – Name for the index dataset. None derives one from the column name and the kind; the name carries no meaning, the linkage is the object reference in the column’s SEARCH_INDEX_LIST.

  • description – Free text stored on the index as its DESCRIPTION attribute.

Raises:
  • ConformanceError – If the table group has no NROWS attribute.

  • SchemaError – If the column’s datatype is unsupported, or SEARCH_INDEXES already holds the bitmap name or its <name>_values name.

  • ReservedNameError – If name (or <name>_values) is a H5Col reserved name.

h5col.indexes.create_chunk_minmax(table_group: Any, column_ds: Any, *, name: str | None = None, description: str | None = None) → Any[source]#

Build a CHUNK_MINMAX index over column_ds and link it.

Building over an already-committed table state writes the index content first and the current-valued tokens last, so a crash mid-build leaves a dataset that fails the validity check.

Parameters:
  • table_group – The table’s HDF5 group. The index is created under its SEARCH_INDEXES group, which is created when absent.

  • column_ds – The column dataset to index.

  • name – Name for the index dataset. None derives one from the column name and the kind; the name carries no meaning, the linkage is the object reference in the column’s SEARCH_INDEX_LIST.

  • description – Free text stored on the index as its DESCRIPTION attribute.

Raises:
  • ConformanceError – If the table group has no NROWS attribute.

  • SchemaError – If the column’s datatype is unsupported, or SEARCH_INDEXES already holds a dataset of the chosen name.

  • ReservedNameError – If name is a H5Col reserved name.

h5col.indexes.create_sorted_rows(table_group: Any, column_ds: Any, *, name: str | None = None, description: str | None = None) → Any[source]#

Build a SORTED_ROWS index over column_ds and link it.

Building over an already-committed table state writes the index content first and the current-valued tokens last, so a crash mid-build leaves a dataset that fails the validity check.

Parameters:
  • table_group – The table’s HDF5 group. The index is created under its SEARCH_INDEXES group, which is created when absent.

  • column_ds – The column dataset to index.

  • name – Name for the index dataset. None derives one from the column name and the kind; the name carries no meaning, the linkage is the object reference in the column’s SEARCH_INDEX_LIST.

  • description – Free text stored on the index as its DESCRIPTION attribute.

Raises:
  • ConformanceError – If the table group has no NROWS attribute.

  • SchemaError – If the column’s datatype is unsupported, or SEARCH_INDEXES already holds a dataset of the chosen name.

  • ReservedNameError – If name is a H5Col reserved name.

h5col.indexes.data_chunk_count(column_ds: Any, nrows: int) → int[source]#

Chunks of column_ds that contain logical-table rows.

ceil(nrows / chunk_len) for a chunked column, 1 for a contiguous column, 0 when nrows == 0. Tail-only chunks are not counted.

Parameters:
  • column_ds – The column dataset.

  • nrows – The table’s row count. Chunks holding only reserved rows above it are not counted.

h5col.indexes.ensure_generation(table_group: Any) → int[source]#

Return the table’s GENERATION, creating or repairing it when needed.

Per the spec, a table acquires GENERATION the first time a search index is built over it (“writing GENERATION first if the table did not previously carry it”); building over an unchanged table is not a mutation, so no increment happens here.

A missing-with-indexes or malformed GENERATION (rule-12 violations some foreign tool left behind) fails the strict validity check, so every token this producer would write against it is dead on arrival. It is repaired as scalar uint64 with a value strictly above the old value and above every index’s SOURCE_GENERATION — a spurious increment is explicitly safe (it can only disable indexes, never validate stale ones), whereas any reused value could equal some index’s token and spuriously validate content nobody has verified.

Parameters:

table_group – The table’s HDF5 group. The attribute is created on it when absent, and repaired when malformed.

h5col.indexes.find_index_column(table_group: Any, index_ds: Any) → Any | None[source]#

The column whose SEARCH_INDEX_LIST references index_ds, or None.

The column-side attribute is the only linkage the spec defines — index datasets carry no back-pointer — so this scans the table’s column datasets. A malformed (non-1-D) SEARCH_INDEX_LIST on some column is skipped, so one bad column cannot break lookups for every other column’s indexes.

An index claimed by more than one column violates the spec’s “a single search-index dataset MUST NOT cover multiple columns”; there is then no correct answer, and silently picking one would make pruning against the wrong column’s data possible — so this raises instead.

Parameters:
  • table_group – The table’s HDF5 group, whose column datasets are scanned.

  • index_ds – The search-index dataset to find the owner of.

h5col.indexes.index_is_valid(index_ds: Any, table_group: Any) → bool[source]#

The consumer validity check for one search-index dataset.

SOURCE_GENERATION == GENERATION AND SOURCE_NROWS == NROWS, with any absent or wrong-datatype token failing the check. A False result means “behave as if the index were not present” — it is never an error.

Parameters:
  • index_ds – The search-index dataset carrying the validity tokens.

  • table_group – The table’s HDF5 group, holding the values they are compared against.

h5col.indexes.index_kind(index_ds: Any) → str | None[source]#

The dataset’s KIND value, or None when absent or not a scalar string.

A malformed KIND (non-string or non-scalar value) yields None so that no kind-dispatched code path ever acts on it; validate flags the malformed attribute separately.

Parameters:

index_ds – A search-index dataset.

h5col.indexes.minmax_dtype(element_dtype: Any) → dtype[source]#

The CHUNK_MINMAX compound dtype for a column of element_dtype.

Parameters:

element_dtype – The column’s element dtype, which the min and max fields take.

h5col.indexes.mutation_generation(table_group: Any) → int | None[source]#

The pre-mutation GENERATION (g_old) for append/truncate.

Strict read: a well-formed token is returned as-is, and an absent one is None (a table that does not carry GENERATION skips the bump steps). A malformed token must not simply be incremented — its lenient integer value bypasses the safety property, because g_old + 1 could equal some index’s residue SOURCE_GENERATION and spuriously validate unverified content once step 5 rewrites the attribute as uint64. It is repaired via ensure_generation(), which picks a value above every source token.

Parameters:

table_group – The table’s HDF5 group.

h5col.indexes.refresh_all_indexes(table_group: Any) → int[source]#

Rebuild every supported stale index against the committed state.

Returns the number of indexes refreshed. Indexes this producer cannot rebuild are left untouched — and stay detectably stale if they already were. Currently valid indexes are also left untouched: they already describe the committed state, and rewriting them in place would open a crash window with torn content behind passing tokens.

Parameters:

table_group – The table’s HDF5 group, whose committed state every index is rebuilt against.

h5col.indexes.refresh_index(table_group: Any, index_ds: Any, column_ds: Any) → None[source]#

Rebuild index_ds against the table’s current committed state.

The current GENERATION/NROWS are already committed, so the order is content first, tokens last (writing current-valued tokens before the content would let a mid-rebuild index pass the check).

A currently valid index is left untouched: its content already describes the committed state (rule 9), and rewriting it in place would open a crash window where torn content sits behind still-passing tokens — the one state the token protocol exists to prevent.

Parameters:
  • table_group – The table’s HDF5 group, whose committed state the index is rebuilt against.

  • index_ds – The index dataset to rebuild.

  • column_ds – The column the index covers, which its content is derived from.

Raises:

ConformanceError – If the table group has no NROWS attribute.

h5col.indexes.search_index_datasets(table_group: Any) → dict[str, Any][source]#

Every search-index dataset (carries KIND) under SEARCH_INDEXES.

A SEARCH_INDEXES child that is not a group (a misuse of the reserved name) holds no index datasets; validate flags it, the consumer paths simply see no indexes.

Parameters:

table_group – The table’s HDF5 group.

h5col.indexes.source_chunk_len(column_ds: Any, nrows: int) → int[source]#

Rows per chunk of the source column (its full extent when contiguous).

Parameters:
  • column_ds – The column dataset.

  • nrows – The table’s row count, used as the chunk length for a contiguous (unchunked) column, which has one notional chunk covering everything.

h5col.indexes.supported_index_dtype(dtype: Any) → bool[source]#

True if this implementation can build its index families over dtype.

A producer subset of is_orderable(), shared by every supported family: the spec also orders variable-length strings and opaque values, but this implementation does not build indexes over them (building an index is always optional for a producer).

Parameters:

dtype – A column’s element dtype.

h5col.indexes.supported_minmax_dtype(dtype: Any) → bool#

Backwards-compatible name from sub-phase 4a; the same predicate applies to every family this implementation builds.

h5col.indexes.table_generation(table_group: Any) → int | None[source]#

The table’s GENERATION, or None when it carries none.

Parameters:

table_group – The table’s HDF5 group.

h5col.indexes.validate_search_indexes(table_group: Any, nrows: int, *, deep: bool = False) → None[source]#

Enforce consistency rules 3, 4, 12, and rule 9 for every supported kind.

Rule 9 is applied to every index family this implementation understands (CHUNK_MINMAX, SORTED_ROWS, BITMAP); an index of an unsupported kind is skipped. Rule 9 applies only to indexes whose validity check passes; a stale index is exempt (consumers treat it as absent). Structural rule-9 checks always run; the semantic check — recomputing the index from its column — is O(index build) and runs only with deep=True.

Parameters:
  • table_group – The table’s HDF5 group.

  • nrows – The table’s committed row count, which the indexes’ extents and tokens are checked against.

  • deep – When True, also recompute each valid index from its column and compare, which costs an index build apiece.

Raises:

ConformanceError – On the first consistency violation found.