Tables and columns#
The three classes below are the reading-and-writing heart of the package.
Table wraps a table group; its
create() / open() classmethods are the
two entry points, and its mapping-style access (table["name"], in,
iteration) hands out the column wrappers.
The subscript, length and iteration behaviour is listed alongside the ordinary
methods: table["name"] hands out a column, and column[17:98] reads rows.
- class h5col.Table(group: Any)[source]#
A H5Col column-oriented table backed by an HDF5 group.
- static is_table_group(group: Any) bool[source]#
True if group is a H5Col table group (lenient CLASS check).
- Parameters:
group – Any h5py group. One that is not a table group answers False rather than raising, so this is safe to use while walking a file.
- classmethod open(group: Any) Table[source]#
Open an existing table group, checking its CLASS and VERSION major.
- Parameters:
group – An h5py group already written as a H5Col table. Opening does not read any column data.
- Raises:
ConformanceError – If group is not a H5Col table group, or its
VERSIONis missing or unparsable.VersionError – If the table’s
VERSIONmajor exceeds the supported major.
- classmethod create(group: Any, columns: TableSpec | Sequence[ColumnSpec | ListColumnSpec], *, title: str | None = None, description: str | None = None, index_columns: Sequence[str] | None = None, column_order: Sequence[str] | None = None, units_vocabulary: str | None = None, encoding_type: str | None = None, encoding_version: str | None = None, default_chunk_bytes: int | None = None) Table[source]#
Create a new, empty table (
NROWS = 0) with the given columns.- Parameters:
group – An h5py group to write the table into. It must not already be a H5Col table group.
columns – A
TableSpec, or a sequence ofColumnSpecandListColumnSpec.title – Human-readable table title, stored as the
TITLEattribute.description – Longer free text, stored as
DESCRIPTION.index_columns – Names of the columns that together identify a row, stored as
INDEX_COLUMNS. Every name must be a declared scalar column; a list column cannot be an index column.column_order – The columns’ logical order, stored as
COLUMN_ORDER. Defaults to the order given in columns, which HDF5 itself does not preserve.units_vocabulary – The vocabulary the columns’
unitscome from, e.g.UDUNITS-2.encoding_type – Producer-defined encoding identifier for the table as a whole.
encoding_version – Version string paired with encoding_type.
default_chunk_bytes – Overrides the automatic (chunk-cache-scaled) target chunk size for columns that do not set an explicit
chunksshape.
- Raises:
SchemaError – If group is already a H5Col table group, a column spec is invalid, or an
index_columnsname is not among the declared columns.ReservedNameError – If a column name is a H5Col reserved name.
FillValueError – If a column’s fill value lies inside its declared valid range.
- classmethod from_arrays(group: Any, arrays: Mapping[str, Any], *, specs: Sequence[ColumnSpec] | None = None, **table_kwargs: Any) Table[source]#
Create a table from column arrays and write them in one call.
- Parameters:
group – An h5py group to write the table into, as for
create().arrays – One array per column, keyed by column name. Every array must hold the same number of rows, and the column order follows this mapping.
specs – Column specs to create the table with. When omitted, one is inferred per array — boolean, a fixed-length string sized to the longest value, or the array’s own dtype.
table_kwargs – Passed through to
create(), sotitle,index_columnsand the rest are available here too.
- classmethod from_arrow(group: Any, table: Any, *, specs: Sequence[ColumnSpec | ListColumnSpec] | None = None, **table_kwargs: Any) Table[source]#
Create a table from a
pyarrow.Tableand write its rows.Arrow’s model is wider than H5Col’s, so anything without an exact equivalent is refused rather than approximated, see
specs_from_arrow(), which decides the mapping and which you can call first to inspect or adjust it:specs = h5col.specs_from_arrow(tbl) specs[2].chunks = 8192 Table.from_arrow(group, tbl, specs=specs)
Rows are written batch by batch rather than all at once, so importing a large table costs about one batch of memory rather than the whole of it.
Two chunks of one dictionary column may carry different dictionaries, in which case the same code stands for two different labels. The codes are never read directly for that reason — the labels are, against the unified category set
specs_from_arrow()derives.The fill-value checks run whichever way the specs arrived: a fill that already occurs in a column would leave those rows reading as missing, so supplying specs cannot skip it.
Needs the optional
pyarrowdependency (pip install h5col[arrow]).Added in version 0.4.0.
- Parameters:
group – An h5py group to write the table into, as for
create().table – The
pyarrow.Tableto import.specs – A complete list of column specs naming exactly the table’s columns. None infers them. Chunking and filters have no Arrow equivalent, so this is the only way to set them.
table_kwargs – Passed through to
create(), sotitle,index_columnsand the rest are available here too.
- Raises:
SchemaError – If a column’s type has no H5Col equivalent, if a fill value occurs in its column’s data, if a boolean column holds nulls, or if specs does not name exactly the table’s columns.
ReservedNameError – If a column name, or a producer metadata key, is one H5Col reserves.
- property nrows: int#
The table’s logical row count (its
NROWSattribute).- Raises:
ConformanceError – If the group carries no
NROWSattribute.
- property generation: int | None#
The table’s
GENERATIONvalidity token, or None when absent.A table acquires
GENERATIONwhen its first search index is built and increments it on every subsequent mutation of committed data.
- property columns: dict[str, Column | ListColumn]#
The table’s columns by name, in column order, as wrapper objects.
- __getitem__(name: str) Column | ListColumn[source]#
The column named name, as a
ColumnorListColumn.- Parameters:
name – A column name. Raises
KeyErrorif the table has no such column.
- __contains__(name: str) bool[source]#
True if the table has a column of this name.
- Parameters:
name – A column name.
- append(data: Mapping[str, Any], *, maintain_indexes: bool = False) None[source]#
Append rows following the H5Col write protocol.
Every provided column must supply the same number of rows
K. A scalar column absent from data is extended and left as its fill value (missing); a boolean column (which has no fill) must always be provided. A list column absent from data must be nullable — its new rows become null lists; a non-nullable list column must always be provided.NROWSis committed last, then the file is flushed.Nonein a column’s values marks that row as missing and is stored as the column’s fill value (for a categorical column, its fill code). A column with no fill to store — a boolean, which H5Col forbids from declaring one — rejectsNoneinstead of coercing it.By default, search indexes are not maintained: the
GENERATIONincrement that publishes the append disables them, detectably, andrefresh_indexes()restores them later — this keeps the hot append path fast. Withmaintain_indexes=True, every supported index is rewritten inside the append protocol (future-valued tokens before content) and remains valid after the commit; indexes this implementation cannot rebuild — unsupported kinds, element dtypes the builder does not handle, non-growable index datasets — are left entirely untouched, tokens included.- Parameters:
data – New rows, one entry per column, keyed by column name. A
numpy.ma.MaskedArrayis accepted and its masked elements mean the same asNone.maintain_indexes – Rebuild the supported search indexes inside the write protocol so they stay valid after the commit, as described above.
- Raises:
OversizedStringError – If a fixed-length string value’s encoding exceeds the column’s byte budget (H5Col never silently truncates).
SchemaError – For unknown columns, values that are not 1-D, unequal column lengths, an omitted fill-less/boolean column, an omitted non-nullable list column, an unknown category label, or a
Nonein a column that declares no fill value.
- truncate(nrows: int, *, maintain_indexes: bool = False) None[source]#
Shrink the logical table to nrows rows (H5Col logical truncation).
The truncation is logical: no column dataset changes extent, and the rows
[nrows, old_NROWS)become reserved storage that consumers ignore. List columns need no extra writes — the smallerNROWSbounds their offsets recursively. Reclaiming physical space would require rewriting each column to its new extent, which this implementation does not do.Index handling mirrors
append()(the spec applies the same steps 4-6 with the new row count): by default every search index is left detectably stale by theGENERATIONbump; withmaintain_indexes=Truethe supported indexes are rebuilt inside the protocol and remain valid after the commit.Truncating to the current row count is a no-op (nothing changes, so nothing is published); growing is an error — that is what
append()is for.- Parameters:
nrows – The new row count. Must be between 0 and the current row count.
maintain_indexes – As for
append().
- Raises:
SchemaError – If nrows is negative or greater than the current row count.
- read(columns: Sequence[str] | None = None, *, where: Any = None, explain: bool = False, masked: bool = True) Any[source]#
Read columns (default all) as
{name: array}over[0, NROWS).- Parameters:
columns – Names to read, in the order given. None (the default) reads every column of the table.
where – Restrict the result to matching rows. Takes the same forms as
select(); None (the default) reads every row.explain – When True the return value becomes a
(result, QueryPlan)pair, the plan describing how the query was evaluated.masked – Return each scalar column as a
numpy.ma.MaskedArraywhose mask marks its missing rows (the default). List columns ignore it, already spelling a null rowNone. Pass False for plain arrays, in which case a missing row holds the column’s fill value with nothing to distinguish it from data.
- Raises:
KeyError – If a requested column name is not a column of the table.
SchemaError – If
where=is malformed or references an unknown column.
- to_arrow(columns: Sequence[str] | None = None, *, where: Any = None) Any[source]#
Convert the table (default all columns) to a
pyarrow.Table.The one export that carries the whole H5Col data model: missing rows become real Arrow nulls instead of the fill value, a categorical becomes a dictionary of the codes and labels already stored, a list column keeps its nulls at every level of nesting, and each column’s
units,descriptionand valid-range attributes ride along as Arrow field metadata under anh5col.prefix.Needs the optional
pyarrowdependency (pip install h5col[arrow]).Added in version 0.2.0.
- Parameters:
columns – Names to convert, in the order given. None (the default) converts every column of the table.
where – Restrict the result to matching rows. Takes the same forms as
select(); None (the default) converts every row.
- Raises:
KeyError – If a requested column name is not a column of the table.
- select(where: Any = None) Selection[source]#
Build a lazy
Selectionover the table.- Parameters:
where – A query
Expression, aList[Tuple](AND), or aList[List[Tuple]](OR-of-ANDs, pyarrow DNF). None (the default) selects every row.
- count(where: Any = None) int[source]#
Number of rows matching where, without reading any column values.
- Parameters:
where – As for
select(). None (the default) counts every row.
- build_index(column: str, kind: str | None = None, *, name: str | None = None, description: str | None = None) SearchIndex[source]#
Build a search index over column (alias of
add_search_index()).- Parameters:
column – As for
add_search_index().kind – As for
add_search_index().name – As for
add_search_index().description – As for
add_search_index().
- info(*, storage: bool = False, full: bool = False) TableInfo[source]#
Table content information, storage and compression details optional.
The result prints as an aligned block at a prompt and renders as a collapsible view in a Jupyter notebook, and its records stay reachable for anyone who wants the numbers rather than the picture.
Nothing here reads a column’s values, and by default nothing measures stored size or compression either, so even inspecting a large table requires a handful of metadata reads.
Added in version 0.5.0.
- Parameters:
storage – Whether to measure what each column costs in the file. Off by default because measuring walks every column’s chunk index.
full – Whether to report each column as a labelled block of everything known about it: description, valid range, per-dataset sizes, etc. This is the same detail a notebook shows behind a disclosure triangle. What a column costs is part of everything, so this sets
storage=Trueas well.
- property search_indexes: dict[str, SearchIndex]#
Every search-index dataset under
SEARCH_INDEXES, wrapped by kind.
- add_search_index(column: str, kind: str | None = None, *, name: str | None = None, description: str | None = None) SearchIndex[source]#
Build a search index over column and link it to the column.
With
kind=Nonethe family is picked automatically:BITMAPfor boolean and categorical columns (low cardinality, exact equality answers),CHUNK_MINMAXfor any other orderable column. Building an index over an unchanged table is not a mutation:GENERATIONis created (0) if absent but never incremented. The default dataset name is<column>__<kind, lowercased>— a readable convention only; the linkage is the object reference in the column’sSEARCH_INDEX_LIST.- Parameters:
column – Name of the column to index. It must be a scalar column; list columns cannot be indexed.
kind –
CHUNK_MINMAX,SORTED_ROWSorBITMAP. None (the default) picks the family that suits the column’s datatype, as above.name – Name for the index dataset under
SEARCH_INDEXES. None uses the<column>__<kind>default described above.description – Free text stored on the index as its
DESCRIPTIONattribute.
- Raises:
KeyError – If column is not a column of the table.
SchemaError – If column is a list column, no index family applies to its dtype (
kind=None), kind is unimplemented, orSEARCH_INDEXESalready holds a dataset of the chosen name.ReservedNameError – If name is a H5Col reserved name.
ConformanceError – If the table carries no
NROWSattribute.
- refresh_indexes() int[source]#
Rebuild every supported search index against the current table state.
Restores indexes left stale by
append(maintain_indexes=False)or by any other mutation. Returns the number of indexes refreshed; indexes of unsupported kinds are left untouched (and stay detectably stale).
- index_is_valid(index: SearchIndex | Any) bool[source]#
The H5Col consumer validity check for index.
- Parameters:
index – A
SearchIndexwrapper or the index dataset itself.
- validate(*, deep: bool = False) None[source]#
Check the H5Col consistency requirements, raising on any violation.
A stale index is never an error: the validity check disables it, as the spec intends.
- Parameters:
deep – When True, additionally re-derive every valid search index from its column and compare (consistency rule 9’s semantic half), which costs an index build apiece. The default run is structural only.
- Raises:
ConformanceError – On the first consistency violation found.
VersionError – If the table’s
VERSIONmajor exceeds the supported major.
- add_column(spec: ColumnSpec | ListColumnSpec, *, default_chunk_bytes: int | None = None) Column | ListColumn[source]#
Add a new column to an existing table (schema evolution).
A scalar column is created and grown to the table’s current extent; its existing rows read as its fill value (missing). A fill-less (boolean) column, or a list column (whose “missing” analogue is a null list that would need explicit backfilling), cannot represent pre-existing rows, so adding one to a table that already has rows is refused.
- Parameters:
spec – The new column’s
ColumnSpecorListColumnSpec. Its name must not already be in use.default_chunk_bytes – As for
create(): the target chunk size when the spec sets no explicitchunksshape.
- Raises:
SchemaError – If a column of that name already exists, the spec is invalid, or the column cannot backfill pre-existing rows (a boolean or list column on a non-empty table).
- class h5col.Column(dataset: Any, table: Table)[source]#
One column of a H5Col table, wrapping its rank-1 HDF5 dataset.
Reads decode to friendly Python values: fixed-length strings become
str, boolean columns become NumPybool, and numeric columns pass through.Rows can be read by subscript —
column[17:98],column[-1],column[[3, 1]]— which isread_rows()with its defaults. Note thatcolumn[...]andcolumn.dataset[...]are not the same read: the second goes straight to h5py, so it skips decoding and can return rows aboveNROWSthattruncate()left behind as reserved storage.- __len__() int[source]#
The number of rows in the column, which is the table’s
NROWS.Added in version 0.3.0.
- __iter__() Iterator[Any][source]#
Iterate the decoded rows.
Defined so that iterating reads the column once. Without it Python falls back on
__getitem__()and fetches every row separately, which is one HDF5 read per row.Added in version 0.3.0.
- __getitem__(key: Any) Any[source]#
Read rows by subscript, decoded and masked.
An integer key returns that one row’s value —
numpy.ma.maskedif the row is missing — and every other key returns an array, the way NumPy behaves.This is
read_rows()with the defaults. Subscript has nowhere to put a keyword, so it always decodes and always masks; callread()orread_rows()when you wantmasked=False.Added in version 0.3.0.
- Parameters:
key – An integer, a slice, a sequence of positions, or a boolean array with one entry per row. A negative position counts back from the last row.
- Raises:
IndexError – If a row position is out of range, a boolean mask is the wrong length, or key is a tuple.
TypeError – If key holds values that are neither integers nor booleans.
- property dtype: dtype#
The column’s NumPy dtype (its stored HDF5 datatype).
For a categorical column this is the integer code dtype, not the category labels (see
categories).
- property is_categorical: bool#
True if this is a categorical column (its values are category codes).
- property categories: NDArray[Any] | None#
The category labels, or None for a non-categorical column.
String labels come back as a compact NumPy string array; numeric labels keep their own dtype.
- property fill_value: Any#
The column’s fill value, or None when it declares none.
None is returned for boolean columns and for any column not in the
H5D_FILL_VALUE_USER_DEFINEDstate (e.g. a full-domain column that declares no missing-row semantics). h5py’s library-default fill value is not a H5Col sentinel and must not be surfaced as one.
- property units_vocabulary: str | None#
The column’s
units_vocabularyattribute, or None when unset.Names the vocabulary
unitsis drawn from — UDUNITS-2, say — so a reader can tell which spelling of a unit was meant. A table declares one for all its columns; this is the per-column override.Added in version 0.4.0.
- read(*, masked: Literal[True] = True) MaskedArray[source]#
- read(*, masked: Literal[False]) NDArray[Any]
- read(*, masked: bool) Any
Read the logical rows
[0, NROWS), decoded to friendly values.- Parameters:
masked – Return a
numpy.ma.MaskedArraywhose mask marks the column’s missing rows (the default). Pass False for the plain array, in which case a missing row holds the column’s fill value with nothing to distinguish it from data.
- read_rows(rows: Any, *, masked: Literal[True] = True) MaskedArray[source]#
- read_rows(rows: Any, *, masked: Literal[False]) NDArray[Any]
- read_rows(rows: Any, *, masked: bool) Any
Read just rows, decoded, in the order given.
A slice is read as a single hyperslab. Anything else is fetched with coalesced, chunk-aligned block reads, so a selection confined to a few chunks costs a few chunks rather than the whole column — in both time and peak memory.
- Parameters:
rows – A slice, a sequence of integer positions, or a boolean mask with one entry per row. A negative position counts back from the end, so
-1is the last row. Integer positions may be given in any order and may repeat; the result follows the order given.masked – As for
read().
- Raises:
IndexError – If a row position is out of range, or a boolean mask is the wrong length.
TypeError – If rows holds values that are neither integers nor booleans.
ValueError – If rows is not one-dimensional.
- to_arrow(rows: Any = None) Any[source]#
Convert the column to an Arrow array, or just rows of it.
Missing rows become real Arrow nulls rather than the fill value, and a categorical column becomes a
DictionaryArrayof the codes and labels H5Col already stores — neither of which NumPy can express.Needs the optional
pyarrowdependency (pip install h5col[arrow]).Added in version 0.2.0.
- Parameters:
rows – Which rows to convert, in any form
read_rows()accepts. None (the default) converts the whole column.
- property filters: FilterPipeline#
The filter pipeline this column’s dataset was created with.
What HDF5 recorded, which is not always what was declared: a plugin may add client data of its own while the dataset is created.
Added in version 0.5.0.
- property storage: Storage#
What this column costs in the file, before and after filtering.
Added in version 0.5.0.
- property search_indexes: list[SearchIndex]#
Search indexes bound to this column, from
SEARCH_INDEX_LIST.The wrappers are bound to this column, so their queries always run against it — even on a non-conformant file where another column also claims the same index dataset.
- Raises:
ConformanceError – If the column’s
SEARCH_INDEX_LISTattribute is malformed (not a 1-D array of object references).
- add_search_index(kind: str | None = None, *, name: str | None = None, description: str | None = None) SearchIndex[source]#
Build a search index over this column (
Table.add_search_index()).- Parameters:
kind – Which index family to build —
CHUNK_MINMAX,SORTED_ROWSorBITMAP. None (the default) picks the family that suits the column’s datatype.name – Name for the index dataset under
SEARCH_INDEXES. None derives one from the column name and the kind.description – Free text stored on the index as its
DESCRIPTIONattribute.
- Raises:
SchemaError – If no index family applies to the column’s dtype (
kind=None), kind is unimplemented, orSEARCH_INDEXESalready holds a dataset of the chosen name.ReservedNameError – If name is a H5Col reserved name.
ConformanceError – If the table carries no
NROWSattribute.
- build_index(kind: str | None = None, *, name: str | None = None, description: str | None = None) SearchIndex[source]#
Build a search index over this column (alias of
add_search_index()).- Parameters:
kind – As for
add_search_index().name – As for
add_search_index().description – As for
add_search_index().
- class h5col.ListColumn(group: Any, table: Table)[source]#
One list column of a H5Col table, wrapping its
CLASS=LIST_COLUMNgroup.Reading returns one Python
listper row (orNonefor a null list), with elements decoded to friendly values — nested lists become nested Python lists, string elements becomestr, and missing leaf elements becomeNone.- __len__() int[source]#
The number of rows in the column, which is the table’s
NROWS.Added in version 0.3.0.
- __getitem__(key: Any) Any[source]#
Read rows by subscript, the same keys a scalar column takes.
An integer returns one row — a list, or
Noneif the row is null — and every other key returns a list of rows. Negative positions count back from the last row.Only the rows asked for are read; see
read_rows()for what that means when the positions are scattered rather than a range.Added in version 0.3.0.
- Parameters:
key – An integer, a slice, a sequence of positions, or a boolean array with one entry per row.
- Raises:
IndexError – If a row position is out of range, a boolean mask is the wrong length, or key is a tuple.
TypeError – If key holds values that are neither integers nor booleans.
- property units_vocabulary: str | None#
The list column’s
units_vocabularyattribute, or None when unset.
- property filters: dict[str, FilterPipeline]#
The filter pipeline of each dataset making up this column.
A list column is several datasets rather than one, so this answers with a pipeline per dataset, keyed by the dataset’s path within the column group:
OFFSETS,VALUES,MASK, and for a column of stringsVALUES/OFFSETSandVALUES/CHARS. A column of lists of lists nests the same way.Added in version 0.5.0.
- property storage: Storage#
What this column costs in the file, its datasets added together.
Added in version 0.5.0.
- read(*, masked: bool = True) list[Any][source]#
Read rows
[0, NROWS)as a list of per-row lists (None= null).- Parameters:
masked – Accepted and ignored, so a caller can pass it uniformly across a table’s columns. A list column is ragged and so cannot be a
numpy.ma.MaskedArray; it already spells a null rowNone.
- read_rows(rows: Any, *, masked: bool = True) list[Any][source]#
Read just rows, in the order given, as a list of (list | None).
A contiguous range is read directly. Scattered positions are served from the range that spans them, from the lowest wanted row to the highest, which is never wider than the column itself: rows that sit near each other cost almost nothing, and rows spread from end to end cost what reading the column costs.
Added in version 0.3.0.
- Parameters:
rows – The same row specs
Column.read_rows()takes — a slice, a sequence of positions, or a boolean mask — so a caller can select rows the same way whatever kind of column it holds.masked – Accepted and ignored, as it is by
read().
Arrow interchange#
Table.to_arrow exports a table and
Table.from_arrow imports one, deciding each
column’s spec for itself. The function below returns those specs without
writing anything, so they can be looked at and adjusted first. It is also the
only way to set chunking and filters, which have no Arrow equivalent to be
inferred from. The guide chapter covers the whole
picture.
- h5col.specs_from_arrow(table: Any) list[Any][source]#
The column specs this package would import an Arrow table with.
Returned so they can be inspected and adjusted before anything is written. Chunking and filters have no Arrow equivalent and so cannot be inferred at all. The string widths and category sets are inferred from the data, which is worth checking on a table you did not write:
specs = h5col.specs_from_arrow(tbl) specs[2].chunks = 8192 specs[2].filters = FilterPipeline([Shuffle(), Deflate(4)])
Nothing is read from or written to a file here.
Arrow’s model is wider than H5Col’s, so a type with no exact H5Col equivalent is refused by name rather than approximated: timestamps, dates, times, durations, decimals, structs, maps, unions, opaque binary and fixed-size lists.
Field metadata under
h5col.becomes the column’s own annotations, the same keyscolumn_metadata()writes. Any other metadata is carried across as a producer attribute, unless its name is one H5Col reserves.Arrow marks a missing value with a null; H5Col marks one with a value drawn from the column’s own domain. The specs come back with
fill_valueunset, which means the recommended value for the datatype, as it does anywhere else. Whether that value is safe for this data is not decided here — a fill that already occurs in a column is refused when the table is written, since setting one is the remedy and it has to be possible to get the specs in order to set it:specs = h5col.specs_from_arrow(tbl) specs[0].fill_value = 254 h5col.Table.from_arrow(group, tbl, specs=specs)
Added in version 0.4.0.
- Parameters:
table – A
pyarrow.Table.- Raises:
SchemaError – If two fields share a name (HDF5 links are unique), if a column’s type has no H5Col equivalent, or if a
h5col.metadata key is not one this importer understands.ReservedNameError – If a column name, or a producer metadata key, is a name H5Col reserves.