Tables and columns#

The three classes below are the reading-and-writing heart of the package. Table wraps a table group; its create() / open() classmethods are the two entry points, and its mapping-style access (table["name"], in, iteration) hands out the column wrappers.

The subscript, length and iteration behaviour is listed alongside the ordinary methods: table["name"] hands out a column, and column[17:98] reads rows.

class h5col.Table(group: Any)[source]#

A H5Col column-oriented table backed by an HDF5 group.

static is_table_group(group: Any) → bool[source]#

True if group is a H5Col table group (lenient CLASS check).

Parameters:

group – Any h5py group. One that is not a table group answers False rather than raising, so this is safe to use while walking a file.

classmethod open(group: Any) → Table[source]#

Open an existing table group, checking its CLASS and VERSION major.

Parameters:

group – An h5py group already written as a H5Col table. Opening does not read any column data.

Raises:
  • ConformanceError – If group is not a H5Col table group, or its VERSION is missing or unparsable.

  • VersionError – If the table’s VERSION major exceeds the supported major.

classmethod create(group: Any, columns: TableSpec | Sequence[ColumnSpec | ListColumnSpec], *, title: str | None = None, description: str | None = None, index_columns: Sequence[str] | None = None, column_order: Sequence[str] | None = None, units_vocabulary: str | None = None, encoding_type: str | None = None, encoding_version: str | None = None, default_chunk_bytes: int | None = None) → Table[source]#

Create a new, empty table (NROWS = 0) with the given columns.

Parameters:
  • group – An h5py group to write the table into. It must not already be a H5Col table group.

  • columns – A TableSpec, or a sequence of ColumnSpec and ListColumnSpec.

  • title – Human-readable table title, stored as the TITLE attribute.

  • description – Longer free text, stored as DESCRIPTION.

  • index_columns – Names of the columns that together identify a row, stored as INDEX_COLUMNS. Every name must be a declared scalar column; a list column cannot be an index column.

  • column_order – The columns’ logical order, stored as COLUMN_ORDER. Defaults to the order given in columns, which HDF5 itself does not preserve.

  • units_vocabulary – The vocabulary the columns’ units come from, e.g. UDUNITS-2.

  • encoding_type – Producer-defined encoding identifier for the table as a whole.

  • encoding_version – Version string paired with encoding_type.

  • default_chunk_bytes – Overrides the automatic (chunk-cache-scaled) target chunk size for columns that do not set an explicit chunks shape.

Raises:
  • SchemaError – If group is already a H5Col table group, a column spec is invalid, or an index_columns name is not among the declared columns.

  • ReservedNameError – If a column name is a H5Col reserved name.

  • FillValueError – If a column’s fill value lies inside its declared valid range.

classmethod from_arrays(group: Any, arrays: Mapping[str, Any], *, specs: Sequence[ColumnSpec] | None = None, **table_kwargs: Any) → Table[source]#

Create a table from column arrays and write them in one call.

Parameters:
  • group – An h5py group to write the table into, as for create().

  • arrays – One array per column, keyed by column name. Every array must hold the same number of rows, and the column order follows this mapping.

  • specs – Column specs to create the table with. When omitted, one is inferred per array — boolean, a fixed-length string sized to the longest value, or the array’s own dtype.

  • table_kwargs – Passed through to create(), so title, index_columns and the rest are available here too.

classmethod from_arrow(group: Any, table: Any, *, specs: Sequence[ColumnSpec | ListColumnSpec] | None = None, **table_kwargs: Any) → Table[source]#

Create a table from a pyarrow.Table and write its rows.

Arrow’s model is wider than H5Col’s, so anything without an exact equivalent is refused rather than approximated, see specs_from_arrow(), which decides the mapping and which you can call first to inspect or adjust it:

specs = h5col.specs_from_arrow(tbl)
specs[2].chunks = 8192
Table.from_arrow(group, tbl, specs=specs)

Rows are written batch by batch rather than all at once, so importing a large table costs about one batch of memory rather than the whole of it.

Two chunks of one dictionary column may carry different dictionaries, in which case the same code stands for two different labels. The codes are never read directly for that reason — the labels are, against the unified category set specs_from_arrow() derives.

The fill-value checks run whichever way the specs arrived: a fill that already occurs in a column would leave those rows reading as missing, so supplying specs cannot skip it.

Needs the optional pyarrow dependency (pip install h5col[arrow]).

Added in version 0.4.0.

Parameters:
  • group – An h5py group to write the table into, as for create().

  • table – The pyarrow.Table to import.

  • specs – A complete list of column specs naming exactly the table’s columns. None infers them. Chunking and filters have no Arrow equivalent, so this is the only way to set them.

  • table_kwargs – Passed through to create(), so title, index_columns and the rest are available here too.

Raises:
  • SchemaError – If a column’s type has no H5Col equivalent, if a fill value occurs in its column’s data, if a boolean column holds nulls, or if specs does not name exactly the table’s columns.

  • ReservedNameError – If a column name, or a producer metadata key, is one H5Col reserves.

property group: Any#

The underlying h5py Group backing the table.

property nrows: int#

The table’s logical row count (its NROWS attribute).

Raises:

ConformanceError – If the group carries no NROWS attribute.

property version: str | None#

The table’s H5Col VERSION string, or None when absent.

property title: str | None#

The table’s title attribute, or None when unset.

property description: str | None#

The table’s description attribute, or None when unset.

property generation: int | None#

The table’s GENERATION validity token, or None when absent.

A table acquires GENERATION when its first search index is built and increments it on every subsequent mutation of committed data.

property column_names: list[str]#

Column names in logical order (column-order if present).

property index_columns: list[str]#

Names of the row-index columns, outermost first.

property columns: dict[str, Column | ListColumn]#

The table’s columns by name, in column order, as wrapper objects.

__getitem__(name: str) → Column | ListColumn[source]#

The column named name, as a Column or ListColumn.

Parameters:

name – A column name. Raises KeyError if the table has no such column.

__contains__(name: str) → bool[source]#

True if the table has a column of this name.

Parameters:

name – A column name.

__iter__() → Iterator[str][source]#

Iterate the column names, in the table’s logical column order.

__len__() → int[source]#

The number of columns in the table (its rows are nrows).

append(data: Mapping[str, Any], *, maintain_indexes: bool = False) → None[source]#

Append rows following the H5Col write protocol.

Every provided column must supply the same number of rows K. A scalar column absent from data is extended and left as its fill value (missing); a boolean column (which has no fill) must always be provided. A list column absent from data must be nullable — its new rows become null lists; a non-nullable list column must always be provided. NROWS is committed last, then the file is flushed.

None in a column’s values marks that row as missing and is stored as the column’s fill value (for a categorical column, its fill code). A column with no fill to store — a boolean, which H5Col forbids from declaring one — rejects None instead of coercing it.

By default, search indexes are not maintained: the GENERATION increment that publishes the append disables them, detectably, and refresh_indexes() restores them later — this keeps the hot append path fast. With maintain_indexes=True, every supported index is rewritten inside the append protocol (future-valued tokens before content) and remains valid after the commit; indexes this implementation cannot rebuild — unsupported kinds, element dtypes the builder does not handle, non-growable index datasets — are left entirely untouched, tokens included.

Parameters:
  • data – New rows, one entry per column, keyed by column name. A numpy.ma.MaskedArray is accepted and its masked elements mean the same as None.

  • maintain_indexes – Rebuild the supported search indexes inside the write protocol so they stay valid after the commit, as described above.

Raises:
  • OversizedStringError – If a fixed-length string value’s encoding exceeds the column’s byte budget (H5Col never silently truncates).

  • SchemaError – For unknown columns, values that are not 1-D, unequal column lengths, an omitted fill-less/boolean column, an omitted non-nullable list column, an unknown category label, or a None in a column that declares no fill value.

truncate(nrows: int, *, maintain_indexes: bool = False) → None[source]#

Shrink the logical table to nrows rows (H5Col logical truncation).

The truncation is logical: no column dataset changes extent, and the rows [nrows, old_NROWS) become reserved storage that consumers ignore. List columns need no extra writes — the smaller NROWS bounds their offsets recursively. Reclaiming physical space would require rewriting each column to its new extent, which this implementation does not do.

Index handling mirrors append() (the spec applies the same steps 4-6 with the new row count): by default every search index is left detectably stale by the GENERATION bump; with maintain_indexes=True the supported indexes are rebuilt inside the protocol and remain valid after the commit.

Truncating to the current row count is a no-op (nothing changes, so nothing is published); growing is an error — that is what append() is for.

Parameters:
  • nrows – The new row count. Must be between 0 and the current row count.

  • maintain_indexes – As for append().

Raises:

SchemaError – If nrows is negative or greater than the current row count.

read(columns: Sequence[str] | None = None, *, where: Any = None, explain: bool = False, masked: bool = True) → Any[source]#

Read columns (default all) as {name: array} over [0, NROWS).

Parameters:
  • columns – Names to read, in the order given. None (the default) reads every column of the table.

  • where – Restrict the result to matching rows. Takes the same forms as select(); None (the default) reads every row.

  • explain – When True the return value becomes a (result, QueryPlan) pair, the plan describing how the query was evaluated.

  • masked – Return each scalar column as a numpy.ma.MaskedArray whose mask marks its missing rows (the default). List columns ignore it, already spelling a null row None. Pass False for plain arrays, in which case a missing row holds the column’s fill value with nothing to distinguish it from data.

Raises:
  • KeyError – If a requested column name is not a column of the table.

  • SchemaError – If where= is malformed or references an unknown column.

to_arrow(columns: Sequence[str] | None = None, *, where: Any = None) → Any[source]#

Convert the table (default all columns) to a pyarrow.Table.

The one export that carries the whole H5Col data model: missing rows become real Arrow nulls instead of the fill value, a categorical becomes a dictionary of the codes and labels already stored, a list column keeps its nulls at every level of nesting, and each column’s units, description and valid-range attributes ride along as Arrow field metadata under an h5col. prefix.

Needs the optional pyarrow dependency (pip install h5col[arrow]).

Added in version 0.2.0.

Parameters:
  • columns – Names to convert, in the order given. None (the default) converts every column of the table.

  • where – Restrict the result to matching rows. Takes the same forms as select(); None (the default) converts every row.

Raises:

KeyError – If a requested column name is not a column of the table.

select(where: Any = None) → Selection[source]#

Build a lazy Selection over the table.

Parameters:

where – A query Expression, a List[Tuple] (AND), or a List[List[Tuple]] (OR-of-ANDs, pyarrow DNF). None (the default) selects every row.

count(where: Any = None) → int[source]#

Number of rows matching where, without reading any column values.

Parameters:

where – As for select(). None (the default) counts every row.

build_index(column: str, kind: str | None = None, *, name: str | None = None, description: str | None = None) → SearchIndex[source]#

Build a search index over column (alias of add_search_index()).

Parameters:
info(*, storage: bool = False, full: bool = False) → TableInfo[source]#

Table content information, storage and compression details optional.

The result prints as an aligned block at a prompt and renders as a collapsible view in a Jupyter notebook, and its records stay reachable for anyone who wants the numbers rather than the picture.

Nothing here reads a column’s values, and by default nothing measures stored size or compression either, so even inspecting a large table requires a handful of metadata reads.

Added in version 0.5.0.

Parameters:
  • storage – Whether to measure what each column costs in the file. Off by default because measuring walks every column’s chunk index.

  • full – Whether to report each column as a labelled block of everything known about it: description, valid range, per-dataset sizes, etc. This is the same detail a notebook shows behind a disclosure triangle. What a column costs is part of everything, so this sets storage=True as well.

property search_indexes: dict[str, SearchIndex]#

Every search-index dataset under SEARCH_INDEXES, wrapped by kind.

add_search_index(column: str, kind: str | None = None, *, name: str | None = None, description: str | None = None) → SearchIndex[source]#

Build a search index over column and link it to the column.

With kind=None the family is picked automatically: BITMAP for boolean and categorical columns (low cardinality, exact equality answers), CHUNK_MINMAX for any other orderable column. Building an index over an unchanged table is not a mutation: GENERATION is created (0) if absent but never incremented. The default dataset name is <column>__<kind, lowercased> — a readable convention only; the linkage is the object reference in the column’s SEARCH_INDEX_LIST.

Parameters:
  • column – Name of the column to index. It must be a scalar column; list columns cannot be indexed.

  • kind – CHUNK_MINMAX, SORTED_ROWS or BITMAP. None (the default) picks the family that suits the column’s datatype, as above.

  • name – Name for the index dataset under SEARCH_INDEXES. None uses the <column>__<kind> default described above.

  • description – Free text stored on the index as its DESCRIPTION attribute.

Raises:
  • KeyError – If column is not a column of the table.

  • SchemaError – If column is a list column, no index family applies to its dtype (kind=None), kind is unimplemented, or SEARCH_INDEXES already holds a dataset of the chosen name.

  • ReservedNameError – If name is a H5Col reserved name.

  • ConformanceError – If the table carries no NROWS attribute.

refresh_indexes() → int[source]#

Rebuild every supported search index against the current table state.

Restores indexes left stale by append(maintain_indexes=False) or by any other mutation. Returns the number of indexes refreshed; indexes of unsupported kinds are left untouched (and stay detectably stale).

index_is_valid(index: SearchIndex | Any) → bool[source]#

The H5Col consumer validity check for index.

Parameters:

index – A SearchIndex wrapper or the index dataset itself.

validate(*, deep: bool = False) → None[source]#

Check the H5Col consistency requirements, raising on any violation.

A stale index is never an error: the validity check disables it, as the spec intends.

Parameters:

deep – When True, additionally re-derive every valid search index from its column and compare (consistency rule 9’s semantic half), which costs an index build apiece. The default run is structural only.

Raises:
  • ConformanceError – On the first consistency violation found.

  • VersionError – If the table’s VERSION major exceeds the supported major.

add_column(spec: ColumnSpec | ListColumnSpec, *, default_chunk_bytes: int | None = None) → Column | ListColumn[source]#

Add a new column to an existing table (schema evolution).

A scalar column is created and grown to the table’s current extent; its existing rows read as its fill value (missing). A fill-less (boolean) column, or a list column (whose “missing” analogue is a null list that would need explicit backfilling), cannot represent pre-existing rows, so adding one to a table that already has rows is refused.

Parameters:
  • spec – The new column’s ColumnSpec or ListColumnSpec. Its name must not already be in use.

  • default_chunk_bytes – As for create(): the target chunk size when the spec sets no explicit chunks shape.

Raises:

SchemaError – If a column of that name already exists, the spec is invalid, or the column cannot backfill pre-existing rows (a boolean or list column on a non-empty table).

class h5col.Column(dataset: Any, table: Table)[source]#

One column of a H5Col table, wrapping its rank-1 HDF5 dataset.

Reads decode to friendly Python values: fixed-length strings become str, boolean columns become NumPy bool, and numeric columns pass through.

Rows can be read by subscript — column[17:98], column[-1], column[[3, 1]] — which is read_rows() with its defaults. Note that column[...] and column.dataset[...] are not the same read: the second goes straight to h5py, so it skips decoding and can return rows above NROWS that truncate() left behind as reserved storage.

__len__() → int[source]#

The number of rows in the column, which is the table’s NROWS.

Added in version 0.3.0.

__iter__() → Iterator[Any][source]#

Iterate the decoded rows.

Defined so that iterating reads the column once. Without it Python falls back on __getitem__() and fetches every row separately, which is one HDF5 read per row.

Added in version 0.3.0.

__getitem__(key: Any) → Any[source]#

Read rows by subscript, decoded and masked.

An integer key returns that one row’s value — numpy.ma.masked if the row is missing — and every other key returns an array, the way NumPy behaves.

This is read_rows() with the defaults. Subscript has nowhere to put a keyword, so it always decodes and always masks; call read() or read_rows() when you want masked=False.

Added in version 0.3.0.

Parameters:

key – An integer, a slice, a sequence of positions, or a boolean array with one entry per row. A negative position counts back from the last row.

Raises:
  • IndexError – If a row position is out of range, a boolean mask is the wrong length, or key is a tuple.

  • TypeError – If key holds values that are neither integers nor booleans.

property name: str#

The column’s name (the final component of its HDF5 path).

property dataset: Any#

The underlying h5py Dataset (for advanced/low-level access).

property dtype: dtype#

The column’s NumPy dtype (its stored HDF5 datatype).

For a categorical column this is the integer code dtype, not the category labels (see categories).

property is_boolean: bool#

True if this is an H5Col boolean column.

property is_string: bool#

True if this is a fixed-length string column.

property is_categorical: bool#

True if this is a categorical column (its values are category codes).

property categories: NDArray[Any] | None#

The category labels, or None for a non-categorical column.

String labels come back as a compact NumPy string array; numeric labels keep their own dtype.

property ordered: bool | None#

The categories’ ordered flag, or None if not categorical/unset.

property codes: NDArray[Any]#

Raw integer category codes over [0, NROWS) (categorical columns).

property fill_value: Any#

The column’s fill value, or None when it declares none.

None is returned for boolean columns and for any column not in the H5D_FILL_VALUE_USER_DEFINED state (e.g. a full-domain column that declares no missing-row semantics). h5py’s library-default fill value is not a H5Col sentinel and must not be surfaced as one.

property units: str | None#

The column’s units attribute, or None when unset.

property units_vocabulary: str | None#

The column’s units_vocabulary attribute, or None when unset.

Names the vocabulary units is drawn from — UDUNITS-2, say — so a reader can tell which spelling of a unit was meant. A table declares one for all its columns; this is the per-column override.

Added in version 0.4.0.

property description: str | None#

The column’s description attribute, or None when unset.

property valid_min: Any#

The column’s valid_min attribute, or None when unset.

property valid_max: Any#

The column’s valid_max attribute, or None when unset.

read(*, masked: Literal[True] = True) → MaskedArray[source]#
read(*, masked: Literal[False]) → NDArray[Any]
read(*, masked: bool) → Any

Read the logical rows [0, NROWS), decoded to friendly values.

Parameters:

masked – Return a numpy.ma.MaskedArray whose mask marks the column’s missing rows (the default). Pass False for the plain array, in which case a missing row holds the column’s fill value with nothing to distinguish it from data.

read_rows(rows: Any, *, masked: Literal[True] = True) → MaskedArray[source]#
read_rows(rows: Any, *, masked: Literal[False]) → NDArray[Any]
read_rows(rows: Any, *, masked: bool) → Any

Read just rows, decoded, in the order given.

A slice is read as a single hyperslab. Anything else is fetched with coalesced, chunk-aligned block reads, so a selection confined to a few chunks costs a few chunks rather than the whole column — in both time and peak memory.

Parameters:
  • rows – A slice, a sequence of integer positions, or a boolean mask with one entry per row. A negative position counts back from the end, so -1 is the last row. Integer positions may be given in any order and may repeat; the result follows the order given.

  • masked – As for read().

Raises:
  • IndexError – If a row position is out of range, or a boolean mask is the wrong length.

  • TypeError – If rows holds values that are neither integers nor booleans.

  • ValueError – If rows is not one-dimensional.

to_arrow(rows: Any = None) → Any[source]#

Convert the column to an Arrow array, or just rows of it.

Missing rows become real Arrow nulls rather than the fill value, and a categorical column becomes a DictionaryArray of the codes and labels H5Col already stores — neither of which NumPy can express.

Needs the optional pyarrow dependency (pip install h5col[arrow]).

Added in version 0.2.0.

Parameters:

rows – Which rows to convert, in any form read_rows() accepts. None (the default) converts the whole column.

property filters: FilterPipeline#

The filter pipeline this column’s dataset was created with.

What HDF5 recorded, which is not always what was declared: a plugin may add client data of its own while the dataset is created.

Added in version 0.5.0.

property storage: Storage#

What this column costs in the file, before and after filtering.

Added in version 0.5.0.

property search_indexes: list[SearchIndex]#

Search indexes bound to this column, from SEARCH_INDEX_LIST.

The wrappers are bound to this column, so their queries always run against it — even on a non-conformant file where another column also claims the same index dataset.

Raises:

ConformanceError – If the column’s SEARCH_INDEX_LIST attribute is malformed (not a 1-D array of object references).

add_search_index(kind: str | None = None, *, name: str | None = None, description: str | None = None) → SearchIndex[source]#

Build a search index over this column (Table.add_search_index()).

Parameters:
  • kind – Which index family to build — CHUNK_MINMAX, SORTED_ROWS or BITMAP. None (the default) picks the family that suits the column’s datatype.

  • name – Name for the index dataset under SEARCH_INDEXES. None derives one from the column name and the kind.

  • description – Free text stored on the index as its DESCRIPTION attribute.

Raises:
  • SchemaError – If no index family applies to the column’s dtype (kind=None), kind is unimplemented, or SEARCH_INDEXES already holds a dataset of the chosen name.

  • ReservedNameError – If name is a H5Col reserved name.

  • ConformanceError – If the table carries no NROWS attribute.

build_index(kind: str | None = None, *, name: str | None = None, description: str | None = None) → SearchIndex[source]#

Build a search index over this column (alias of add_search_index()).

Parameters:
is_missing() → NDArray[bool][source]#

Boolean mask of missing rows over [0, NROWS).

A column with no user-defined fill value (boolean columns, or full-domain columns per H5Col) declares no missing-row semantics, so every row reads as present.

class h5col.ListColumn(group: Any, table: Table)[source]#

One list column of a H5Col table, wrapping its CLASS=LIST_COLUMN group.

Reading returns one Python list per row (or None for a null list), with elements decoded to friendly values — nested lists become nested Python lists, string elements become str, and missing leaf elements become None.

__len__() → int[source]#

The number of rows in the column, which is the table’s NROWS.

Added in version 0.3.0.

__iter__() → Iterator[Any][source]#

Iterate the rows, each a list or None.

Added in version 0.3.0.

__getitem__(key: Any) → Any[source]#

Read rows by subscript, the same keys a scalar column takes.

An integer returns one row — a list, or None if the row is null — and every other key returns a list of rows. Negative positions count back from the last row.

Only the rows asked for are read; see read_rows() for what that means when the positions are scattered rather than a range.

Added in version 0.3.0.

Parameters:

key – An integer, a slice, a sequence of positions, or a boolean array with one entry per row.

Raises:
  • IndexError – If a row position is out of range, a boolean mask is the wrong length, or key is a tuple.

  • TypeError – If key holds values that are neither integers nor booleans.

property name: str#

The list column’s name (the final component of its HDF5 path).

property group: Any#

The underlying h5py Group (for advanced/low-level access).

property nullable: bool#

True if the top level carries a MASK (null lists are possible).

property units: str | None#

The list column’s units attribute, or None when unset.

property units_vocabulary: str | None#

The list column’s units_vocabulary attribute, or None when unset.

property description: str | None#

The list column’s description attribute, or None when unset.

property filters: dict[str, FilterPipeline]#

The filter pipeline of each dataset making up this column.

A list column is several datasets rather than one, so this answers with a pipeline per dataset, keyed by the dataset’s path within the column group: OFFSETS, VALUES, MASK, and for a column of strings VALUES/OFFSETS and VALUES/CHARS. A column of lists of lists nests the same way.

Added in version 0.5.0.

property storage: Storage#

What this column costs in the file, its datasets added together.

Added in version 0.5.0.

read(*, masked: bool = True) → list[Any][source]#

Read rows [0, NROWS) as a list of per-row lists (None = null).

Parameters:

masked – Accepted and ignored, so a caller can pass it uniformly across a table’s columns. A list column is ragged and so cannot be a numpy.ma.MaskedArray; it already spells a null row None.

read_rows(rows: Any, *, masked: bool = True) → list[Any][source]#

Read just rows, in the order given, as a list of (list | None).

A contiguous range is read directly. Scattered positions are served from the range that spans them, from the lowest wanted row to the highest, which is never wider than the column itself: rows that sit near each other cost almost nothing, and rows spread from end to end cost what reading the column costs.

Added in version 0.3.0.

Parameters:
  • rows – The same row specs Column.read_rows() takes — a slice, a sequence of positions, or a boolean mask — so a caller can select rows the same way whatever kind of column it holds.

  • masked – Accepted and ignored, as it is by read().

is_missing() → NDArray[bool][source]#

Boolean mask of null-list rows over [0, NROWS).

A list column with no top-level MASK cannot mark a row missing, so every row reads as present.

Arrow interchange#

Table.to_arrow exports a table and Table.from_arrow imports one, deciding each column’s spec for itself. The function below returns those specs without writing anything, so they can be looked at and adjusted first. It is also the only way to set chunking and filters, which have no Arrow equivalent to be inferred from. The guide chapter covers the whole picture.

h5col.specs_from_arrow(table: Any) → list[Any][source]#

The column specs this package would import an Arrow table with.

Returned so they can be inspected and adjusted before anything is written. Chunking and filters have no Arrow equivalent and so cannot be inferred at all. The string widths and category sets are inferred from the data, which is worth checking on a table you did not write:

specs = h5col.specs_from_arrow(tbl)
specs[2].chunks = 8192
specs[2].filters = FilterPipeline([Shuffle(), Deflate(4)])

Nothing is read from or written to a file here.

Arrow’s model is wider than H5Col’s, so a type with no exact H5Col equivalent is refused by name rather than approximated: timestamps, dates, times, durations, decimals, structs, maps, unions, opaque binary and fixed-size lists.

Field metadata under h5col. becomes the column’s own annotations, the same keys column_metadata() writes. Any other metadata is carried across as a producer attribute, unless its name is one H5Col reserves.

Arrow marks a missing value with a null; H5Col marks one with a value drawn from the column’s own domain. The specs come back with fill_value unset, which means the recommended value for the datatype, as it does anywhere else. Whether that value is safe for this data is not decided here — a fill that already occurs in a column is refused when the table is written, since setting one is the remedy and it has to be possible to get the specs in order to set it:

specs = h5col.specs_from_arrow(tbl)
specs[0].fill_value = 254
h5col.Table.from_arrow(group, tbl, specs=specs)

Added in version 0.4.0.

Parameters:

table – A pyarrow.Table.

Raises:
  • SchemaError – If two fields share a name (HDF5 links are unique), if a column’s type has no H5Col equivalent, or if a h5col. metadata key is not one this importer understands.

  • ReservedNameError – If a column name, or a producer metadata key, is a name H5Col reserves.