Writing an Arrow table into H5Col#
Arrow is where a lot of tabular data already is. Parquet reads into it, DuckDB and Polars hand it out, pandas converts to it, and a table that arrives over the wire during interprocess communication is Arrow before it is anything else. So the other direction matters as much as the export to Arrow.
To store a pyarrow.Table as an H5Col table is one call:
import h5py
import h5col
with h5py.File("weather.h5", "w") as f:
table = h5col.Table.from_arrow(f.create_group("observations"), arrow_tbl)
The rows are written batch by batch rather than all at once, so importing a large table costs about one batch of memory.
What cannot come across#
Arrow’s datatype system is richer than H5Col’s, and the differences cannot be silently handled. A timestamp column has no H5Col equivalent, and neither do dates, times, durations, decimals, structs, maps, or unions.
None of these are approximated. Each is refused with an exception which recommends an alternative:
SchemaError: column 'observed_at': Arrow type timestamp[us] cannot be stored
in H5Col — H5Col has no datetime type; store the epoch offsets as an integer
column and describe the encoding in its attributes
Silently storing a timestamp as an integer would produce a file that reads back
as numbers no one can interpret. Converting the column yourself, and saying in
its description or some other attributes what the numbers mean, produces one
that survives the trip.
A fixed-size list is refused for a different reason, and it is worth knowing
which. HDF5 has the datatype for it: an array datatype (H5T_ARRAY) holds a
fixed count of elements per row. This is what Arrow’s fixed_size_list is, and
the convention lets a column dataset carry any HDF5 datatype. However, h5py
cannot set a fill value for this datatype, which is how H5Col marks a missing
row. Converting such an Arrow column to an H5Col variable list type is left for
the user rather than silently done, because it drops a guarantee: a list column
does not fix its row lengths, so nothing afterwards holds the imported column to
a fixed count.
Variable-length binary is refused for a third reason, and this one cannot be
worked around at all. Padding the blobs to a common width would not survive the
trip: no byte is safe to strip from the end of a blob, since any byte can
legitimately be part of one. What is missing is the length, and no padding
scheme carries it. Convert to fixed_size_binary if the values do share a
width — that maps exactly, as the next section explains.
Raw bytes#
fixed_size_binary[n] becomes an opaque column: n bytes per row, stored
back to back with no offsets, which is byte-for-byte what Arrow holds. The
buffer is handed over rather than converted, in both directions.
Opaque columns raise the missing-value question in its sharpest form. A fill
value has to be a value the column’s data will not contain, and for raw bytes
there is no such value in principle — any byte string might be real data. H5Col
picks one that is merely very unlikely: the ASCII marker FILL followed by
rising byte values, so an eight-byte column’s fill is
46 49 4c 4c 01 02 03 04 "FILL...."
The rising tail is the part that earns its keep. All zeros and all 0xFF are
what zero padding, erased flash and uninitialized memory leave behind, and a
fill made of either would collide constantly. A counting sequence is something
data almost never is.
Unlikely is not impossible, so the collision check applies here as everywhere else: a column that does contain the pattern is refused, and you name a fill of your own through the specs. One case deserves care — a one-byte opaque column has only 256 possible values and the recommended fill claims one of them, which is a real risk rather than a remote one.
What has no home is Arrow’s variable-length binary. With no fixed width, there
is nothing to size a column to.
Nulls have to become values#
This is the part worth understanding before importing anything.
Arrow marks a missing value with a null: a bit in a separate buffer, holding no opinion about what the value would have been. H5Col marks one with a fill value drawn from the column’s own datatype domain. The two models do not line up on their own, so every column that can hold a null needs a fill value chosen for it.
from_arrow chooses the recommended fill for the column’s datatype, and then
checks it against the data. If the fill already occurs as a real value, the
import is refused:
SchemaError: column 'delta': the recommended fill value np.int8(-127) occurs
in the data, so those rows would read as missing; np.int8(-126) does not
occur — pass a spec with that as its fill_value if it lies outside the
column's logical range
This check is worth the trouble it occasionally causes. Without it you get a
file that passes validate(deep=True) and reads back with rows quietly
missing that were never missing at all, the kind of error that surfaces
months later in someone’s analysis.
The suggested value is a genuine offer and not a decision made for you. Note
what it does not claim: a value absent from the data is not the same as a
value outside the column’s logical range, which is what H5Col asks for. If
-126 is a reading your instrument can produce, take a different one or widen
the datatype, which is what the convention prescribes for a column with no value
to spare. That is also what the message says when nothing near the limits of the
type is free:
... nor is any other value near the limits of uint8; a column using its whole
datatype has no value left to mark absence, so widen the datatype, which is
what H5Col prescribes for this
Choosing a fill value on your behalf would be the wrong kind of helpful. The
recommended value is documented, so a reader knows what a uint8 column’s fill
is without opening the file; one picked per import from whatever the data
happened to contain is knowable only after the fact. And if a later append
brought in a value we had quietly reserved, those rows would start reading as
missing.
The same rule applies inside a list column, at every level of nesting: a null element becomes the leaf’s fill value, so a leaf that already contains that value is refused, too.
Two cases have no fill available at all:
Boolean columns. H5Col forbids a boolean from declaring a fill value, so a boolean column holding nulls is refused. Drop the nulls, or import the column as an integer.
String columns containing an empty string. The recommended fill for a string column is the empty string, so a column holding one needs a different fill named explicitly. That is what the specs below are for.
Taking the specs into your own hands#
Two things about a H5Col column have no Arrow equivalent whatsoever: chunking and filters. Nothing in an Arrow schema says how the data should be laid out on disk or what should compress it, and those are exactly the decisions that make a stored column fast or slow to read.
So from_arrow on its own gives you an unfiltered table with default chunking.
To decide otherwise, ask what it would do, adjust, and hand it back:
specs = h5col.specs_from_arrow(arrow_tbl)
for spec in specs:
spec.chunks = 65_536
spec.filters = h5col.FilterPipeline([h5col.Shuffle(), h5col.Deflate(4)])
table = h5col.Table.from_arrow(group, arrow_tbl, specs=specs)
Nothing is read from or written to a file by specs_from_arrow, so this is a
cheap thing to do first and examine its findings. It is also where you set a
fill value the data does not contain, widen a string column beyond what its
current values need, or fix anything else the inference got merely defensible
rather than right.
specs_from_arrow answers what the columns would look like, and stops there.
It does not check a fill value against the data, which is deliberate: setting
one is the way out of a collision, so getting hold of the specs cannot itself
be the thing that fails. A spec whose fill_value is unset means the
recommended value for its datatype, as it does anywhere else in the package.
What you cannot do is skip the check. It runs when the table is written, whichever way the specs arrived, because a fill value that occurs in its own column is the one importing mistake that produces a conformant file with unreadable rows.
What the inference reads from the data#
Most of a column’s spec comes from its Arrow type. Three things cannot:
String widths. H5Col string columns are fixed-length; Arrow’s are not. The column is scanned and sized to its widest value, which means a later
appendof a longer string raises rather than truncating. If you expect longer values later, set a wider budget in the spec.Category labels. A dictionary column’s labels are collected across all its chunks, since two chunks of one Arrow column may carry different dictionaries. The dictionary’s index type is kept as the code datatype, so a column that arrives with
int32codes is stored withint32codes.Which levels of a list column hold nulls. A
MASKdataset is created only at the levels that actually need one.
Metadata#
Field metadata beginning with h5col. is read back as the column’s own
annotations: units, units_vocabulary, description, valid_min,
valid_max and ordered. These are the same keys the export writes, so a
table that made the trip in the other direction arrives with its descriptions
intact.
Any other field metadata is carried across as an ordinary HDF5 attribute on the
column, so a producer’s own keys are kept rather than dropped. The one
restriction is that such a key may not be an H5Col reserved name, or one it
writes itself. Such filed metadata raises an
ReservedNameError.
Column names get the same treatment. A name Arrow permits but HDF5 cannot store as a link, or one H5Col reserves, is refused rather than mangled into something storable.
The round trip#
Export, import, and export again gives back an equal Arrow table — same types, same values, same field metadata — for every column kind H5Col can store. The test suite checks this per kind, which is the honest way to state what the two directions guarantee together.
What it does not claim is that the first export is unchanged by the trip.
Arrow has types H5Col stores as something close but not identical: a string
column comes back as large_string, and a list as large_list, because
those are the layouts H5Col’s own storage matches. The pair settles after one
hop, and stays there.
pyarrow is not required to use h5col. Install it alongside if you want to
import or export Arrow tables:
pip install h5col[arrow]