Reading into Python#
An H5Col table holds more than a plain NumPy array can express. Some rows are missing. A categorical column stores small integer codes that stand for labels. A list column holds a different number of values in every row. When you read a table, all of that has to arrive in some Python object, and no single object is the right answer for every purpose.
So h5col offers two, and it is worth knowing which one you want before you
start:
Table.readgives a dictionary of NumPy arrays. It needs nothing beyond NumPy and it is the right default for almost everything.Table.to_arrowgives an Apache Arrow table. It is the form that carries everything the table holds, and it is the bridge to pandas, Polars, DuckDB and Parquet. It needs the optionalpyarrowpackage.
This chapter goes through what each one hands back, column type by column type, then how to read part of a column rather than all of it, and finally where each form falls short.
Reading into NumPy#
Take a small table of weather observations, with a missing temperature, a missing station kind, and a mix of list values:
table.append({
"station": ["KBOS", "KJFK", "KLGA"],
"t_air": [21.5, None, 23.1],
"kind": ["manned", "automatic", None],
"checked": [True, True, False],
"samples": [[1.0, 2.0], None, []],
})
read() returns one entry per column:
Column type |
What you get |
|---|---|
numeric |
a masked array of the column’s own dtype |
boolean |
a masked array of |
fixed-length string |
a masked array of |
categorical |
a masked array of the labels, not the codes |
list column |
a Python list, one entry per row |
Every column except a list column comes back as a
numpy.ma.MaskedArray. That is the part most worth explaining.
Why missing values arrive masked#
A missing row is stored as the column’s fill value: for t_air above, that is
-999. If reading simply handed you the stored numbers, you would get this:
table["t_air"].read(masked=False)
array([ 21.5, -999. , 23.1], dtype=float32)
Nothing in that array says the middle value is not a real measurement. Take the
average and you get -318.1, which is not a temperature anyone recorded. The
mistake is easy to make and produces a plausible-looking number, which is the
worst kind of mistake.
By default you get the same values with a mask alongside them:
table["t_air"].read()
masked_array(data=[21.5, --, 23.100000381469727],
mask=[False, True, False],
fill_value=-999.0,
dtype=float32)
Now the average is 22.3, because NumPy skips the masked entry. Sums, counts,
minimums and the rest all behave the same way. You did not have to remember
anything, which is the point.
The mask itself is an ordinary boolean array, True where the row is missing:
table["t_air"].read().mask
array([False, True, False])
It always agrees with Column.is_missing(),
which is the same question asked directly.
Every column is masked, even when it cannot be#
Boolean columns cannot have missing rows because the H5Col convention does not
allow a fill value for them. And some columns simply have no missing rows in
them. Those still come back as masked arrays, with a mask that is False
everywhere.
This is deliberate. If the type of result["checked"] depended on whether that
particular column could be missing, then any code looping over the columns of a
table would have to check before it could do anything, and would break the
first time it met a table whose author made a different choice. One type per
kind of column is easier to write against than one type per column.
Turning it off#
Pass masked=False to get plain arrays:
table.read(masked=False)
Missing rows then hold the fill value, with nothing to mark them, and it is
back to you to call is_missing() and apply it. The keyword
works the same way on
Column.read,
Column.read_rows,
Table.read and
Selection.read.
Two habits worth picking up#
Use .tolist(), not list(). They differ, and only one of them does what
you probably want:
table["t_air"].read().tolist()
[21.5, None, 23.100000381469727]
list(table["t_air"].read())
[np.float32(21.5), masked, np.float32(23.1)]
.tolist() turns a missing row into None. Plain list() gives NumPy’s
masked marker, which is not None and will not compare equal to it. If you
have code that tests if value is None, reach for .tolist().
Know which NumPy functions keep the mask. A good many do not.
np.concatenate, np.stack, np.append and np.where all return something
that is still a masked array by type but has quietly lost its mask, and
np.asarray hands back the underlying values. NumPy provides masked versions —
np.ma.concatenate and friends — and those are the ones to use.
When the mask does get dropped, what you are left with is the stored values,
fill value and all. That is the same thing masked=False would have given you,
so a lost mask never leaves you worse off than not having asked for one. It is
still worth avoiding.
Getting the plain values back#
If plain fill values are deliberately needed, .filled() is the proper way to get them:
table["t_air"].read().filled()
array([ 21.5, -999. , 23.1], dtype=float32)
Default NumPy printing of masked array values shows all the digits, whereas for
plain arrays it adjusts based on the dtype. The numbers are identical and only
the display changes, but if you are showing values to somebody, rounding first
is worth the trouble. .tolist() behaves the same way.
One small oddity: for a column with no missing rows at all, the
fill_value shown in the array’s display is a NumPy placeholder such as
'N/A' rather than the column’s own. It is never used — .filled() on an
array with nothing masked returns the values untouched — so it is display noise
rather than anything to act on.
Strings#
A fixed-length string column reads back as real Python strings, held in a NumPy
array of numpy.dtypes.StringDType:
table["station"].read().tolist()
['KBOS', 'KJFK', 'KLGA']
StringDType keeps the text packed in one block of memory instead of building a
separate Python string object for every row. It supports comparing, sorting, and
finding unique values. However, one consequence of the NumPy implementation is
that the text is only checked for valid UTF-8 when reading specific values out,
not when the column is read into the array. If a string contains invalid
bytes, the error appears only when that particular value is accessed. String
columns written by h5col cannot get into this state because there are checks
on the way in.
Categorical columns#
You get labels, not codes:
table["kind"].read().tolist()
['manned', 'automatic', None]
The codes are still there if you want them, as
Column.codes, and the label set as
Column.categories.
There is one place where categoricals do not follow the general rule.
.filled() on other column types puts the column’s fill value into the masked
slots; for a categorical there is no label to put there, because a missing row
has no category. .tolist() gives you None, which is the answer you want.
List columns#
A list column reads back as a plain Python list, one entry per row:
table["samples"].read()
[[np.float64(1.0), np.float64(2.0)], None, []]
There is no masked array here, and there cannot be. Rows hold different numbers
of values, and a NumPy array needs them all to be the same shape. A list column
already reports missing rows in the clearest way available: the row is
None. Note that None and [] mean different things. None is a row
with no value at all; [] is a row whose value is a list that happens to be
empty. The list columns chapter goes into that distinction.
The values inside each row are NumPy scalars rather than plain Python numbers,
which is why the output above reads np.float64(1.0) and not 1.0. They
compare and calculate exactly as expected.
These columns accept the masked keyword and ignore it, so you can pass it
across a whole table without special-casing.
Reading into Arrow#
Three things a table can hold have no NumPy equivalent at all:
a missing value that is genuinely absent, rather than a particular number standing in for absence;
a categorical column with its complete semantics — a small set of labels plus one code per row, rather than expanded to a full label for every row;
a list column, with its own missing values at every level of nesting.
Arrow enables preserving the entire H5Col table semantics. Table.to_arrow gives you a
pyarrow.Table:
table.to_arrow()
station: large_string
t_air: float
kind: dictionary<values=string, indices=int8, ordered=0>
checked: bool
samples: large_list<item: double>
----
station: [["KBOS","KJFK","KLGA"]]
t_air: [[21.5,null,23.1]]
kind: [ -- dictionary: ["manned","automatic"] -- indices: [0,1,null]]
checked: [[true,true,false]]
samples: [[[1,2],null,[]]]
The missing temperature is null, not -999. The kind column is still a
dictionary of two labels with one code per row. The samples column keeps its
rows and its missing row.
Each column’s units, units_vocabulary, description and valid-range
attributes travel along as Arrow field metadata, under names beginning with
h5col., and they survive being written to Parquet and read back. So a table
exported this way does not lose the descriptions that made it understandable.
The most of the tabular ecosystem is just one call away from an Arrow table:
.to_pandas(), Polars, DuckDB, pyarrow.parquet.write_table. Arrow is also the
faster path for list columns, by a wide margin. h5col stores them in nearly
the layout Arrow uses. Where reading a list column into Python lists has to
build every row as an object, the Arrow export mostly hands the same blocks of
memory straight over. On a column of two hundred thousand rows that is roughly
twenty to thirty times faster.
The trip runs both ways: Table.from_arrow
writes a pyarrow.Table as an H5Col table.
pyarrow is not required to use h5col. Install it alongside if you want this
data export feature:
pip install h5col[arrow]
Reading part of a column#
Everything so far has read whole columns. When you want some of one, say so when you ask, rather than reading it all and slicing the result — the difference is real, and on a large column it is not small.
Column.read_rows takes whatever describes the
rows you want — a slice, a list of positions, or a boolean array with one entry
per row — and subscript is the shorter spelling of the same thing:
col = table["t_air"]
col[17:98] # a range
col[-1] # one value, not an array of one
col[[3, 1, 3]] # any order, repeats allowed
col[col.is_missing()] # only the missing rows
A range is read in one pass, so asking for part of a column really is cheaper
than asking for all of it. A negative position counts back from the last row,
the way it does in a slice. An integer key returns that row’s value on its own
— numpy.ma.masked if the row is missing — while every other key returns an
array, which is how NumPy behaves.
Subscript has nowhere to put a keyword, so it always decodes and always masks.
Reach for read or read_rows when you want masked=False.
For a whole table rather than one column, select the rows first with
Table.select and read from the selection; the
queries section covers that properly.
Ranges in a list column#
List columns take the same keys and also read only the rows asked for, though
they get there differently. A list column’s rows are reached through its
OFFSETS, so a range read looks up where that range’s values begin and end and
reads only that span, at every level of nesting. Scattered positions are served
from the range that spans them — the lowest wanted row to the highest — which
is never wider than the column itself, so rows near each other cost almost
nothing and rows at opposite ends cost what the whole column costs.
Reading fifty rows of a 200,000-row list column pulls 251 values rather than a million, and takes about 1.5 ms rather than 79 ms.
Where to_arrow fits#
to_arrow is a different trade. It reads the whole column, so it saves no
reading at all, but it hands the stored buffers straight to Arrow rather than
building a Python list for every row — around 4 ms for that same 200,000-row
column. It suits wanting most of a large column, where a range read suits
wanting a small part of one.
column[...] is not column.dataset[...]#
The two are a letter apart and are not the same read. The second goes straight
to h5py, so it skips decoding, ignores missing values, and can hand back rows
above NROWS that truncate() left behind as reserved
storage. Use it when you want the stored bytes and nothing else.
Where these forms fall short#
Worth knowing to avoid any surprises.
An empty string reads as missing. The default fill value for a string
column is the empty string, so a row genuinely containing "" cannot be told
apart from a row with no value. If empty strings are real data in your model,
choose a different fill value when creating the column. The
missing values chapter covers the choice.
A mask can be lost quietly. np.concatenate, np.stack, np.append and
np.where all return arrays that have silently dropped it. The np.ma
equivalents do not.
List columns cannot carry a mask, so a table read into NumPy is not uniform: most columns are masked arrays, list columns are Python lists. Arrow does not have this split.
A read hands back data, not a handle. An h5py dataset is a lazy thing: you
can hold one, slice it later, and nothing is read until you do. What read()
and to_arrow() return is the opposite — every row, already in memory, before
the call comes back. If a column is larger than you want to hold, ask for part
of it up front, as reading part of a column
describes, rather than reading it whole and slicing afterwards.
Arrow needs a dependency. It is optional on purpose: a file written by
h5col can be read with nothing but HDF5, and requiring a large package for
the base case would undercut that.