Reading into Python#

An H5Col table holds more than a plain NumPy array can express. Some rows are missing. A categorical column stores small integer codes that stand for labels. A list column holds a different number of values in every row. When you read a table, all of that has to arrive in some Python object, and no single object is the right answer for every purpose.

So h5col offers two, and it is worth knowing which one you want before you start:

  • Table.read gives a dictionary of NumPy arrays. It needs nothing beyond NumPy and it is the right default for almost everything.

  • Table.to_arrow gives an Apache Arrow table. It is the form that carries everything the table holds, and it is the bridge to pandas, Polars, DuckDB and Parquet. It needs the optional pyarrow package.

This chapter goes through what each one hands back, column type by column type, then how to read part of a column rather than all of it, and finally where each form falls short.

Reading into NumPy#

Take a small table of weather observations, with a missing temperature, a missing station kind, and a mix of list values:

table.append({
    "station": ["KBOS", "KJFK", "KLGA"],
    "t_air":   [21.5, None, 23.1],
    "kind":    ["manned", "automatic", None],
    "checked": [True, True, False],
    "samples": [[1.0, 2.0], None, []],
})

read() returns one entry per column:

Column type

What you get

numeric

a masked array of the column’s own dtype

boolean

a masked array of bool

fixed-length string

a masked array of numpy.dtypes.StringDType

categorical

a masked array of the labels, not the codes

list column

a Python list, one entry per row

Every column except a list column comes back as a numpy.ma.MaskedArray. That is the part most worth explaining.

Why missing values arrive masked#

A missing row is stored as the column’s fill value: for t_air above, that is -999. If reading simply handed you the stored numbers, you would get this:

table["t_air"].read(masked=False)
array([  21.5, -999. ,   23.1], dtype=float32)

Nothing in that array says the middle value is not a real measurement. Take the average and you get -318.1, which is not a temperature anyone recorded. The mistake is easy to make and produces a plausible-looking number, which is the worst kind of mistake.

By default you get the same values with a mask alongside them:

table["t_air"].read()
masked_array(data=[21.5, --, 23.100000381469727],
             mask=[False,  True, False],
       fill_value=-999.0,
            dtype=float32)

Now the average is 22.3, because NumPy skips the masked entry. Sums, counts, minimums and the rest all behave the same way. You did not have to remember anything, which is the point.

The mask itself is an ordinary boolean array, True where the row is missing:

table["t_air"].read().mask
array([False,  True, False])

It always agrees with Column.is_missing(), which is the same question asked directly.

Every column is masked, even when it cannot be#

Boolean columns cannot have missing rows because the H5Col convention does not allow a fill value for them. And some columns simply have no missing rows in them. Those still come back as masked arrays, with a mask that is False everywhere.

This is deliberate. If the type of result["checked"] depended on whether that particular column could be missing, then any code looping over the columns of a table would have to check before it could do anything, and would break the first time it met a table whose author made a different choice. One type per kind of column is easier to write against than one type per column.

Turning it off#

Pass masked=False to get plain arrays:

table.read(masked=False)

Missing rows then hold the fill value, with nothing to mark them, and it is back to you to call is_missing() and apply it. The keyword works the same way on Column.read, Column.read_rows, Table.read and Selection.read.

Two habits worth picking up#

Use .tolist(), not list(). They differ, and only one of them does what you probably want:

table["t_air"].read().tolist()
[21.5, None, 23.100000381469727]
list(table["t_air"].read())
[np.float32(21.5), masked, np.float32(23.1)]

.tolist() turns a missing row into None. Plain list() gives NumPy’s masked marker, which is not None and will not compare equal to it. If you have code that tests if value is None, reach for .tolist().

Know which NumPy functions keep the mask. A good many do not. np.concatenate, np.stack, np.append and np.where all return something that is still a masked array by type but has quietly lost its mask, and np.asarray hands back the underlying values. NumPy provides masked versions — np.ma.concatenate and friends — and those are the ones to use.

When the mask does get dropped, what you are left with is the stored values, fill value and all. That is the same thing masked=False would have given you, so a lost mask never leaves you worse off than not having asked for one. It is still worth avoiding.

Getting the plain values back#

If plain fill values are deliberately needed, .filled() is the proper way to get them:

table["t_air"].read().filled()
array([  21.5, -999. ,   23.1], dtype=float32)

Default NumPy printing of masked array values shows all the digits, whereas for plain arrays it adjusts based on the dtype. The numbers are identical and only the display changes, but if you are showing values to somebody, rounding first is worth the trouble. .tolist() behaves the same way.

One small oddity: for a column with no missing rows at all, the fill_value shown in the array’s display is a NumPy placeholder such as 'N/A' rather than the column’s own. It is never used — .filled() on an array with nothing masked returns the values untouched — so it is display noise rather than anything to act on.

Strings#

A fixed-length string column reads back as real Python strings, held in a NumPy array of numpy.dtypes.StringDType:

table["station"].read().tolist()
['KBOS', 'KJFK', 'KLGA']

StringDType keeps the text packed in one block of memory instead of building a separate Python string object for every row. It supports comparing, sorting, and finding unique values. However, one consequence of the NumPy implementation is that the text is only checked for valid UTF-8 when reading specific values out, not when the column is read into the array. If a string contains invalid bytes, the error appears only when that particular value is accessed. String columns written by h5col cannot get into this state because there are checks on the way in.

Categorical columns#

You get labels, not codes:

table["kind"].read().tolist()
['manned', 'automatic', None]

The codes are still there if you want them, as Column.codes, and the label set as Column.categories.

There is one place where categoricals do not follow the general rule. .filled() on other column types puts the column’s fill value into the masked slots; for a categorical there is no label to put there, because a missing row has no category. .tolist() gives you None, which is the answer you want.

List columns#

A list column reads back as a plain Python list, one entry per row:

table["samples"].read()
[[np.float64(1.0), np.float64(2.0)], None, []]

There is no masked array here, and there cannot be. Rows hold different numbers of values, and a NumPy array needs them all to be the same shape. A list column already reports missing rows in the clearest way available: the row is None. Note that None and [] mean different things. None is a row with no value at all; [] is a row whose value is a list that happens to be empty. The list columns chapter goes into that distinction.

The values inside each row are NumPy scalars rather than plain Python numbers, which is why the output above reads np.float64(1.0) and not 1.0. They compare and calculate exactly as expected.

These columns accept the masked keyword and ignore it, so you can pass it across a whole table without special-casing.

Reading into Arrow#

Three things a table can hold have no NumPy equivalent at all:

  • a missing value that is genuinely absent, rather than a particular number standing in for absence;

  • a categorical column with its complete semantics — a small set of labels plus one code per row, rather than expanded to a full label for every row;

  • a list column, with its own missing values at every level of nesting.

Arrow enables preserving the entire H5Col table semantics. Table.to_arrow gives you a pyarrow.Table:

table.to_arrow()
station: large_string
t_air: float
kind: dictionary<values=string, indices=int8, ordered=0>
checked: bool
samples: large_list<item: double>
----
station: [["KBOS","KJFK","KLGA"]]
t_air: [[21.5,null,23.1]]
kind: [  -- dictionary: ["manned","automatic"]  -- indices: [0,1,null]]
checked: [[true,true,false]]
samples: [[[1,2],null,[]]]

The missing temperature is null, not -999. The kind column is still a dictionary of two labels with one code per row. The samples column keeps its rows and its missing row.

Each column’s units, units_vocabulary, description and valid-range attributes travel along as Arrow field metadata, under names beginning with h5col., and they survive being written to Parquet and read back. So a table exported this way does not lose the descriptions that made it understandable.

The most of the tabular ecosystem is just one call away from an Arrow table: .to_pandas(), Polars, DuckDB, pyarrow.parquet.write_table. Arrow is also the faster path for list columns, by a wide margin. h5col stores them in nearly the layout Arrow uses. Where reading a list column into Python lists has to build every row as an object, the Arrow export mostly hands the same blocks of memory straight over. On a column of two hundred thousand rows that is roughly twenty to thirty times faster.

The trip runs both ways: Table.from_arrow writes a pyarrow.Table as an H5Col table.

pyarrow is not required to use h5col. Install it alongside if you want this data export feature:

pip install h5col[arrow]

Reading part of a column#

Everything so far has read whole columns. When you want some of one, say so when you ask, rather than reading it all and slicing the result — the difference is real, and on a large column it is not small.

Column.read_rows takes whatever describes the rows you want — a slice, a list of positions, or a boolean array with one entry per row — and subscript is the shorter spelling of the same thing:

col = table["t_air"]
col[17:98]            # a range
col[-1]               # one value, not an array of one
col[[3, 1, 3]]        # any order, repeats allowed
col[col.is_missing()] # only the missing rows

A range is read in one pass, so asking for part of a column really is cheaper than asking for all of it. A negative position counts back from the last row, the way it does in a slice. An integer key returns that row’s value on its own — numpy.ma.masked if the row is missing — while every other key returns an array, which is how NumPy behaves.

Subscript has nowhere to put a keyword, so it always decodes and always masks. Reach for read or read_rows when you want masked=False.

For a whole table rather than one column, select the rows first with Table.select and read from the selection; the queries section covers that properly.

Ranges in a list column#

List columns take the same keys and also read only the rows asked for, though they get there differently. A list column’s rows are reached through its OFFSETS, so a range read looks up where that range’s values begin and end and reads only that span, at every level of nesting. Scattered positions are served from the range that spans them — the lowest wanted row to the highest — which is never wider than the column itself, so rows near each other cost almost nothing and rows at opposite ends cost what the whole column costs.

Reading fifty rows of a 200,000-row list column pulls 251 values rather than a million, and takes about 1.5 ms rather than 79 ms.

Where to_arrow fits#

to_arrow is a different trade. It reads the whole column, so it saves no reading at all, but it hands the stored buffers straight to Arrow rather than building a Python list for every row — around 4 ms for that same 200,000-row column. It suits wanting most of a large column, where a range read suits wanting a small part of one.

column[...] is not column.dataset[...]#

The two are a letter apart and are not the same read. The second goes straight to h5py, so it skips decoding, ignores missing values, and can hand back rows above NROWS that truncate() left behind as reserved storage. Use it when you want the stored bytes and nothing else.

Where these forms fall short#

Worth knowing to avoid any surprises.

An empty string reads as missing. The default fill value for a string column is the empty string, so a row genuinely containing "" cannot be told apart from a row with no value. If empty strings are real data in your model, choose a different fill value when creating the column. The missing values chapter covers the choice.

A mask can be lost quietly. np.concatenate, np.stack, np.append and np.where all return arrays that have silently dropped it. The np.ma equivalents do not.

List columns cannot carry a mask, so a table read into NumPy is not uniform: most columns are masked arrays, list columns are Python lists. Arrow does not have this split.

A read hands back data, not a handle. An h5py dataset is a lazy thing: you can hold one, slice it later, and nothing is read until you do. What read() and to_arrow() return is the opposite — every row, already in memory, before the call comes back. If a column is larger than you want to hold, ask for part of it up front, as reading part of a column describes, rather than reading it whole and slicing afterwards.

Arrow needs a dependency. It is optional on purpose: a file written by h5col can be read with nothing but HDF5, and requiring a large package for the base case would undercut that.