Changelog#
All notable changes to h5col are recorded here. The format follows Keep a Changelog, and the project aims to follow Semantic Versioning.
[0.5.0] - 2026-08-24#
Added#
Table.info()reports table content with optional full information (storage, compression, metadata). The result prints as an aligned block at a prompt and renders interactive guide with disclosure triangles in a Jupyter notebook. Both renderings draw on the same records, which stay reachable asTableInfo,ColumnInfoandIndexInfofor anyone who wants to inspect the information rather than just its pretty rendering.By default it reads object headers only: no column values, no stored sizes.
storage=Trueadds the measurements, andfull=True(includesstorage=True) replaces the summary with a labelled block per column.It does not raise on a malformed file, since that is when content information may be valuable. What column property is missing is reported as absent.
Tablekeeps its shortrepr, so a baretableshows the same one line in a notebook as at a prompt.A column reports the filter plugin pipeline it was created with.
Column.filtersgives it as aFilterPipeline, andListColumn.filtersgives one per dataset making up the column, keyed by path within its group.FilterPipeline.from_dataset(ds)does the same for any h5py dataset, andFilterPipeline.from_dcpl(dcpl)for a dataset-creation property list.What comes back is what HDF5 stored, which is not always what was declared: a plugin may add client data of its own while the dataset is created. A
Shuffle()declared with no parameters reads back asshuffle(4)on a 4-byte column, because that is the element size the plugin recorded.A column reports its bytes stored in the file.
Column.storagegives aStorageholding the bytes the values would occupy unfiltered, the bytes HDF5 wrote, and their ratio, e.g.:<h5col.Storage 80.0 kB -> 30.6 kB (2.6x)>.ListColumn.storageadds up the datasets making up the column, anddataset_storage/group_storagedo the same for any h5py dataset or group.The ratio is
None, printed as—, when there is nothing to divide. An empty column and one left entirely at its fill value both leave HDF5 with no chunks to allocate. Careful when invoking this report for a wide table in object store unless the file is cloud optimized.Filter plugins can be named.
from_name("zstd", 5)builds a filter from the canonical name kept by The HDF Group’s filter plugin registry,plugin_name(id)goes the other way, and a pipeline accepts bare names —FilterPipeline(["shuffle", "zstd"])— for plugin defaults. Names are lower case and matched exactly, as the registry specifies; an identifier the registry does not name answersplugin_<id>.A filter reports whether it can run here.
Filter.availablesays whether the plugin implementing it is registered with the HDF5 library, andFilter.can_encode/Filter.can_decodesay whether that plugin was built able to write and to read. A column that declares a missing plugin as a mandatory filter is now refused when it is created, with a message naming the plugin, rather than later when storing data. Marking the filter optional still works as HDF5 skips a filter it cannot apply and writes the column without it.
Changed#
Every column is created the same way. Fixed-length string columns used to be built through h5py’s high-level dataset API, which holds one compressor, puts the filters in an order of its own choosing, and drops per-filter optional flags. They now go through the dataset-creation property list like every other column, so a string column keeps the pipeline it was given.
A filter plugin prints its name and configuration data.
repr(Deflate(4))is nowdeflate(4)and a pipeline is<h5col.FilterPipeline shuffle(4) | zstd(5)>, rather than the generated dataclass forms naming every field. Configuration data (cd_values[]) are printed whole and undecoded, the values a plugin added to the ones it was given are not separated.FilterPipeline.to_h5py_kwargs()is deprecated and will be removed in a later release. Nothing in h5col creates datasets through h5py’s filter keywords any more. Code calling h5py’screate_datasetdirectly can keep using it for now.FilterPipeline.apply()is the replacement.
Fixed#
A string column’s fill value is stored correctly. h5py’s low-level
set_fill_valuebuilds its HDF5 type from the NumPy dtype and hands the buffer over unconverted, so a fixed-length string fill went into the file as a pointer rather than as characters — silently, with no error. The value is now converted the way h5py’s own high-level path converts it. (h5py pull request 2964 fixes the underlying bug.)An oversized fill value is refused instead of truncated. A fill wider than its column used to be quietly cut down to fit, so
b"WAY_TOO_LONG_VALUE"becameb"WAY_TOO_". It now raisesOversizedStringError, as an oversized row value already did.
[0.4.0] - 2026-08-15#
Added#
Arrow tables can be imported.
Table.from_arrow(group, tbl)writes apyarrow.Tableas a H5Col table, the inverse of the existingto_arrowexport. Rows go in batch by batch, so importing a large table costs about one batch of memory rather than the whole of it.Arrow’s type system is wider than H5Col’s, and the difference is not approximated: a timestamp, date, time, duration, decimal, struct, map, union, opaque binary or fixed-size-list column is refused by name, with the alternative spelled out in the message.
The harder mismatch is that Arrow marks a missing value with a null while H5Col marks one with a value from the column’s own domain. Every column that can hold a null therefore gets a fill value, and a fill that already occurs in the data is refused — that combination would otherwise produce a conformant file whose rows read as missing when they are not. The check runs at every level of a list column, since a null element is stored the same way a null row is. A boolean column holding nulls is refused outright: the convention gives booleans no fill value to store them in.
specs_from_arrow(tbl)returns the column specs the import would use, so they can be inspected and adjusted before anything is written. Chunking and filters have no Arrow equivalent and cannot be inferred at all, so this is the only way to set them. String widths and category label sets are read from the data, which is worth a look on a table you did not write.Field metadata under
h5col.becomes the column’s own annotations. Any other metadata is carried across as an ordinary HDF5 attribute, unless its name is one H5Col reserves or writes itself, which raisesReservedNameError.Needs the optional
pyarrowdependency (pip install h5col[arrow]). A new guide chapter, Writing an Arrow table into H5Col, covers all of this in prose.Opaque columns work, and cross to Arrow and back. An opaque column holds a fixed number of raw bytes per row.
h5col.opaque_fill_bytessupplies the fill value recommendation: the ASCII markerFILLfollowed by rising byte values, so an eight-byte column’s fill is46 49 4c 4c 01 02 03 04and reads asFILL....in a hex dump. No byte string can be reserved by being out of range, since any of them might be real data, so the aim is a value that is vanishingly unlikely rather than impossible. The rising tail is what makes it so. A counting byte sequence is something opaque data almost never has. The collision check applies as it does to every other datatype, so a column that does contain the pattern is still refused. Consumers read a fill value from the dataset creation property list and never recompute it, so this is a writer-side default: a file written this way is readable by anything conformant. Arrow’sfixed_size_binary[n]maps to ann-byte opaque column exactly. Variable-length Arrowbinaryremains refused.Column.units_vocabularyreads the column’sunits_vocabularyattribute, whichColumnSpeccould already write but nothing could read back.ListColumnhas had the property since 0.1.0.
Changed#
The API reference records when each entry arrived. Anything added after the first release carries a New in version note, and an entry without one has been there since 0.1.0. These pages are built from
mainand published as a single version, so that note is what tells you whether what you are reading exists in the release you installed.The documentation site shows the latest release in its navigation bar, read from this file, so a page built from
mainsays which release it is ahead of.The documentation build now refuses a
versionaddedorversionchangedthat names a version which is neither a release nor the one in development, or that sits below a numpydoc section — where Sphinx reads it as one more parameter and renders it alongside the real ones instead of as a note.
Fixed#
A fill value that occurs in its own column was a dead end.
from_arrowrefused such a column and told the caller to supply aColumnSpecwith a different fill. Butspecs_from_arrowrefused it too, so the advice could not be followed. Auint8column containing 255, or alist<uint8>containing one, was simply unimportable.Deciding a fill value has moved to where the table is written.
specs_from_arrownow answers what the columns would look like and stops there, returning specs whosefill_valueis unset, meaning the recommended value for the datatype, as it does everywhere else. The check itself is unchanged and still unskippable. It runs on write whichever way the specs arrived. The refusal also names a value that would work, found by walking inward from the limits of the datatype, where H5Col puts its recommendations for the same reason. Signed types skip their own minimum. The suggestion is offered rather than applied. A value absent from the data at hand is not the same as one outside the column’s logical range, and picking silently would mean a laterappendcould turn real rows into missing ones. Where nothing near the limits is free, the message says so and offers the remedy: widen the datatype.A
Noneelement written into a fixed-length string list leaf that declares no fill value was stored as empty bytes rather than refused. With no fill value to compare against, that element reads back as an empty string which is present — the row was written to say “missing” and says the opposite. Every other non-boolean leaf already raised in that situation; the check now covers all of them, so the two cannot drift apart again. Leavesh5colcreates are unaffected, having always been given a fill value; this is reachable through a file whose leaf came from elsewhere.The Arrow export left the
orderedflag of a categorical column’s Arrow type at 0 even for an ordered categorical, recording the fact only in theh5col.orderedmetadata key. A consumer reading the type rather than the metadata — which is where Arrow puts this — saw an unordered dictionary. The exported type now carries the flag, andh5col.orderedstays as the key that survives a Parquet round trip.The Arrow export dropped a scalar column’s
units_vocabulary. Onlyunits,descriptionand the valid range travelled, so the vocabulary the units were drawn from was lost on the way out. List columns were unaffected.to_arrowcrashed on a column whose datatype Arrow has no type for, withArrowNotImplementedError: Unsupported numpy type 20raised from inside pyarrow — on a tableh5colitself wrote andvalidate(deep=True)passes. The convention lets a column dataset carry any HDF5 datatype, so opaque, compound, array and complex columns are all legitimate; they are now refused by name, saying which column it is and that the values are still readable throughread()or the column’sdataset. The same guard covers a list column’s leaf values.
[0.3.0] - 2026-08-11#
Added#
Columns read by subscript.
column[17:98],column[-1],column[[3, 1, 3]]andcolumn[mask]areread_rowswith its defaults, so a row selection reads the way it does in h5py and NumPy. An integer key returns that row’s value on its own —numpy.ma.maskedwhen the row is missing — and every other key returns an array. Subscript has nowhere to put a keyword, so it always decodes and always masks;readandread_rowsremain the forms that takemasked=False.len(column)is the row count and iterating a column yields its decoded rows, reading the column once rather than once per row. Note that having a length makes a column with no rows falsy, as it does for a list or an h5py dataset, soif column:now asks whether the column has rows rather than whether the object exists.List columns accept the same keys, and gained
read_rowsto match.Note that
column[...]andcolumn.dataset[...]are not the same read. The second goes straight to h5py, so it skips decoding, ignores missing values, and can return rows aboveNROWSthattruncateleft behind as reserved storage.List columns read only the rows asked for. A list column’s rows are reached through its
OFFSETS, so reading a range means looking up where that range’s values start and end and narrowing the child read to that span — recursively, so a nested list or aSTRING_VALUESgroup narrows too. Scattered positions are served from the range that spans them, which is never wider than the column itself: rows near each other cost almost nothing, and rows at opposite ends cost what reading the column costs.Reading 50 rows of a 200,000-row list column pulls 251 values rather than a million, taking 1.5 ms rather than 79 ms. A single row costs 6 values.
Selection.readno longer sends list columns down the read-whole path, so a query matching a few rows of a table with a list column went from 84 ms to 1.6 ms, of which the list column is now a rounding error.to_arrowis unchanged and still reads the whole column: it saves no reading but skips building a Python list per row, which suits wanting most of a large column, where a range read suits wanting a small part of one.Row selections take slices, boolean masks and negative positions.
Column.read_rows,Column.to_arrowand everything built on them accept a slice (read_rows(slice(17, 98))), a boolean array with one entry per row (read_rows(col.is_missing())), and positions counting back from the end (read_rows([-1, -2])) alongside the integer sequences they already took.A slice is read as a single hyperslab rather than going through the gather path, which had to sort the positions and scatter the result back into the caller’s order — work that is pure overhead when the positions were a range to begin with. Reading a million contiguous rows from a two-million-row column drops from 6.9 ms to 1.3 ms, which is what the same read costs in h5py directly. Asking for half a column is now cheaper than asking for all of it; previously it was three times more expensive.
Example notebook Reading part of a table: subscript and
read_rowsover a 200,000-row table, with every HDF5 read counted so the saving is shown rather than asserted. Scalar columns fetch the chunks their rows land in — two rows at opposite ends cost two chunks, not the span; list columns are served from the range that spans the wanted rows, where reading fifty of them costs 251 values against a million. Also covers masked results from subscript,masked=Falseviaread_rows, and the difference betweencolumn[...]andcolumn.dataset[...].Every parameter of every public callable is documented. An audit of the source found 114 public functions and methods whose docstring left at least one parameter of its signature undescribed — 207 parameters in all, of which 86 were on callables the API reference renders.
Table.createdocumented none of its ten;Table.readandColumn.read_rowsdocumentedmaskedbut not the parameters that decide what gets read. All of them now have aParameterssection covering the whole signature.Ruff’s D417, which catches a
Parameterssection that covers only part of a signature, is switched off by the numpy docstring convention. It is now named explicitly in the lint configuration so it runs, and the gate fails if a parameter description goes missing again.
Changed#
The version is declared in one place.
pyproject.tomlnow takes it fromsrc/h5col/__init__.pyrather than repeating it, so the two cannot drift. Arelease-checkworkflow runs onv*tags and fails when the commit a tag points at does not report the version the tag claims, or when a release tag carries a pre-release version.The missing-value documentation no longer uses “sentinel”. H5Col’s marker for a missing row is simply the column’s fill value. The reference pages keep the convention’s own term, “the canonical missing-value test”.
The Reading into Python chapter gained a “Reading part of a column” section, so row selection is no longer buried in a list of limitations, and its Arrow half is now titled to match its NumPy half rather than reading as the chapter’s conclusion.
Fixed#
A boolean mask passed to
read_rowsselected the wrong rows. The mask was cast to an integer dtype, turning it into a run of ones and zeros, so a mask marking rows 3, 7 and 9 read rows 1 and 0 over and over. A mask now selects the rows it marks, and one of the wrong length raisesIndexErrorrather than being reinterpreted.Non-integer row positions were truncated silently —
read_rows([1.5])read row 1. They now raiseTypeError.
[0.2.0] - 2026-08-10#
Added#
Arrow export —
Table.to_arrow(),Column.to_arrow()andSelection.to_arrow(), behind the optionalpyarrowdependency (pip install h5col[arrow]). This is the one representation that carries the whole H5Col data model, because NumPy has no type for three of the things H5Col stores:missing rows become real Arrow nulls, so the fill value never reaches a consumer as data;
a categorical becomes a
DictionaryArrayof exactly the codes and labels already on disk, rather than being expanded to one label per row;a list column keeps its nulls at every level of nesting, including an inner null no top-level mask can express.
Numeric columns hand their data buffer to Arrow unchanged. Fixed-length string columns are converted to
large_string— a fixed-width column has no offsets to lend — with the offsets computed by array arithmetic rather than a Python loop. Each column’sunits,description,valid_min,valid_maxand (for categoricals)orderedattributes ride along as Arrow field metadata under anh5col.prefix, and survive a Parquet round trip.List columns are the case Arrow fits best, because H5Col already stores them in Arrow’s layout.
OFFSETS(uint64) is reinterpreted as Arrow’s int64 offsets without touching the memory, and aSTRING_VALUESgroup’sCHARSgoes across as the string payload. Only the null masks are converted, from H5Col’s byte per row to Arrow’s bit. Against the Pythonread()at 200,000 rows this is 19x forlist<float64>, 33x forlist<string>and 26x forlist<list<float64>>.User guide chapter Reading into Python, covering what each column type hands back from
read()andto_arrow(), why missing values arrive masked,.tolist()versuslist(), which NumPy functions drop a mask, and where each form falls short.Example notebook Exporting to Arrow:
to_arrow()end to end — real nulls in place of fill values, categoricals as Arrow dictionaries, list columns with their nesting and null-versus-empty distinction intact, column attributes carried as field metadata, and the hop to pandas and Parquet.
Changed#
read()now returns masked arrays by default.Column.read,Column.read_rows,Table.readandSelection.readtakemasked=True, and every scalar column comes back as anumpy.ma.MaskedArraywhose mask marks its missing rows. Previously a missing row was handed back as the column’s fill value with nothing to distinguish it from data, so a mean over a column with missing rows was silently wrong. Passmasked=Falsefor the previous behaviour, unchanged.Uniform across scalar columns: boolean columns, which H5Col forbids from declaring a fill, and columns that declare none still come back masked with an all-False mask, so code written over the returned dict never has to branch on whether a given column can be missing.
List columns are unchanged — ragged, so they cannot be masked arrays — and already spell a null row
None. They acceptmasked=and ignore it.fill_valueis the column’s own decoded fill value, soread().filled()reproducesread(masked=False)and any operation that drops the mask degrades to the previous behaviour rather than to NumPy’s defaults (999999 for an int8 column, the stringN/Afor a string one).Note that
list(...)over a masked array yieldsnumpy.ma.maskedwhere.tolist()yieldsNone, and thatnp.concatenate,np.stackand friends drop the mask silently — use thenp.ma.*equivalents.NumPy prints a masked array by spelling out each value rather than formatting for the dtype, so a
float32column that displayed as[21.4 27.9]now displays as[21.399999618530273 ...]. The values are unchanged.
String and categorical columns now decode into NumPy 2’s
numpy.dtypes.StringDTypeinstead of adtype=objectarray of Python strings. Categorical columns use its nullable form, since a missing row has no label to carry; categories whose labels are not strings keep an object array. For 400,000 short strings the decoded column costs 6.4 MB rather than 25.8 MB of resident memory, andColumn.readon that column drops from 23 ms to 2 ms.Column.categoriesreturns aStringDTypearray for string labels.The
StoStringDTypecast validates lazily, so a non-conformant producer’s invalid UTF-8 now raisesUnicodeDecodeErrorwhen the offending value is read out of the array rather than when the column is read.Minimum NumPy raised to 2.0, which introduced
StringDType; the conda dependency on HDF5 raised to 2.1.The documentation workflow builds only when something the site is actually built from changes, rather than on every push.
GitHub Actions updated to versions running on Node 24, clearing the Node 20 deprecation warnings.
Fixed#
Categorical decoding invented missing values.
decode_codesread the fill value without confirming theH5D_FILL_VALUE_USER_DEFINEDstate, and mapped any code outside[0, ncategories)toNone. A single column could then contradict itself —read()reporting a row asNonewhileis_missing()and the query layer reported it present — on a file that passesvalidate(deep=True). An unindexable code now raisesConformanceError, and a column declaring no fill value no longer has h5py’s library default read as one.A
numpy.ma.MaskedArraypassed toappend()had its mask ignored, so a masked row was written as whatever value sat beneath the mask. For a fixed-length string column that produced the literal characters--; for a numeric column it silently stored stale data as though it were real. A masked element now means the same asNoneon write.
[0.1.0] - 2026-08-05#
Added#
Tables backed by HDF5 groups — create, open, append, read, and
truncate, tracking the logical row count (NROWS).Column datatypes — numeric, fixed-length strings (oversized values raise
OversizedStringError; never silently truncated), the H5Col boolean enum, and categorical columns (integer codes with a label mapping).List columns —
LIST_COLUMN/STRING_VALUESwith null masks and arbitrary nesting.Filter pipelines —
Filter/FilterPipelineand the built-inDeflate/Shuffle/Fletcher32, plus adaptation ofhdf5pluginfilters, applied in declared order through the dataset-creation property list.Search indexes with validity tokens (
GENERATION/SOURCE_*) —CHUNK_MINMAX,SORTED_ROWS, andBITMAP: build, refresh, and query primitives.Query layer —
field()expressions and pyarrow-style DNF tuple filters,Table.select/read(where=)/count, with three-valued (Kleene) missing-value semantics matching Arrow, and anexplainplan.Conformance validation —
validate()and the semanticvalidate(deep=True).Friendly
FixedStringhandler for HDF5 fixed-length strings, and a full H5Col exception family rooted atH5ColError.Column.read_rows(rows)reads just the given rows, decoded, using coalesced chunk-aligned block reads. Rows may be in any order and may repeat.Documentation site under
docs/(Sphinx, Markdown via MyST, pydata theme), published to GitHub Pages: getting-started pages, a user guide, the query syntax reference, the rendered example notebooks, and an API reference. The theme is restyled with self-hosted IBM Plex Sans/Mono (SIL OFL 1.1) and a deep-teal palette tuned for both light and dark schemes.Project logo, applied across the README, the documentation navbar and landing page, the favicon and Apple touch icon, and the
og:imageused for link previews.GitHub workflows:
ci.ymlruns the gate (tests, lint, format check, type check) on Linux and macOS;docs.ymlbuilds the documentation on every push and pull request and deploys it to GitHub Pages frommain.CI runs the gate on Windows as well as Linux and macOS, and the pixi lockfile pins
win-64alongsidelinux-64andosx-arm64.CI and documentation build badges in the README.
Known limitations#
CHUNK_BLOOMsearch indexes are deferred.Object references are written as
H5T_STD_REF_OBJrather than the convention’sH5T_STD_REF, because h5py cannot yet create the unified type.