Inspecting a table#
info() answers two questions about a table: what is in it,
and how all that content is stored.
>>> table.info()
h5col.Table '/measurements' nrows=20,000 version=1.0
"Ocean buoy hourly observations"
Columns (6)
name kind dtype units chunk filters notes
-------- ----------- --------- ----- --------- ----------------- ---------------------
time numeric int64 s 4,096 shuffle(8) | zstd
station categorical int8 4,096 3 categories, ordered
temp_c numeric float32 degC 4,096 shuffle(4) | zstd
qc boolean bool 4,096
notes string S16 utf-8 4,096
readings list float32 1,048,576 nullable
Search indexes (1)
column name kind state
------ ------------------ ------------ -------
time time__chunk_minmax CHUNK_MINMAX current
A display column that no table column uses is left out, so a table with no units does not print a column of blanks.
The last column, notes, holds whichever kind-specific extra applies: a
category count, an ordering, whether a list column’s rows may be null.
In a notebook#
The same report renders with disclosure triangles, one per column, each opening onto that column’s detail. The triangles are checkboxes and labels styled with CSS and not JavaScript, which is what makes them suitable for nbconvert, nbviewer, VS Code and Quarto, all of which strip scripts out of a rich display but keep an inline stylesheet.
What it costs to ask#
By default info() reads object headers and nothing else. It does not read a
column’s values, and it does not measure stored size, so asking a very large
table what it contains is cheap even over a network.
storage=True adds the measurements, which means walking every column’s chunk
index — quick on a local file, likely much slower for a remote file in object
storage unless cloud optimized:
>>> table.info(storage=True)
...
Storage
what logical stored ratio
-------------- -------- ------- -----
columns 260.0 kB 34.6 kB 7.5x
categories 3 B 3 B 1.0x
search indexes 200 B 65.5 kB 0.00x
The totals stay on separate lines rather than folding into one number, so an index-heavy table cannot make its columns look better compressed than they are.
A ratio below 1 is not a mistake. It means the column occupies more room than its values need, which is usually a chunk sized far larger than the data that went into it.
Everything about one column#
full=True replaces the summary with a labelled block per column, which is the
same detail a notebook shows behind a triangle:
>>> table.info(full=True)
...
temp_c
kind numeric
dtype float32
fill value 9.96921e+36
units degC (vocabulary: UDUNITS-2)
valid range -5.0 .. 40.0
chunk 4,096 rows
filters shuffle(4) | zstd
storage 80.0 kB -> 12.4 kB (6.4x)
description Sea surface temperature
On a file that is wrong#
info() is what to reach for when a file looks broken, before working with that
table raises an error. What it cannot read is reported as absent: a table
missing its VERSION reports version as nothing and lists its columns anyway,
and a categorical column whose CATEGORIES reference does not resolve is still
listed, without its category count.
The records behind the report#
info() returns a TableInfo holding
ColumnInfo and IndexInfo records. The printed
forms are renderings of those, so anything the report shows can be reached as a
value and inspected:
report = table.info(storage=True)
worst = min(
(c for c in report.columns if c.storage and c.storage.ratio),
key=lambda c: c.storage.ratio,
)
print(worst.name, worst.storage)