Inspecting a table#

info() answers two questions about a table: what is in it, and how all that content is stored.

>>> table.info()
h5col.Table '/measurements'  nrows=20,000  version=1.0
"Ocean buoy hourly observations"

Columns (6)
  name      kind         dtype      units  chunk      filters            notes
  --------  -----------  ---------  -----  ---------  -----------------  ---------------------
  time      numeric      int64      s      4,096      shuffle(8) | zstd
  station   categorical  int8              4,096                         3 categories, ordered
  temp_c    numeric      float32    degC   4,096      shuffle(4) | zstd
  qc        boolean      bool              4,096
  notes     string       S16 utf-8         4,096
  readings  list         float32           1,048,576                     nullable

Search indexes (1)
  column  name                kind          state
  ------  ------------------  ------------  -------
  time    time__chunk_minmax  CHUNK_MINMAX  current

A display column that no table column uses is left out, so a table with no units does not print a column of blanks.

The last column, notes, holds whichever kind-specific extra applies: a category count, an ordering, whether a list column’s rows may be null.

In a notebook#

The same report renders with disclosure triangles, one per column, each opening onto that column’s detail. The triangles are checkboxes and labels styled with CSS and not JavaScript, which is what makes them suitable for nbconvert, nbviewer, VS Code and Quarto, all of which strip scripts out of a rich display but keep an inline stylesheet.

What it costs to ask#

By default info() reads object headers and nothing else. It does not read a column’s values, and it does not measure stored size, so asking a very large table what it contains is cheap even over a network.

storage=True adds the measurements, which means walking every column’s chunk index — quick on a local file, likely much slower for a remote file in object storage unless cloud optimized:

>>> table.info(storage=True)
...
Storage
  what            logical   stored   ratio
  --------------  --------  -------  -----
  columns         260.0 kB  34.6 kB  7.5x
  categories      3 B       3 B      1.0x
  search indexes  200 B     65.5 kB  0.00x

The totals stay on separate lines rather than folding into one number, so an index-heavy table cannot make its columns look better compressed than they are.

A ratio below 1 is not a mistake. It means the column occupies more room than its values need, which is usually a chunk sized far larger than the data that went into it.

Everything about one column#

full=True replaces the summary with a labelled block per column, which is the same detail a notebook shows behind a triangle:

>>> table.info(full=True)
...
  temp_c
      kind         numeric
      dtype        float32
      fill value   9.96921e+36
      units        degC   (vocabulary: UDUNITS-2)
      valid range  -5.0 .. 40.0
      chunk        4,096 rows
      filters      shuffle(4) | zstd
      storage      80.0 kB -> 12.4 kB (6.4x)
      description  Sea surface temperature

On a file that is wrong#

info() is what to reach for when a file looks broken, before working with that table raises an error. What it cannot read is reported as absent: a table missing its VERSION reports version as nothing and lists its columns anyway, and a categorical column whose CATEGORIES reference does not resolve is still listed, without its category count.

The records behind the report#

info() returns a TableInfo holding ColumnInfo and IndexInfo records. The printed forms are renderings of those, so anything the report shows can be reached as a value and inspected:

report = table.info(storage=True)
worst = min(
    (c for c in report.columns if c.storage and c.storage.ratio),
    key=lambda c: c.storage.ratio,
)
print(worst.name, worst.storage)