spacr.qt.widgets.pivot_spec

Tabulate — the pivot table’s engine: rows, columns, aggregations, and n.

JMP’s Tabulate is a drag-and-drop pivot: put plateID down the rows, gene across the columns, tick mean and sd, and read the table. It is the thing people reach for before they plot anything, because a number they can copy into a slide is often the whole answer.

Like spacr.qt.widgets.graph_spec, this half is pure pandas and numpy with no Qt — testable without a display, usable from a notebook, and available to whatever screen wants a summary table without inheriting a drag-and-drop panel.

n is not optional

Every cell of a PivotResult carries its n, whether or not the user ticked it, and the panel prints it. A mean over four objects and a mean over four thousand are the same three digits on screen, and the difference between them is the difference between a result and a coincidence. Making n opt-in would make the most important number on the table the one nobody turns on.

For the same reason sd and sem of a single object are NaN, not zero. One measurement has no spread; a zero there reads as “perfectly reproducible”, which is the opposite of what it means.

Empty is not zero

A cell for a combination that has no rows at all is empty in every layer, including n. 0 is a count of something, and a grid where “no objects were measured in D7” looks identical to “some were measured and none survived the filter” is a grid nobody can read. PivotResult.present keeps the two apart, and PivotResult.is_empty() is what the renderer asks.

The distinction that is real is kept: a cell with rows whose value column is entirely NaN reads n = 0 — objects were measured there and none of them produced this measurement. That is a fact worth showing, and it is different from the well not existing.

The full grid, for the same reason the facet grid is full

Rows and columns are the cartesian product of their keys’ levels, empty combinations included, exactly as spacr.qt.widgets.graph_spec.facet_grid() draws empty panels: a table read by position needs its positions to line up, and one that closes up its gaps tells the reader “there is no row H” when the truth is “row H was measured and is empty”. Above MAX_ROWS × MAX_COLS the product is abandoned for the observed combinations only, and PivotResult.notice says so — an unreadable table is not more honest than a smaller one that admits what it left out.

The plate hierarchy

plateID / rowID / columnID / fieldID is what spaCR users actually pivot on, so multiple keys per axis nest in the order they were dropped and WELL_HIERARCHY is offered as a preset. Nothing here is special-cased for them; they are ordinary columns that happen to be the common answer.

Feeding the chart

PivotResult.to_long() returns one row per cell with one column per statistic, which is exactly the shape GraphBuilderPanel wants: x = plateID, y = mean, size = n. The pivot does not draw anything.

Exceptions

PivotError

A pivot that cannot mean anything, with the reason in the message.

Classes

PivotResult

A computed table: one 2-D array per (value, agg) layer, plus n.

PivotSpec

What goes down the rows, across the columns, and into the cells.

Functions

format_value(→ str)

One number for a table cell, or '' for a missing one.

pivot(→ PivotResult)

Compute the table spec describes over frame.

Module Contents

exception spacr.qt.widgets.pivot_spec.PivotError[source]

Bases: ValueError

A pivot that cannot mean anything, with the reason in the message.

Initialize self. See help(type(self)) for accurate signature.

class spacr.qt.widgets.pivot_spec.PivotResult[source]

A computed table: one 2-D array per (value, agg) layer, plus n.

Parameters:
  • row_keys – the columns nesting down the rows, outermost first, as in the spec’s rows.

  • col_keys – the columns nesting across, outermost first, as in the spec’s cols.

  • row_levels – one tuple of level strings per displayed row, aligned with row_keys.

  • col_levels – likewise across.

  • layers – {(value, agg): array}, each (nrow, ncol) of float, NaN where the cell is empty or the statistic does not exist (sd at n=1).

  • sizes – rows of the source frame in each cell — the row count, not the value count. sizes == 0 and present == False are the same thing here, and both are what makes a cell blank.

  • present – whether the combination has any rows at all.

  • spec – the PivotSpec the table was computed from; its layers name the arrays in layers.

  • n_source_rows – rows in the frame the table was computed over.

  • hidden_rows – source rows whose keys fell outside the displayed levels. Non-zero means the table is not the whole frame.

  • notice – the computation’s notes joined with "; " – a grid that was cut down, rows left outside the shown levels, a value column that is not numeric – or empty when there were none.

col_label(col: int, sep: str = ' · ') → str[source]

One column’s grouping levels, joined for display.

Parameters:
  • col – the column’s position.

  • sep – what to join the levels with.

Returns:

the label.

is_empty(row: int, col: int) → bool[source]

No rows of the source frame landed here. Renders blank, not 0.

Parameters:
  • row – displayed row index, 0-based.

  • col – displayed column index, 0-based.

low_n_cells(threshold: int = LOW_N) → int[source]

Non-empty cells whose n is at or below threshold.

n_at(value: str, row: int, col: int) → int | None[source]

The n behind a cell’s statistics, or None when it is empty.

None and 0 are different: None is “nothing was measured in this combination”, 0 is “objects were measured and none of them has a value for this feature”.

Parameters:
  • value – the value column the layer aggregates, or COUNT_ONLY ("") for a table with no value columns.

  • row – displayed row index, 0-based.

  • col – displayed column index, 0-based.

n_range() → Tuple[int, int] | None[source]

(smallest, largest) n over the non-empty cells, or None.

The n behind the statistics — values, not rows — because that is the number a mean was taken over. They differ whenever a value column has NaN in it, and the smaller one is the one that matters.

row_label(row: int, sep: str = ' · ') → str[source]

One row’s grouping levels, joined for display.

Parameters:
  • row – the row’s position.

  • sep – what to join the levels with.

Returns:

the label.

summary() → str[source]

One line under the table: shape, n range, and what was left out.

to_csv(path: str) → str[source]

Write to_frame() to path. Returns the path.

Parameters:

path – the CSV file to write, without an index column; an existing file is overwritten.

to_frame() → pandas.DataFrame[source]

The spreadsheet shape: row keys as leading columns, one column per (column level × value × agg).

What the CSV export writes and what a user pastes into a slide. The multi-level column header is flattened into one readable string because a CSV has one header row, and a reader who has to reassemble three of them has been handed a puzzle rather than a table.

to_long() → pandas.DataFrame[source]

One row per non-empty cell; one column per statistic.

The frame to hand the Graph Builder: x = plateID, y = mean, size = n, colour = gene. Empty cells are omitted rather than carried as NaN rows — a scatter of nothing is a mark at the origin waiting to happen, and present is where “which combinations were empty” lives.

With more than one value column, value_column names which one the row is about, so the frame stays tidy instead of growing a column per (value × agg) pair.

value_at(value: str, agg: str, row: int, col: int) → float[source]

One statistic, or NaN when it is empty or does not exist.

Parameters:
  • value – the value column the layer aggregates, or COUNT_ONLY ("") for a table with no value columns.

  • agg – the aggregation, one of AGGREGATIONS; a (value, agg) layer the table does not hold raises PivotError.

  • row – displayed row index, 0-based.

  • col – displayed column index, 0-based.

property layer_keys: Tuple[Tuple[str, str], ...][source]

The (value, agg) pairs this result holds one array for.

Returns:

the layer keys, in the spec’s order.

property n_cells: int[source]

How many cells the table has, across all layers.

Returns:

rows times columns.

property shape: Tuple[int, int][source]

The table’s size as (rows, columns) of DISPLAYED levels.

Not the source frame’s shape: a pivot’s rows are groups, so this is how big the answer is rather than how much went into it.

Returns:

the row and column counts.

class spacr.qt.widgets.pivot_spec.PivotSpec[source]

What goes down the rows, across the columns, and into the cells.

Frozen and JSON round-tripping like GraphSpec, and for the same reason: a table is something a settings file, a report or a macro should be able to carry, and every edit returning a new spec makes “undo the last drag” trivial.

Parameters:
  • rows – keys nesting down the rows, outermost first.

  • cols – keys nesting across the columns, outermost first.

  • values – the columns to aggregate. Empty means a contingency table: n counts rows and nothing else is computed.

  • aggs – which aggregations to compute. N is always added — see the module docstring.

  • quantile – the fraction QUANTILE reports, in [0, 1].

Raises:

PivotError – on an unknown aggregation, a quantile outside [0, 1], or a column used on two axes at once — at the point the spec is built rather than when the table comes out wrong.

__post_init__() → None[source]

Normalise the axes and validate the aggregations.

n is always present and always first: a table where the user could turn it off is a table where a mean over four objects looks like a mean over four thousand.

Raises:

PivotError – if a column is on both the row and the column axis – one column cannot nest inside itself, and every cell off the diagonal would be empty by construction; if an aggregation is unknown; or if quantile is not a fraction in [0, 1].

describe() → str[source]

The pivot in one line: what goes down, across, and into the cells.

Returns:

a one-line description.

classmethod from_dict(payload: Mapping[str, Any]) → PivotSpec[source]

Rebuild a spec from plain data.

UNKNOWN KEYS ARE IGNORED rather than raising, so a spec saved by a later version still opens with the parts this one knows.

Parameters:

payload – what to_dict() produced.

Returns:

the rebuilt spec.

classmethod from_json(text: str) → PivotSpec[source]

Rebuild a spec from JSON text.

Parameters:

text – the JSON text.

Returns:

the rebuilt spec.

to_dict() → Dict[str, Any][source]

This spec as plain data.

Returns:

a JSON-safe dict.

to_json() → str[source]

This spec as JSON text, keys sorted so the file is diffable.

Returns:

the JSON text.

used_columns() → Tuple[str, ...][source]

Every column this pivot reads, deduplicated and in order.

What lets a spec be validated against a table before it is computed, so a pivot saved on one experiment says which columns are missing here rather than failing part-way through the aggregation.

Returns:

the column names.

with_aggs(aggs: Sequence[str]) → PivotSpec[source]

A copy using different aggregations.

Parameters:

aggs – the aggregation names, such as mean or median.

Returns:

the new spec.

with_cols(cols: Sequence[str]) → PivotSpec[source]

A copy with a different set of column groupings.

Parameters:

cols – the columns to group across the columns.

Returns:

the new spec.

with_rows(rows: Sequence[str]) → PivotSpec[source]

A copy with a different set of row groupings.

A COPY: a spec is a value, so the one a view is already showing is never edited underneath it.

Parameters:

rows – the columns to group down the rows.

Returns:

the new spec.

with_values(values: Sequence[str]) → PivotSpec[source]

A copy aggregating different columns into the cells.

Parameters:

values – the columns to aggregate.

Returns:

the new spec.

property is_empty: bool[source]

no table yet.

Type:

Nothing on any axis and nothing to count

property layers: Tuple[Tuple[str, str], ...][source]

(value, agg) pairs, in the order a cell stacks them.

With no value column there is one layer, (COUNT_ONLY, N) — the contingency count.

spacr.qt.widgets.pivot_spec.format_value(value: float, *, digits: int = 4) → str[source]

One number for a table cell, or '' for a missing one.

Blank rather than nan: a cell that reads nan is read as an error in the software, and a cell that reads 0 is worse. See the module docstring — an sd of a single object is genuinely nothing, and blank is what nothing looks like.

Parameters:
  • value – the cell value; None, NaN and infinities give '', and whole numbers are written without decimals.

  • digits – significant digits for a non-integer value.

spacr.qt.widgets.pivot_spec.pivot(frame: pandas.DataFrame, spec: PivotSpec | None = None) → PivotResult[source]

Compute the table spec describes over frame.

The policy is in the module docstring; the short version is that every cell carries its n, an empty combination is empty rather than zero, sd and sem are ddof=1 and therefore blank at n=1, and the grid is the full cartesian product until that stops being readable.

Parameters:
  • frame – the source table; every value column and row/column key the spec names must be one of its columns.

  • spec – what to compute; None uses a default PivotSpec.

Raises:

PivotError – for a spec that cannot describe a table over this frame, with the reason in the message.