spacr.qt.widgets.pivot_spec¶
Tabulate — the pivot table’s engine: rows, columns, aggregations, and n.
JMP’s Tabulate is a drag-and-drop pivot: put plateID down the rows,
gene across the columns, tick mean and sd, and read the table. It is
the thing people reach for before they plot anything, because a number they can
copy into a slide is often the whole answer.
Like spacr.qt.widgets.graph_spec, this half is pure pandas and numpy
with no Qt — testable without a display, usable from a notebook, and
available to whatever screen wants a summary table without inheriting a
drag-and-drop panel.
n is not optional¶
Every cell of a PivotResult carries its n, whether or not the user
ticked it, and the panel prints it. A mean over four objects and a mean over
four thousand are the same three digits on screen, and the difference between
them is the difference between a result and a coincidence. Making n opt-in
would make the most important number on the table the one nobody turns on.
For the same reason sd and sem of a single object are NaN, not
zero. One measurement has no spread; a zero there reads as “perfectly
reproducible”, which is the opposite of what it means.
Empty is not zero¶
A cell for a combination that has no rows at all is empty in every layer,
including n. 0 is a count of something, and a grid where “no objects were
measured in D7” looks identical to “some were measured and none survived the
filter” is a grid nobody can read. PivotResult.present keeps the two
apart, and PivotResult.is_empty() is what the renderer asks.
The distinction that is real is kept: a cell with rows whose value column is
entirely NaN reads n = 0 — objects were measured there and none of them
produced this measurement. That is a fact worth showing, and it is different
from the well not existing.
The full grid, for the same reason the facet grid is full¶
Rows and columns are the cartesian product of their keys’ levels, empty
combinations included, exactly as
spacr.qt.widgets.graph_spec.facet_grid() draws empty panels: a table
read by position needs its positions to line up, and one that closes up its
gaps tells the reader “there is no row H” when the truth is “row H was
measured and is empty”. Above MAX_ROWS × MAX_COLS the product
is abandoned for the observed combinations only, and PivotResult.notice
says so — an unreadable table is not more honest than a smaller one that
admits what it left out.
The plate hierarchy¶
plateID / rowID / columnID / fieldID is what spaCR users
actually pivot on, so multiple keys per axis nest in the order they were
dropped and WELL_HIERARCHY is offered as a preset. Nothing here is
special-cased for them; they are ordinary columns that happen to be the common
answer.
Feeding the chart¶
PivotResult.to_long() returns one row per cell with one column per
statistic, which is exactly the shape
GraphBuilderPanel wants: x =
plateID, y = mean, size = n. The pivot does not draw anything.
Exceptions¶
A pivot that cannot mean anything, with the reason in the message. |
Classes¶
A computed table: one 2-D array per |
|
What goes down the rows, across the columns, and into the cells. |
Functions¶
|
One number for a table cell, or |
|
Compute the table |
Module Contents¶
- exception spacr.qt.widgets.pivot_spec.PivotError[source]¶
Bases:
ValueErrorA pivot that cannot mean anything, with the reason in the message.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.qt.widgets.pivot_spec.PivotResult[source]¶
A computed table: one 2-D array per
(value, agg)layer, plus n.- Parameters:
row_keys – the columns nesting down the rows, outermost first, as in the spec’s
rows.col_keys – the columns nesting across, outermost first, as in the spec’s
cols.row_levels – one tuple of level strings per displayed row, aligned with
row_keys.col_levels – likewise across.
layers –
{(value, agg): array}, each(nrow, ncol)of float, NaN where the cell is empty or the statistic does not exist (sdat n=1).sizes – rows of the source frame in each cell — the row count, not the value count.
sizes == 0andpresent == Falseare the same thing here, and both are what makes a cell blank.present – whether the combination has any rows at all.
spec – the
PivotSpecthe table was computed from; itslayersname the arrays inlayers.n_source_rows – rows in the frame the table was computed over.
hidden_rows – source rows whose keys fell outside the displayed levels. Non-zero means the table is not the whole frame.
notice – the computation’s notes joined with
"; "– a grid that was cut down, rows left outside the shown levels, a value column that is not numeric – or empty when there were none.
- col_label(col: int, sep: str = ' · ') str[source]¶
One column’s grouping levels, joined for display.
- Parameters:
col – the column’s position.
sep – what to join the levels with.
- Returns:
the label.
- is_empty(row: int, col: int) bool[source]¶
No rows of the source frame landed here. Renders blank, not 0.
- Parameters:
row – displayed row index, 0-based.
col – displayed column index, 0-based.
- n_at(value: str, row: int, col: int) int | None[source]¶
The n behind a cell’s statistics, or
Nonewhen it is empty.Noneand0are different:Noneis “nothing was measured in this combination”,0is “objects were measured and none of them has a value for this feature”.- Parameters:
value – the value column the layer aggregates, or
COUNT_ONLY("") for a table with no value columns.row – displayed row index, 0-based.
col – displayed column index, 0-based.
- n_range() Tuple[int, int] | None[source]¶
(smallest, largest)n over the non-empty cells, orNone.The n behind the statistics — values, not rows — because that is the number a mean was taken over. They differ whenever a value column has NaN in it, and the smaller one is the one that matters.
- row_label(row: int, sep: str = ' · ') str[source]¶
One row’s grouping levels, joined for display.
- Parameters:
row – the row’s position.
sep – what to join the levels with.
- Returns:
the label.
- to_csv(path: str) str[source]¶
Write
to_frame()topath. Returns the path.- Parameters:
path – the CSV file to write, without an index column; an existing file is overwritten.
- to_frame() pandas.DataFrame[source]¶
The spreadsheet shape: row keys as leading columns, one column per
(column level × value × agg).What the CSV export writes and what a user pastes into a slide. The multi-level column header is flattened into one readable string because a CSV has one header row, and a reader who has to reassemble three of them has been handed a puzzle rather than a table.
- to_long() pandas.DataFrame[source]¶
One row per non-empty cell; one column per statistic.
The frame to hand the Graph Builder:
x = plateID,y = mean,size = n,colour = gene. Empty cells are omitted rather than carried as NaN rows — a scatter of nothing is a mark at the origin waiting to happen, andpresentis where “which combinations were empty” lives.With more than one value column,
value_columnnames which one the row is about, so the frame stays tidy instead of growing a column per (value × agg) pair.
- value_at(value: str, agg: str, row: int, col: int) float[source]¶
One statistic, or NaN when it is empty or does not exist.
- Parameters:
value – the value column the layer aggregates, or
COUNT_ONLY("") for a table with no value columns.agg – the aggregation, one of
AGGREGATIONS; a(value, agg)layer the table does not hold raisesPivotError.row – displayed row index, 0-based.
col – displayed column index, 0-based.
- property layer_keys: Tuple[Tuple[str, str], ...][source]¶
The
(value, agg)pairs this result holds one array for.- Returns:
the layer keys, in the spec’s order.
- class spacr.qt.widgets.pivot_spec.PivotSpec[source]¶
What goes down the rows, across the columns, and into the cells.
Frozen and JSON round-tripping like
GraphSpec, and for the same reason: a table is something a settings file, a report or a macro should be able to carry, and every edit returning a new spec makes “undo the last drag” trivial.- Parameters:
rows – keys nesting down the rows, outermost first.
cols – keys nesting across the columns, outermost first.
values – the columns to aggregate. Empty means a contingency table:
ncounts rows and nothing else is computed.aggs – which aggregations to compute.
Nis always added — see the module docstring.quantile – the fraction
QUANTILEreports, in[0, 1].
- Raises:
PivotError – on an unknown aggregation, a quantile outside
[0, 1], or a column used on two axes at once — at the point the spec is built rather than when the table comes out wrong.
- __post_init__() None[source]¶
Normalise the axes and validate the aggregations.
nis always present and always first: a table where the user could turn it off is a table where a mean over four objects looks like a mean over four thousand.- Raises:
PivotError – if a column is on both the row and the column axis – one column cannot nest inside itself, and every cell off the diagonal would be empty by construction; if an aggregation is unknown; or if
quantileis not a fraction in[0, 1].
- describe() str[source]¶
The pivot in one line: what goes down, across, and into the cells.
- Returns:
a one-line description.
- classmethod from_dict(payload: Mapping[str, Any]) PivotSpec[source]¶
Rebuild a spec from plain data.
UNKNOWN KEYS ARE IGNORED rather than raising, so a spec saved by a later version still opens with the parts this one knows.
- Parameters:
payload – what
to_dict()produced.- Returns:
the rebuilt spec.
- classmethod from_json(text: str) PivotSpec[source]¶
Rebuild a spec from JSON text.
- Parameters:
text – the JSON text.
- Returns:
the rebuilt spec.
- to_json() str[source]¶
This spec as JSON text, keys sorted so the file is diffable.
- Returns:
the JSON text.
- used_columns() Tuple[str, ...][source]¶
Every column this pivot reads, deduplicated and in order.
What lets a spec be validated against a table before it is computed, so a pivot saved on one experiment says which columns are missing here rather than failing part-way through the aggregation.
- Returns:
the column names.
- with_aggs(aggs: Sequence[str]) PivotSpec[source]¶
A copy using different aggregations.
- Parameters:
aggs – the aggregation names, such as
meanormedian.- Returns:
the new spec.
- with_cols(cols: Sequence[str]) PivotSpec[source]¶
A copy with a different set of column groupings.
- Parameters:
cols – the columns to group across the columns.
- Returns:
the new spec.
- with_rows(rows: Sequence[str]) PivotSpec[source]¶
A copy with a different set of row groupings.
A COPY: a spec is a value, so the one a view is already showing is never edited underneath it.
- Parameters:
rows – the columns to group down the rows.
- Returns:
the new spec.
- spacr.qt.widgets.pivot_spec.format_value(value: float, *, digits: int = 4) str[source]¶
One number for a table cell, or
''for a missing one.Blank rather than
nan: a cell that readsnanis read as an error in the software, and a cell that reads0is worse. See the module docstring — an sd of a single object is genuinely nothing, and blank is what nothing looks like.- Parameters:
value – the cell value;
None, NaN and infinities give'', and whole numbers are written without decimals.digits – significant digits for a non-integer value.
- spacr.qt.widgets.pivot_spec.pivot(frame: pandas.DataFrame, spec: PivotSpec | None = None) PivotResult[source]¶
Compute the table
specdescribes overframe.The policy is in the module docstring; the short version is that every cell carries its n, an empty combination is empty rather than zero,
sdandsemareddof=1and therefore blank at n=1, and the grid is the full cartesian product until that stops being readable.- Parameters:
frame – the source table; every value column and row/column key the spec names must be one of its columns.
spec – what to compute;
Noneuses a defaultPivotSpec.
- Raises:
PivotError – for a spec that cannot describe a table over this frame, with the reason in the message.