spacr.qt.widgets.graph_spec

The Graph Builder’s spec: what is on which channel, and nothing else.

A JMP-style graph builder is two things that want very much to be one thing: a pile of drag-and-drop chrome, and a small declarative object saying “area is on x, gene is on colour, facet by plate down the rows”. This module is the second one, kept deliberately apart from the first.

Why the separation is the important decision

Four later screens — small multiples, the gate editor, the feature explorer and the campaign control charts — are all “a Graph Builder with one extra rule”. If the axis and facet logic lives inside a QWidget, each of them either re-derives it or inherits a widget it does not want. So everything in here is

  • pure pandas and numpy, with no Qt, like spacr.selection — testable without a display, usable from a notebook or from spacr-run;

  • serialisable — GraphSpec.to_dict() round-trips through JSON, so a chart is something a settings file, a report or a macro can carry;

  • immutable — every channel change returns a new spec, which is what makes “undo the last drag” and “compare two specs” trivial rather than a diffing problem.

The four pieces

GraphSpec

Which column is on which of the six channels (x, y, colour, size, facet-row, facet-column), the plot type (or None for “infer it”), and the handful of options that change what is computed rather than what it looks like.

facet_grid()

The panel layout — the full cartesian product of the row and column levels, empty combinations included. A missing panel and an empty panel say different things (“this table has no plate 3 / row H” versus “plate 3 row H was measured and nothing survived the filter”), and only one of them is true; drawing the grid complete is what keeps the reader from guessing.

scales_for()

One set of limits, bin edges, category orders and colour levels for every panel. Shared axes are the default because comparing across panels is the entire point of faceting, and two panels whose y axes differ by an order of magnitude look identical.

prepare_data()

The large-data policy, stated rather than hidden. See below.

Large data

spaCR measurement tables run to 10^5–10^6 object rows and nobody wants a scatter of a million overlapping dots. Three strategies, chosen by what the plot actually needs, and the chosen one is always named in RenderData.notice so a subset can never be mistaken for the whole:

  • aggregate plots use every row, always. A histogram, bar, box, violin or heatmap is already a reduction — sampling before aggregating would change the answer for no gain, so AGGREGATE_KINDS never sample regardless of size.

  • scatter/line up to DEFAULT_POINT_BUDGET rows: every row is a mark.

  • scatter above the budget: 2-D density binning — every row is counted into a shared-edge 2-D histogram and the panel is drawn as a raster. Nothing is dropped, the density is quantitative, and a brush is still exact because a rectangle brush is a predicate on x and y evaluated against the full frame, not against the pixels.

  • scatter above the budget where binning cannot answer the question — a categorical colour or a size channel needs per-point marks — falls back to a seeded uniform sample, with the count, the fraction and the word “sampled” in the notice.

The one thing not on offer is quietly plotting the head of the frame.

Exceptions

SpecError

A spec that cannot mean anything — an unknown channel or plot type.

Classes

FacetGrid

The complete panel layout — including the combinations with no rows.

FacetPanel

One panel of the grid, and the rows that belong in it.

GraphSpec

Which column is on which channel, and what to draw.

RenderData

What the renderer should draw, and what it must say about it.

Scales

One set of limits and orders for every panel.

Functions

brush_mask(→ numpy.ndarray)

Rows of frame inside the rectangle a user dragged on a panel.

column_kinds(→ Dict[str, str])

Sort frame's columns into CONTINUOUS / CATEGORICAL

facet_grid(→ FacetGrid)

Split frame into the grid spec's facet channels describe.

infer_kind(→ str)

The plot type the dropped columns imply.

plottable_columns(→ Tuple[str, ...])

The columns worth offering in the drag well, sorted.

prepare_data(→ RenderData)

Decide how many rows get drawn, and say so.

scales_for(→ Scales)

Limits, bin edges, level orders and colour levels shared by every panel.

value_axes(→ Tuple[Optional[str], Optional[str]])

Which column each axis actually carries, once the kind is known.

Module Contents

exception spacr.qt.widgets.graph_spec.SpecError[source]

Bases: ValueError

A spec that cannot mean anything — an unknown channel or plot type.

Raised at the point the spec is built rather than at render time. A misspelled channel that fell through to “no column on x” would draw a plausible-looking chart of the wrong thing, which is worse than a traceback next to the line that caused it.

Initialize self. See help(type(self)) for accurate signature.

class spacr.qt.widgets.graph_spec.FacetGrid[source]

The complete panel layout — including the combinations with no rows.

An empty panel is drawn empty, never skipped. “Plate 3 / row H has no surviving cells” and “there is no plate 3 / row H” are different facts, and a grid that closes up the gaps tells the reader the second one when the first is true.

Parameters:
  • row_column – column split into panel rows, or None when rows are not faceted.

  • col_column – column split into panel columns, or None when columns are not faceted.

  • row_levels – level of each panel row, in order; (None,) when rows are not faceted.

  • col_levels – level of each panel column, in order; (None,) when columns are not faceted.

  • panels – every panel, row by row, including those with no rows.

  • hidden_rows – rows excluded because their facet level did not make the MAX_FACET_LEVELS cut. Non-zero means the grid is not the whole table, and notice says so.

panel(row: int, col: int) → FacetPanel[source]

The panel at one grid position.

Parameters:
  • row – the grid row, from 0.

  • col – the grid column, from 0.

Returns:

the panel.

property is_faceted: bool[source]

Whether this is a grid rather than a single chart.

Returns:

True when either axis has levels.

property n_panels: int[source]

Rows × columns — the count including empty panels.

property shape: Tuple[int, int][source]

The grid’s size as (rows, columns) of facet levels.

Returns:

the row and column counts.

class spacr.qt.widgets.graph_spec.FacetPanel[source]

One panel of the grid, and the rows that belong in it.

index holds positional indices into the frame facet_grid() was given, not label indices: a measurement frame carries a duplicated or reset index often enough that positions are the only safe currency.

Parameters:
  • row – zero-based grid row of the panel.

  • col – zero-based grid column of the panel.

  • row_level – facet-row level this panel shows, or None when rows are not faceted.

  • col_level – facet-column level this panel shows, or None when columns are not faceted.

  • index – positional indices of the panel’s rows in the faceted frame.

frame(source: pandas.DataFrame) → pandas.DataFrame[source]

This panel’s rows out of source.

Parameters:

source – the frame the grid was built from, or one with the same row positions; rows are taken by position.

title() → str[source]

The panel’s own label, empty when the grid is not faceted.

property is_empty: bool[source]

No rows. Still drawn — see FacetGrid.

property n: int[source]

How many rows landed in this panel.

Returns:

the row count.

class spacr.qt.widgets.graph_spec.GraphSpec[source]

Which column is on which channel, and what to draw.

Frozen: every edit returns a new spec (with_channel(), with_kind()), so the panel can keep a history for undo and two specs can be compared with ==.

Parameters:
  • x – column on the horizontal axis, or None.

  • y – column on the vertical axis, or None.

  • colour – column mapped to hue. Categorical → the fixed eight-colour order; continuous → one-hue light-to-dark ramp.

  • size – column mapped to mark area. Only meaningful for point marks; an aggregate plot ignores it rather than pretending.

  • facet_row – column whose levels become rows of panels.

  • facet_col – column whose levels become columns of panels.

  • kind – one of PLOT_KINDS, or None to infer it from what was dropped (infer_kind()). None and “the inferred value” are kept apart on purpose: a spec that inferred scatter and one that pinned scatter behave differently the moment the user drags a categorical column onto x.

  • roles – per-column override of column_kinds() — {"cell_count": "continuous"}. The table-wide rule is a good guess and not always the right one, and the alternative to an override here is the user editing the table.

  • bins – histogram / density bin count per axis.

  • shared_x – every panel gets the same x limits. Default on.

  • shared_y – likewise for y.

  • point_budget – individual marks before prepare_data() switches to binning or sampling.

  • seed – the sampler’s seed, so “the same chart” is the same chart — a screenshot in a report and the screen it came from must not differ by a random draw.

Raises:

SpecError – on an unknown plot kind, a non-positive bin count or a role that is neither continuous nor categorical.

__post_init__() → None[source]

Normalise the channels and validate the plot settings.

"" and None both mean an empty zone and are normalised to None, which is what lets if spec.x: be the whole test everywhere else.

Raises:

SpecError – if the plot kind is not one this module offers – None is allowed and means “infer it from the columns dropped”; if a role override is neither continuous nor categorical; or if bins is below 1.

binding_error(kinds: Mapping[str, str]) → str[source]

Explain incomplete or incompatible bindings, or return an empty string.

A pinned plot kind survives channel edits. Its temporary inability to draw is an editable state, not an invalid spec or a reason to silently choose another chart. The map must describe all available columns, including role overrides, as returned by kinds_for().

Parameters:

kinds – column kinds for the current table.

Returns:

an actionable canvas message, or "" when drawable.

Explicit bars also support a low-cardinality numeric Y (classified as categorical) and fall back to counts when Y has no numeric values.

column_for(channel: str) → str | None[source]

The column on channel.

Parameters:

channel – one of CHANNELS.

Raises:

SpecError – on an unknown channel.

describe(kinds: Mapping[str, str] | None = None) → str[source]

One human line, for the chart’s caption and the window title.

classmethod from_dict(payload: Mapping[str, Any]) → GraphSpec[source]

Rebuild from to_dict().

Unknown keys are ignored and missing keys take their defaults, so a spec written by an older (or newer) build still opens. A spec that will not load is a chart the user cannot get back.

Parameters:

payload – mapping written by to_dict(); only channel names and kind, roles, bins, shared_x, shared_y, point_budget and seed are read, and the constructor still validates them.

classmethod from_json(text: str) → GraphSpec[source]

Rebuild a spec from JSON text.

Parameters:

text – the JSON text.

Returns:

the rebuilt spec.

kinds_for(frame: pandas.DataFrame) → Dict[str, str][source]

column_kinds() of frame with this spec’s overrides applied.

Parameters:

frame – the table whose columns are classified; overrides for columns it lacks are ignored.

resolved_kind(kinds: Mapping[str, str]) → str[source]

The kind that will actually be drawn.

THE PIN, THEN THE SETTING, THEN THE INFERENCE. An explicit pin is the user saying “this chart, now” and still wins outright. With no pin, the Default Graph Type preference chooses among the forms that fit, and the inference answers when there is no preference or the chosen form cannot be drawn here – see _kind_and_note().

Parameters:

kinds – {column: kind} map as returned by column_kinds() or GraphSpec.kinds_for().

to_dict() → Dict[str, Any][source]

A plain JSON-able dict. Every field, always — a stable schema beats a compact one for something later screens read.

to_json() → str[source]

This spec as JSON text, keys sorted so the file is diffable.

Returns:

the JSON text.

used_columns() → Tuple[str, ...][source]

Every column named by a channel, de-duplicated, in channel order.

with_channel(channel: str, column: str | None) → GraphSpec[source]

A copy with column on channel (None empties the zone).

Parameters:
  • channel – one of CHANNELS; anything else raises SpecError.

  • column – column name, converted with str(); None or an empty string empties the zone.

with_kind(kind: str | None) → GraphSpec[source]

A copy pinned to kind, or back to inferring when None.

Parameters:

kind – one of PLOT_KINDS to pin, or None to infer it again; the constructor raises SpecError for any other value.

with_role(column: str, role: str | None) → GraphSpec[source]

A copy treating column as role; None restores the rule.

Parameters:
  • column – column name to override; converted with str().

  • role – "continuous" or "categorical", or None to drop the override; the constructor raises SpecError for any other value.

property channels: Dict[str, str | None][source]

{channel: column or None} for all six, in CHANNELS order.

property is_empty: bool[source]

there is no chart to draw yet.

Type:

Nothing on x and nothing on y

class spacr.qt.widgets.graph_spec.RenderData[source]

What the renderer should draw, and what it must say about it.

notice is not optional decoration. It is the difference between a chart of a million cells and a chart of fifty thousand of them, and a screenshot that does not carry it is a result nobody can check.

Parameters:
  • frame – the rows to draw: the whole table, or a random sample of it when strategy is SAMPLED.

  • strategy – FULL, BINNED or SAMPLED: how the rows are drawn.

  • n_total – number of rows in the table given to prepare_data().

  • n_shown – number of rows actually drawn or, for a binned density, counted.

property is_complete: bool[source]

Whether every row is represented — as a mark or in a bin.

class spacr.qt.widgets.graph_spec.Scales[source]

One set of limits and orders for every panel.

Computed over the whole frame, never per panel: that is what “shared axes” means, and it is why the later trellis screen can reuse this instead of re-deriving it. When spec.shared_x is off, the renderer autoscales x per panel and x_limits is ignored — the field is still filled, so a caption can say what the shared limits would have been.

Parameters:

count_limit – the tallest bar or bin across all panels, so a histogram grid’s count axis is comparable too. Sharing the value axis of an aggregate is the same rule as sharing a data axis; forgetting it is the usual way a faceted histogram lies.

x_positions() → Dict[str, int] | None[source]

Level → tick position, so a categorical x lines up across panels.

y_positions() → Dict[str, int] | None[source]

Where each categorical y level sits on the axis.

None for a continuous y, which is the distinction a renderer needs: a category is drawn at an integer position it was assigned, a number at the position it IS.

Returns:

{level: position}, or None when y is continuous.

spacr.qt.widgets.graph_spec.brush_mask(frame: pandas.DataFrame, spec: GraphSpec, kinds: Mapping[str, str], x0: float, y0: float, x1: float, y1: float, scales: Scales | None = None) → numpy.ndarray[source]

Rows of frame inside the rectangle a user dragged on a panel.

A brush is a predicate, not a hit test against drawn marks, which is why it stays exact when the panel was binned or sampled: the rectangle is evaluated against whatever frame it is handed, and the renderer hands it the unsampled one.

A categorical axis is matched by the level under the swept tick positions, so brushing three boxes of a box plot selects those three groups.

On a histogram or a bar chart the vertical axis is a count, not a variable, so only the horizontal sweep constrains anything — brushing across four bins selects the rows in those four bins, whatever height the drag happened to start at.

Parameters:
  • frame – the rows to test; hand it the unsampled panel rows so the result names every row in the rectangle.

  • spec – the chart specification; it decides which column each axis carries.

  • kinds – {column: kind} map as returned by column_kinds() or GraphSpec.kinds_for().

  • x0 – horizontal start of the dragged rectangle in the panel’s data coordinates; on a categorical axis these are tick positions, one per level. The two ends of each axis may come in either order.

  • y0 – vertical start of the rectangle; ignored, like y1, on a histogram or bar chart.

  • x1 – horizontal end of the rectangle.

  • y1 – vertical end of the rectangle.

spacr.qt.widgets.graph_spec.column_kinds(frame: pandas.DataFrame) → Dict[str, str][source]

Sort frame’s columns into CONTINUOUS / CATEGORICAL / UNPLOTTABLE.

A thin re-reading of spacr.qt.widgets.data_filter_panel.classify_columns() — the Local Data Filter’s rule — rather than a second classifier. The two screens must agree about what cell_count is, or a column offered as a tick list in the filter and as a continuous axis in the plot would give a user two different mental models of the same table.

The translation is one-to-one: a column the filter offers as a range is continuous; one it offers as ticks is categorical; one it skips (high-cardinality free text, or a key that identifies rather than describes) is not worth an axis either.

Parameters:

frame – the table whose columns are classified.

spacr.qt.widgets.graph_spec.facet_grid(frame: pandas.DataFrame, spec: GraphSpec, *, levels_source: pandas.DataFrame | None = None, max_levels: int = MAX_FACET_LEVELS, max_panels: int = MAX_PANELS) → FacetGrid[source]

Split frame into the grid spec’s facet channels describe.

Parameters:
  • frame – the rows to place into panels (post-filter, post-sample).

  • spec – graph channels and facet columns that define the grid.

  • levels_source – where the levels come from, when that is not frame. The renderer passes the pre-sample frame, so a level that exists in the population but drew no rows in the sample still gets its panel — drawn empty, which is the honest picture — instead of the grid silently changing shape with the sample.

  • max_levels – per axis; beyond it levels are cut and counted.

  • max_panels – hard ceiling on rows × columns.

Returns:

a FacetGrid whose panels are the full cartesian product in row-major order.

spacr.qt.widgets.graph_spec.infer_kind(spec: GraphSpec, kinds: Mapping[str, str]) → str[source]

The plot type the dropped columns imply.

Only x and y decide. Colour, size and the facets never change what kind of chart this is — dragging gene onto colour must not silently turn a scatter into something else, or the chart the user built stops being the chart they are looking at.

x, y

kind

nothing

EMPTY

one continuous

HISTOGRAM

one categorical

BAR (counts per level)

two continuous

SCATTER

one of each

BOX

two categorical

HEATMAP (a contingency count)

VIOLIN and LINE are reachable only as an explicit override: a violin claims a density estimate the data may not support at small n, and joining points with a line asserts an order between them that a measurement table does not have.

Parameters:
  • spec – the chart specification; only its x and y columns are read.

  • kinds – {column: kind} map as returned by column_kinds() or GraphSpec.kinds_for(); a column missing from it is treated as categorical.

spacr.qt.widgets.graph_spec.plottable_columns(frame: pandas.DataFrame) → Tuple[str, ...][source]

The columns worth offering in the drag well, sorted.

Same rule, same reason as the filter panel’s picker: a measurement table has hundreds of columns, and offering all of them is the same as offering none.

Parameters:

frame – the table whose columns are offered.

spacr.qt.widgets.graph_spec.prepare_data(frame: pandas.DataFrame, spec: GraphSpec, kinds: Mapping[str, str]) → RenderData[source]

Decide how many rows get drawn, and say so.

See the module docstring for the policy. The short version: aggregates use everything, point plots use everything up to spec.point_budget, and above that they either bin (nothing lost) or sample (said out loud).

Parameters:
  • frame – the filtered table to draw.

  • spec – the chart specification; its kind, point_budget, bins, colour and size channels and seed decide the strategy.

  • kinds – {column: kind} map as returned by column_kinds() or GraphSpec.kinds_for().

spacr.qt.widgets.graph_spec.scales_for(frame: pandas.DataFrame, spec: GraphSpec, kinds: Mapping[str, str], grid: FacetGrid | None = None) → Scales[source]

Limits, bin edges, level orders and colour levels shared by every panel.

Parameters:
  • frame – the rows that will be drawn — post-filter and post-sample, so the limits bound what is actually on screen.

  • spec – graph channels, kind and bin count to scale.

  • kinds – mapping from column names to continuous or categorical axis kinds.

  • grid – needed only for Scales.count_limit, which is the maximum over panels and therefore cannot be computed from the frame alone.

spacr.qt.widgets.graph_spec.value_axes(spec: GraphSpec, kinds: Mapping[str, str]) → Tuple[str | None, str | None][source]

Which column each axis actually carries, once the kind is known.

Almost always (spec.x, spec.y). The exception is a lone column on Y that infers a histogram or a bar chart: those draw the column along the horizontal axis and counts up the vertical one, whichever zone it was dropped in. Without this the scales would be computed for an axis nothing is drawn on, panels would autoscale independently, and a faceted “histogram of Y” would quietly stop sharing its bins.

Parameters:

Nested helpers

facet_grid.axis(column: str | None) → Tuple[Tuple[str | None, ...], int]

One facet axis’s levels and how many there are.

spacr/qt/widgets/graph_spec.py:850