spacr.qt.widgets.graph_spec¶
The Graph Builder’s spec: what is on which channel, and nothing else.
A JMP-style graph builder is two things that want very much to be one thing:
a pile of drag-and-drop chrome, and a small declarative object saying “area
is on x, gene is on colour, facet by plate down the rows”. This module is
the second one, kept deliberately apart from the first.
Why the separation is the important decision¶
Four later screens — small multiples, the gate editor, the feature explorer and
the campaign control charts — are all “a Graph Builder with one extra rule”.
If the axis and facet logic lives inside a QWidget, each of them either
re-derives it or inherits a widget it does not want. So everything in here is
pure pandas and numpy, with no Qt, like
spacr.selection— testable without a display, usable from a notebook or fromspacr-run;serialisable —
GraphSpec.to_dict()round-trips through JSON, so a chart is something a settings file, a report or a macro can carry;immutable — every channel change returns a new spec, which is what makes “undo the last drag” and “compare two specs” trivial rather than a diffing problem.
The four pieces¶
GraphSpecWhich column is on which of the six channels (x, y, colour, size, facet-row, facet-column), the plot type (or
Nonefor “infer it”), and the handful of options that change what is computed rather than what it looks like.facet_grid()The panel layout — the full cartesian product of the row and column levels, empty combinations included. A missing panel and an empty panel say different things (“this table has no plate 3 / row H” versus “plate 3 row H was measured and nothing survived the filter”), and only one of them is true; drawing the grid complete is what keeps the reader from guessing.
scales_for()One set of limits, bin edges, category orders and colour levels for every panel. Shared axes are the default because comparing across panels is the entire point of faceting, and two panels whose y axes differ by an order of magnitude look identical.
prepare_data()The large-data policy, stated rather than hidden. See below.
Large data¶
spaCR measurement tables run to 10^5–10^6 object rows and nobody wants a
scatter of a million overlapping dots. Three strategies, chosen by what the
plot actually needs, and the chosen one is always named in
RenderData.notice so a subset can never be mistaken for the whole:
aggregate plots use every row, always. A histogram, bar, box, violin or heatmap is already a reduction — sampling before aggregating would change the answer for no gain, so
AGGREGATE_KINDSnever sample regardless of size.scatter/line up to
DEFAULT_POINT_BUDGETrows: every row is a mark.scatter above the budget: 2-D density binning — every row is counted into a shared-edge 2-D histogram and the panel is drawn as a raster. Nothing is dropped, the density is quantitative, and a brush is still exact because a rectangle brush is a predicate on x and y evaluated against the full frame, not against the pixels.
scatter above the budget where binning cannot answer the question — a categorical colour or a size channel needs per-point marks — falls back to a seeded uniform sample, with the count, the fraction and the word “sampled” in the notice.
The one thing not on offer is quietly plotting the head of the frame.
Exceptions¶
A spec that cannot mean anything — an unknown channel or plot type. |
Classes¶
The complete panel layout — including the combinations with no rows. |
|
One panel of the grid, and the rows that belong in it. |
|
Which column is on which channel, and what to draw. |
|
What the renderer should draw, and what it must say about it. |
|
One set of limits and orders for every panel. |
Functions¶
|
Rows of |
|
Sort |
|
Split |
|
The plot type the dropped columns imply. |
|
The columns worth offering in the drag well, sorted. |
|
Decide how many rows get drawn, and say so. |
|
Limits, bin edges, level orders and colour levels shared by every panel. |
|
Which column each axis actually carries, once the kind is known. |
Module Contents¶
- exception spacr.qt.widgets.graph_spec.SpecError[source]¶
Bases:
ValueErrorA spec that cannot mean anything — an unknown channel or plot type.
Raised at the point the spec is built rather than at render time. A misspelled channel that fell through to “no column on x” would draw a plausible-looking chart of the wrong thing, which is worse than a traceback next to the line that caused it.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.qt.widgets.graph_spec.FacetGrid[source]¶
The complete panel layout — including the combinations with no rows.
An empty panel is drawn empty, never skipped. “Plate 3 / row H has no surviving cells” and “there is no plate 3 / row H” are different facts, and a grid that closes up the gaps tells the reader the second one when the first is true.
- Parameters:
row_column – column split into panel rows, or None when rows are not faceted.
col_column – column split into panel columns, or None when columns are not faceted.
row_levels – level of each panel row, in order;
(None,)when rows are not faceted.col_levels – level of each panel column, in order;
(None,)when columns are not faceted.panels – every panel, row by row, including those with no rows.
hidden_rows – rows excluded because their facet level did not make the
MAX_FACET_LEVELScut. Non-zero means the grid is not the whole table, andnoticesays so.
- panel(row: int, col: int) FacetPanel[source]¶
The panel at one grid position.
- Parameters:
row – the grid row, from 0.
col – the grid column, from 0.
- Returns:
the panel.
- class spacr.qt.widgets.graph_spec.FacetPanel[source]¶
One panel of the grid, and the rows that belong in it.
indexholds positional indices into the framefacet_grid()was given, not label indices: a measurement frame carries a duplicated or reset index often enough that positions are the only safe currency.- Parameters:
row – zero-based grid row of the panel.
col – zero-based grid column of the panel.
row_level – facet-row level this panel shows, or None when rows are not faceted.
col_level – facet-column level this panel shows, or None when columns are not faceted.
index – positional indices of the panel’s rows in the faceted frame.
- frame(source: pandas.DataFrame) pandas.DataFrame[source]¶
This panel’s rows out of
source.- Parameters:
source – the frame the grid was built from, or one with the same row positions; rows are taken by position.
- class spacr.qt.widgets.graph_spec.GraphSpec[source]¶
Which column is on which channel, and what to draw.
Frozen: every edit returns a new spec (
with_channel(),with_kind()), so the panel can keep a history for undo and two specs can be compared with==.- Parameters:
x – column on the horizontal axis, or
None.y – column on the vertical axis, or
None.colour – column mapped to hue. Categorical → the fixed eight-colour order; continuous → one-hue light-to-dark ramp.
size – column mapped to mark area. Only meaningful for point marks; an aggregate plot ignores it rather than pretending.
facet_row – column whose levels become rows of panels.
facet_col – column whose levels become columns of panels.
kind – one of
PLOT_KINDS, orNoneto infer it from what was dropped (infer_kind()).Noneand “the inferred value” are kept apart on purpose: a spec that inferredscatterand one that pinnedscatterbehave differently the moment the user drags a categorical column onto x.roles – per-column override of
column_kinds()—{"cell_count": "continuous"}. The table-wide rule is a good guess and not always the right one, and the alternative to an override here is the user editing the table.bins – histogram / density bin count per axis.
shared_x – every panel gets the same x limits. Default on.
shared_y – likewise for y.
point_budget – individual marks before
prepare_data()switches to binning or sampling.seed – the sampler’s seed, so “the same chart” is the same chart — a screenshot in a report and the screen it came from must not differ by a random draw.
- Raises:
SpecError – on an unknown plot kind, a non-positive bin count or a role that is neither continuous nor categorical.
- __post_init__() None[source]¶
Normalise the channels and validate the plot settings.
""andNoneboth mean an empty zone and are normalised toNone, which is what letsif spec.x:be the whole test everywhere else.- Raises:
SpecError – if the plot kind is not one this module offers –
Noneis allowed and means “infer it from the columns dropped”; if a role override is neither continuous nor categorical; or ifbinsis below 1.
- binding_error(kinds: Mapping[str, str]) str[source]¶
Explain incomplete or incompatible bindings, or return an empty string.
A pinned plot kind survives channel edits. Its temporary inability to draw is an editable state, not an invalid spec or a reason to silently choose another chart. The map must describe all available columns, including role overrides, as returned by
kinds_for().- Parameters:
kinds – column kinds for the current table.
- Returns:
an actionable canvas message, or
""when drawable.
Explicit bars also support a low-cardinality numeric Y (classified as categorical) and fall back to counts when Y has no numeric values.
- column_for(channel: str) str | None[source]¶
The column on
channel.- Parameters:
channel – one of
CHANNELS.- Raises:
SpecError – on an unknown channel.
- describe(kinds: Mapping[str, str] | None = None) str[source]¶
One human line, for the chart’s caption and the window title.
- classmethod from_dict(payload: Mapping[str, Any]) GraphSpec[source]¶
Rebuild from
to_dict().Unknown keys are ignored and missing keys take their defaults, so a spec written by an older (or newer) build still opens. A spec that will not load is a chart the user cannot get back.
- Parameters:
payload – mapping written by
to_dict(); only channel names andkind,roles,bins,shared_x,shared_y,point_budgetandseedare read, and the constructor still validates them.
- classmethod from_json(text: str) GraphSpec[source]¶
Rebuild a spec from JSON text.
- Parameters:
text – the JSON text.
- Returns:
the rebuilt spec.
- kinds_for(frame: pandas.DataFrame) Dict[str, str][source]¶
column_kinds()offramewith this spec’s overrides applied.- Parameters:
frame – the table whose columns are classified; overrides for columns it lacks are ignored.
- resolved_kind(kinds: Mapping[str, str]) str[source]¶
The kind that will actually be drawn.
THE PIN, THEN THE SETTING, THEN THE INFERENCE. An explicit pin is the user saying “this chart, now” and still wins outright. With no pin, the Default Graph Type preference chooses among the forms that fit, and the inference answers when there is no preference or the chosen form cannot be drawn here – see
_kind_and_note().- Parameters:
kinds –
{column: kind}map as returned bycolumn_kinds()orGraphSpec.kinds_for().
- to_dict() Dict[str, Any][source]¶
A plain JSON-able dict. Every field, always — a stable schema beats a compact one for something later screens read.
- to_json() str[source]¶
This spec as JSON text, keys sorted so the file is diffable.
- Returns:
the JSON text.
- used_columns() Tuple[str, ...][source]¶
Every column named by a channel, de-duplicated, in channel order.
- with_channel(channel: str, column: str | None) GraphSpec[source]¶
A copy with
columnonchannel(Noneempties the zone).- Parameters:
channel – one of
CHANNELS; anything else raisesSpecError.column – column name, converted with
str(); None or an empty string empties the zone.
- with_kind(kind: str | None) GraphSpec[source]¶
A copy pinned to
kind, or back to inferring whenNone.- Parameters:
kind – one of
PLOT_KINDSto pin, or None to infer it again; the constructor raisesSpecErrorfor any other value.
- with_role(column: str, role: str | None) GraphSpec[source]¶
A copy treating
columnasrole;Nonerestores the rule.- Parameters:
column – column name to override; converted with
str().role –
"continuous"or"categorical", or None to drop the override; the constructor raisesSpecErrorfor any other value.
- class spacr.qt.widgets.graph_spec.RenderData[source]¶
What the renderer should draw, and what it must say about it.
noticeis not optional decoration. It is the difference between a chart of a million cells and a chart of fifty thousand of them, and a screenshot that does not carry it is a result nobody can check.- Parameters:
frame – the rows to draw: the whole table, or a random sample of it when
strategyisSAMPLED.strategy –
FULL,BINNEDorSAMPLED: how the rows are drawn.n_total – number of rows in the table given to
prepare_data().n_shown – number of rows actually drawn or, for a binned density, counted.
- class spacr.qt.widgets.graph_spec.Scales[source]¶
One set of limits and orders for every panel.
Computed over the whole frame, never per panel: that is what “shared axes” means, and it is why the later trellis screen can reuse this instead of re-deriving it. When
spec.shared_xis off, the renderer autoscales x per panel andx_limitsis ignored — the field is still filled, so a caption can say what the shared limits would have been.- Parameters:
count_limit – the tallest bar or bin across all panels, so a histogram grid’s count axis is comparable too. Sharing the value axis of an aggregate is the same rule as sharing a data axis; forgetting it is the usual way a faceted histogram lies.
- x_positions() Dict[str, int] | None[source]¶
Level → tick position, so a categorical x lines up across panels.
- y_positions() Dict[str, int] | None[source]¶
Where each categorical y level sits on the axis.
None for a continuous y, which is the distinction a renderer needs: a category is drawn at an integer position it was assigned, a number at the position it IS.
- Returns:
{level: position}, or None when y is continuous.
- spacr.qt.widgets.graph_spec.brush_mask(frame: pandas.DataFrame, spec: GraphSpec, kinds: Mapping[str, str], x0: float, y0: float, x1: float, y1: float, scales: Scales | None = None) numpy.ndarray[source]¶
Rows of
frameinside the rectangle a user dragged on a panel.A brush is a predicate, not a hit test against drawn marks, which is why it stays exact when the panel was binned or sampled: the rectangle is evaluated against whatever frame it is handed, and the renderer hands it the unsampled one.
A categorical axis is matched by the level under the swept tick positions, so brushing three boxes of a box plot selects those three groups.
On a histogram or a bar chart the vertical axis is a count, not a variable, so only the horizontal sweep constrains anything — brushing across four bins selects the rows in those four bins, whatever height the drag happened to start at.
- Parameters:
frame – the rows to test; hand it the unsampled panel rows so the result names every row in the rectangle.
spec – the chart specification; it decides which column each axis carries.
kinds –
{column: kind}map as returned bycolumn_kinds()orGraphSpec.kinds_for().x0 – horizontal start of the dragged rectangle in the panel’s data coordinates; on a categorical axis these are tick positions, one per level. The two ends of each axis may come in either order.
y0 – vertical start of the rectangle; ignored, like
y1, on a histogram or bar chart.x1 – horizontal end of the rectangle.
y1 – vertical end of the rectangle.
- spacr.qt.widgets.graph_spec.column_kinds(frame: pandas.DataFrame) Dict[str, str][source]¶
Sort
frame’s columns intoCONTINUOUS/CATEGORICAL/UNPLOTTABLE.A thin re-reading of
spacr.qt.widgets.data_filter_panel.classify_columns()— the Local Data Filter’s rule — rather than a second classifier. The two screens must agree about whatcell_countis, or a column offered as a tick list in the filter and as a continuous axis in the plot would give a user two different mental models of the same table.The translation is one-to-one: a column the filter offers as a range is continuous; one it offers as ticks is categorical; one it skips (high-cardinality free text, or a key that identifies rather than describes) is not worth an axis either.
- Parameters:
frame – the table whose columns are classified.
- spacr.qt.widgets.graph_spec.facet_grid(frame: pandas.DataFrame, spec: GraphSpec, *, levels_source: pandas.DataFrame | None = None, max_levels: int = MAX_FACET_LEVELS, max_panels: int = MAX_PANELS) FacetGrid[source]¶
Split
frameinto the gridspec’s facet channels describe.- Parameters:
frame – the rows to place into panels (post-filter, post-sample).
spec – graph channels and facet columns that define the grid.
levels_source – where the levels come from, when that is not
frame. The renderer passes the pre-sample frame, so a level that exists in the population but drew no rows in the sample still gets its panel — drawn empty, which is the honest picture — instead of the grid silently changing shape with the sample.max_levels – per axis; beyond it levels are cut and counted.
max_panels – hard ceiling on rows × columns.
- Returns:
a
FacetGridwhosepanelsare the full cartesian product in row-major order.
- spacr.qt.widgets.graph_spec.infer_kind(spec: GraphSpec, kinds: Mapping[str, str]) str[source]¶
The plot type the dropped columns imply.
Only x and y decide. Colour, size and the facets never change what kind of chart this is — dragging
geneonto colour must not silently turn a scatter into something else, or the chart the user built stops being the chart they are looking at.x, y
kind
nothing
EMPTYone continuous
HISTOGRAMone categorical
BAR(counts per level)two continuous
SCATTERone of each
BOXtwo categorical
HEATMAP(a contingency count)VIOLINandLINEare reachable only as an explicit override: a violin claims a density estimate the data may not support at small n, and joining points with a line asserts an order between them that a measurement table does not have.- Parameters:
spec – the chart specification; only its
xandycolumns are read.kinds –
{column: kind}map as returned bycolumn_kinds()orGraphSpec.kinds_for(); a column missing from it is treated as categorical.
- spacr.qt.widgets.graph_spec.plottable_columns(frame: pandas.DataFrame) Tuple[str, ...][source]¶
The columns worth offering in the drag well, sorted.
Same rule, same reason as the filter panel’s picker: a measurement table has hundreds of columns, and offering all of them is the same as offering none.
- Parameters:
frame – the table whose columns are offered.
- spacr.qt.widgets.graph_spec.prepare_data(frame: pandas.DataFrame, spec: GraphSpec, kinds: Mapping[str, str]) RenderData[source]¶
Decide how many rows get drawn, and say so.
See the module docstring for the policy. The short version: aggregates use everything, point plots use everything up to
spec.point_budget, and above that they either bin (nothing lost) or sample (said out loud).- Parameters:
frame – the filtered table to draw.
spec – the chart specification; its kind,
point_budget,bins, colour and size channels andseeddecide the strategy.kinds –
{column: kind}map as returned bycolumn_kinds()orGraphSpec.kinds_for().
- spacr.qt.widgets.graph_spec.scales_for(frame: pandas.DataFrame, spec: GraphSpec, kinds: Mapping[str, str], grid: FacetGrid | None = None) Scales[source]¶
Limits, bin edges, level orders and colour levels shared by every panel.
- Parameters:
frame – the rows that will be drawn — post-filter and post-sample, so the limits bound what is actually on screen.
spec – graph channels, kind and bin count to scale.
kinds – mapping from column names to continuous or categorical axis kinds.
grid – needed only for
Scales.count_limit, which is the maximum over panels and therefore cannot be computed from the frame alone.
- spacr.qt.widgets.graph_spec.value_axes(spec: GraphSpec, kinds: Mapping[str, str]) Tuple[str | None, str | None][source]¶
Which column each axis actually carries, once the kind is known.
Almost always
(spec.x, spec.y). The exception is a lone column on Y that infers a histogram or a bar chart: those draw the column along the horizontal axis and counts up the vertical one, whichever zone it was dropped in. Without this the scales would be computed for an axis nothing is drawn on, panels would autoscale independently, and a faceted “histogram of Y” would quietly stop sharing its bins.- Parameters:
spec – the chart specification whose
xandycolumns are read.kinds –
{column: kind}map as returned bycolumn_kinds()orGraphSpec.kinds_for().