spacr.qt.widgets.pca_model

PCA over a measurement table — the statistics, with the decisions written down.

spaCR measures hundreds of features per object and, until this module, had no principal components anywhere. The reason it is worth having is not dimensionality reduction for its own sake: it is the one view that answers “these two populations separate — on what?” in a single picture, because the loadings say which measurements carry the separation.

PCA is also easy to get subtly wrong on exactly this kind of table, in ways that produce a confident-looking plot of nothing. Three decisions do all the damage, so all three are made here, in the open, and reported back to the user with every result.

1. Centring and scaling

Always centred. Scaled to unit variance by default (SCALE_ZSCORE), with SCALE_NONE available and never implicit.

PCA maximises variance, and variance has units. In a spaCR object table cell_area is px² and runs to 10^4 with a variance near 10^6; eccentricity is dimensionless and lives in [0, 1]; channel_1_mean_intensity is counts. An unscaled (covariance) PCA of those three returns “PC1 = cell_area” every time — not because size is the interesting axis but because px² is a big number. Worse, the answer is not even stable: recording the same areas in µm² instead of px² would rotate every component.

Standardising each feature to mean 0 and sample SD 1 (a correlation PCA) makes the result invariant to any per-feature linear rescaling, which is the only defensible default when the columns have no common unit.

It is not free, and the cost is stated rather than hidden: equal weight means forty near-duplicate Zernike moments outvote one carefully chosen intensity ratio, and a feature that is pure noise gets the same unit of variance as a feature that is pure signal. SCALE_NONE is there for the case where the user has already put the features in a common unit and wants the big ones to dominate.

2. NaN

Never imputed by default, and never zero-filled at all.

A NaN in a spaCR measurement table is usually structural, not missing at random — the same reasoning spacr.anndata_export sets out at length. A pathogen_* column is NaN for a cell with no pathogen in it; that is a fact about the cell, not a gap in the record. Mean-imputing it invents a pathogen of average size and moves an uninfected cell into the middle of the infected distribution, invisibly.

Dropping rows instead is not neutral either: complete-case analysis over a feature set containing one pathogen_* column silently deletes every uninfected cell, so the result is the PCA of the infected cells presented as the PCA of the cells.

The three failure modes are different in kind, which is what the policies are built around:

  • dropping a feature changes the space — a smaller question, answered exactly;

  • dropping a row changes the population — the same question, answered about a self-selected subset;

  • imputing changes the data — the original question, answered about numbers nobody measured.

NAN_AUTO, the default, picks between the first two by how the missingness is shaped, because in this table the two shapes are far apart: a structural NaN affects tens of percent of the rows (every uninfected cell), a sporadic one affects a handful (one object where mahotas returned nothing). So a feature missing in more than DEFAULT_STRUCTURAL_MISSING of the rows is treated as structural and dropped by name; rows still carrying a NaN among the survivors are dropped as sporadic and counted. Nothing is imputed, and both actions appear in PCAResult.notes and in PCAResult.report().

NAN_COMPLETE, NAN_DROP_FEATURES and NAN_MEAN are the explicit versions of the three failure modes above, for a user who knows which one they want. NAN_MEAN is offered and never chosen for them.

Infinities are missing under every policy. Ratio features produce ±inf from a zero denominator, an inf survives dropna, and one of them destroys the scaling and therefore every component. They become NaN before the policy runs and are counted separately in PCAResult.n_infinite.

3. Constant and collinear columns

A constant column has SD 0, so standardising it is 0/0. Left in, it contributes a NaN that spreads through the whole decomposition; “fixed” by substituting a 1, it becomes a zero-variance direction that the solver is free to return as a component with an arbitrary loading pattern. It is dropped before the decomposition and named in PCAResult.dropped_features — a column with one value cannot explain variation in anything.

Perfectly collinear columns are a different problem and get a different answer: they are kept, because the feature list is the user’s statement of what they care about and picking which of area and area_um2 to discard is not this module’s call. What is refused is inventing components out of the redundancy. The numerical rank is computed from the singular values and no more than rank components are ever returned, so a table of 12 features spanning 9 real directions yields 9 components, not 12 — the last three would have ~0 explained variance and a direction chosen by floating-point noise. Groups of features that are perfectly correlated are also detected and reported, since a duplicated feature silently doubles its own weight in a correlation PCA.

Honest explained variance

PCAResult.explained_variance_ratio is over the total variance of the analysed matrix, so it sums to 1 across all rank components rather than across the ones that were kept. Alongside it, every result carries how much of the original table it is actually about (PCAResult.n_rows_in versus len(result), PCAResult.n_features_in versus its features): “PC1 explains 71%” means something different when the matrix is three features and 8% of the objects.

And the single most common way a PCA of morphology features misleads anyone is that PC1 is just size. When every feature loads with the same sign, PC1 is a general-magnitude axis — big objects at one end, small at the other — which is usually a fact about segmentation rather than biology. PCAResult.is_size_like() detects it from the loading signs and PCAResult.headline() says it in words, unprompted.

No Qt in here

Pure numpy and pandas, like spacr.qt.widgets.graph_spec and spacr.selection: usable from a notebook, testable without a display, and free for the later screens (the feature explorer, the gate editor) to reuse without inheriting a widget.

Exceptions

PCAError

A PCA that cannot mean anything, with the reason in the message.

Classes

PCAResult

A decomposition, plus everything needed to read it honestly.

PCASpec

What to decompose and how. Frozen and JSON round-tripping, like

Functions

candidate_features(→ Tuple[str, ...])

The columns worth offering as PCA features, sorted.

component_index(→ Optional[int])

'PC1' -> 0, and None for anything that is not a PC column.

component_name(→ str)

0 -> 'PC1'. The column name a score lands in, in one place.

pca(→ PCAResult)

Decompose frame's features under spec.

Module Contents

exception spacr.qt.widgets.pca_model.PCAError[source]

Bases: ValueError

A PCA that cannot mean anything, with the reason in the message.

Raised rather than returned as an empty result: every one of these is a sentence the user can act on (“every row has a NaN in pathogen_area”), and a caller that swallowed it would draw an empty scatter with no explanation. Screens catch it and show the message.

Initialize self. See help(type(self)) for accurate signature.

class spacr.qt.widgets.pca_model.PCAResult[source]

A decomposition, plus everything needed to read it honestly.

Parameters:
  • features – the columns actually decomposed, in matrix order. Not the ones asked for — see dropped_features.

  • rows – positional indices into the frame pca() was given, naming the objects that survived the NaN policy. Positional because a measurement frame carries a duplicated or reset index often enough that positions are the only safe currency, the same reason FacetPanel uses them.

  • scores – (len(rows), k) — each object’s coordinate on each component, in the units of the analysed (post-scaling) matrix.

  • loadings – (len(features), k), unit-norm columns. The direction of each component in feature space.

  • correlations – (len(features), k) Pearson r between each feature and each component score. This is what the biplot arrows are: unlike a unit-norm loading, an r is a number with a meaning on its own, and for a standardised PCA the squared row sums are each feature’s communality.

  • explained_variance – per component, in the analysed matrix’s units.

  • explained_variance_ratio – over the total variance of the analysed matrix — so it sums to 1 across all rank components, not across the ones returned.

  • total_variance – the analysed matrix’s total variance, summed over all components (not only the ones returned), in its own units.

  • centre – per feature, the mean subtracted before the decomposition, aligned with features.

  • scale – per feature, the divisor applied after centring: the standard deviation (ddof=1) under z-score scaling, else 1.

  • rank – the numerical rank. No more components than this exist.

  • n_rows_in – rows in the frame pca() was given, before the NaN policy.

  • n_features_in – requested features that exist in the frame, before the NaN policy and constant columns removed any.

  • scaling – the scaling mode used, one of SCALE_MODES.

  • nan_policy – the missing-value policy used, one of NAN_POLICIES.

  • dropped_features – {column: why}, for everything that did not reach the decomposition.

  • dropped_rows – objects removed by the NaN policy.

  • n_infinite – cells that were ±inf and were treated as missing.

  • collinear_groups – features that are perfectly correlated with each other. Kept in the analysis, reported because each such group weighs as many times as it has members.

  • notes – plain-language notes on what the analysis did, e.g. rank deficiency.

__len__() → int[source]

Objects in the analysis.

caveats() → Tuple[str, ...][source]

Everything a reader needs before believing the picture.

dominant(k: int = 0) → Tuple[str, float][source]

(feature, share) — the feature with the largest |loading| on component k, and how much of the component’s squared loading it carries.

A share near 1 means the component is that feature under another name, which is worth saying before anyone reads biology into it.

headline(k: int = 0) → str[source]

One sentence about component k, including the bad news.

is_degenerate(k: int = 0) → bool[source]

Whether component k holds so little variance that its direction is not identified.

Near-collinear features leave a numerically full-rank matrix whose last directions are floating-point residue: the solver returns a direction because it must, not because the data has one. A component under DEGENERATE_RATIO of the total variance is reported with that said, so nobody reads a loading pattern out of rounding error.

is_size_like(k: int = 0) → bool[source]

Whether component k is a general-magnitude axis.

Every feature loading with the same sign is the classic “size” factor of morphometrics: one end big objects, the other end small ones. It is real, it is usually PC1, and it is usually a statement about segmentation and focus rather than about biology — so it is detected and said out loud rather than left for the reader to notice.

Two guards keep it from crying wolf. It needs three features, because with two “same sign” is a coin flip. And it needs the loading spread over more than one of them: a component that is a single feature trivially agrees with itself about sign, and calling that a size axis would be a statement about arithmetic rather than about the objects.

loadings_frame() → pandas.DataFrame[source]

One row per feature: its loading and its correlation on each component. What “export the loadings” writes.

plane_features(kx: int = 0, ky: int = 1, count: int = 8) → Tuple[int, ...][source]

Indices of the features best represented in the (kx, ky) plane.

Ranked by r_x² + r_y² — how much of the feature is visible in this plane at all. Drawing the other four hundred arrows would say nothing except that the figure is full; drawing the longest ones is the standard choice because a short arrow means “this feature points somewhere you are not looking”, not “this feature does not matter”.

report() → str[source]

The whole story, as the panel shows it and a report file writes it.

scores_frame(source: pandas.DataFrame, *, components: int | None = None) → pandas.DataFrame[source]

source’s surviving rows with PC1…PCk columns added.

Every original column is kept, which is the whole point: it is what lets the scores plot colour by gene, facet by plateID and publish a real object-key selection when a cluster is brushed. The Graph Builder then draws it with no knowledge that it is a PCA.

Parameters:
  • source – the same frame pca() was given; its surviving rows are taken by position.

  • components – how many PC columns to add; None adds all, and other values are clamped between 1 and n_components.

sign_agreement(k: int = 0) → float[source]

How one-sided component k’s loadings are, in [0.5, 1].

The larger share of squared loading sitting on a single sign. 1.0 means every feature moves the same way along this axis.

top_features(k: int = 0, count: int = 5) → Tuple[Tuple[str, float], ...][source]

The count features loading hardest on component k, as (feature, loading) with the sign kept, strongest first.

variance_frame() → pandas.DataFrame[source]

One row per component — the scree plot’s data, and its CSV.

property component_names: Tuple[str, ...][source]

The components’ display names, in order.

Returns:

one name per component.

property cumulative_ratio: numpy.ndarray[source]

Running total of explained_variance_ratio.

property n_components: int[source]

How many components were kept.

Returns:

the component count.

property n_features: int[source]

How many columns went into the decomposition.

The columns ACTUALLY used, which can be fewer than the spec asked for: a constant or all-null column contributes no variance and is dropped before fitting.

Returns:

the feature count.

property retained_ratio: float[source]

Variance share of the components actually returned.

property row_share: float[source]

Fraction of the given objects the analysis is about.

class spacr.qt.widgets.pca_model.PCASpec[source]

What to decompose and how. Frozen and JSON round-tripping, like GraphSpec, so a saved analysis is something a settings file or a report can carry.

Parameters:
  • features – the columns to decompose. Empty means candidate_features() of whatever frame it is run against — which is a default, not a promise: the result records the features it actually used.

  • n_components – how many to compute. Capped at the numerical rank, always: components past the rank have no direction.

  • scaling – SCALE_ZSCORE (default) or SCALE_NONE.

  • nan_policy – one of NAN_POLICIES.

  • structural_missing – the NAN_AUTO threshold.

Raises:

PCAError – on an unknown scaling or policy, or a non-positive component count — at the point the spec is built, not at render time.

__post_init__() → None[source]

De-duplicate the features and validate the decomposition settings.

Raises:

PCAError – if scaling or nan_policy is not one this module offers, if n_components is below 1, or if structural_missing is not a fraction of rows in [0, 1]. The component count is capped at the module’s maximum rather than refused: asking for more components than that is a request for noise, not an error.

describe() → str[source]

One line, for a caption.

classmethod from_dict(payload: Mapping[str, Any]) → PCASpec[source]

Rebuild from to_dict(); unknown keys ignored, missing keys defaulted, so an analysis written by another build still opens.

Parameters:

payload – a mapping as to_dict() writes it; features, n_components, scaling, nan_policy and structural_missing are read and any other key is ignored.

classmethod from_json(text: str) → PCASpec[source]

Rebuild a spec from JSON text.

Parameters:

text – the JSON text.

Returns:

the rebuilt spec.

to_dict() → Dict[str, Any][source]

This spec as plain data.

Returns:

a JSON-safe dict.

to_json() → str[source]

This spec as JSON text, keys sorted so the file is diffable.

Returns:

the JSON text.

with_components(n: int) → PCASpec[source]

A copy keeping a different number of components.

Parameters:

n – how many to keep.

Returns:

the new spec.

with_features(features: Sequence[str]) → PCASpec[source]

A copy decomposing a different set of columns.

A COPY: a spec is a value, so the one a view is showing is never edited underneath it.

Parameters:

features – the columns to decompose.

Returns:

the new spec.

with_nan_policy(policy: str) → PCASpec[source]

A copy handling missing values differently.

Parameters:

policy – the policy’s name.

Returns:

the new spec.

with_scaling(scaling: str) → PCASpec[source]

A copy using a different scaling.

SCALING IS NOT COSMETIC HERE. PCA maximises variance, so unscaled columns in different units let the largest unit dominate every component – which is why this is a spec field and not a display option.

Parameters:

scaling – the scaling’s name.

Returns:

the new spec.

spacr.qt.widgets.pca_model.candidate_features(frame: pandas.DataFrame) → Tuple[str, ...][source]

The columns worth offering as PCA features, sorted.

Continuous by spacr.qt.widgets.graph_spec.column_kinds() — which is a re-reading of the Local Data Filter’s classifier, the one column classifier in this codebase. That rule already excludes object keys (they identify rather than describe) and small-cardinality numeric codes like cell_count and a class label, which are counts and labels rather than measured quantities and have no business setting the scale of a component.

A user who disagrees about a particular column puts it in PCASpec.features explicitly; nothing here refuses it.

Parameters:

frame – the measurement table whose columns are classified.

spacr.qt.widgets.pca_model.component_index(name: str) → int | None[source]

'PC1' -> 0, and None for anything that is not a PC column.

Parameters:

name – a column name such as "PC1"; stripped and matched case-insensitively.

spacr.qt.widgets.pca_model.component_name(index: int) → str[source]

0 -> 'PC1'. The column name a score lands in, in one place.

Parameters:

index – the 0-based component index; converted to int.

spacr.qt.widgets.pca_model.pca(frame: pandas.DataFrame, spec: PCASpec | None = None) → PCAResult[source]

Decompose frame’s features under spec.

The whole policy is in the module docstring; the short version is that features are standardised unless told otherwise, NaN is never imputed unless told to, constant columns are dropped by name, collinear ones are kept and reported, and no more components are returned than the data has independent directions.

Parameters:
  • frame – the measurement table; the features are read from its columns and coerced to numbers, with ±inf treated as missing.

  • spec – what to decompose and how; None uses a default PCASpec (every candidate feature, z-scored).

Raises:

PCAError – whenever the answer would be meaningless — with the reason and the way out in the message.