spacr.qt.widgets.pca_model¶
PCA over a measurement table — the statistics, with the decisions written down.
spaCR measures hundreds of features per object and, until this module, had no principal components anywhere. The reason it is worth having is not dimensionality reduction for its own sake: it is the one view that answers “these two populations separate — on what?” in a single picture, because the loadings say which measurements carry the separation.
PCA is also easy to get subtly wrong on exactly this kind of table, in ways that produce a confident-looking plot of nothing. Three decisions do all the damage, so all three are made here, in the open, and reported back to the user with every result.
1. Centring and scaling¶
Always centred. Scaled to unit variance by default
(SCALE_ZSCORE), with SCALE_NONE available and never implicit.
PCA maximises variance, and variance has units. In a spaCR object table
cell_area is px² and runs to 10^4 with a variance near 10^6;
eccentricity is dimensionless and lives in [0, 1];
channel_1_mean_intensity is counts. An unscaled (covariance) PCA of those
three returns “PC1 = cell_area” every time — not because size is the
interesting axis but because px² is a big number. Worse, the answer is not
even stable: recording the same areas in µm² instead of px² would rotate every
component.
Standardising each feature to mean 0 and sample SD 1 (a correlation PCA) makes the result invariant to any per-feature linear rescaling, which is the only defensible default when the columns have no common unit.
It is not free, and the cost is stated rather than hidden: equal weight means
forty near-duplicate Zernike moments outvote one carefully chosen intensity
ratio, and a feature that is pure noise gets the same unit of variance as a
feature that is pure signal. SCALE_NONE is there for the case where
the user has already put the features in a common unit and wants the big
ones to dominate.
2. NaN¶
Never imputed by default, and never zero-filled at all.
A NaN in a spaCR measurement table is usually structural, not missing at
random — the same reasoning spacr.anndata_export sets out at length. A
pathogen_* column is NaN for a cell with no pathogen in it; that is a fact
about the cell, not a gap in the record. Mean-imputing it invents a pathogen of
average size and moves an uninfected cell into the middle of the infected
distribution, invisibly.
Dropping rows instead is not neutral either: complete-case analysis over a
feature set containing one pathogen_* column silently deletes every
uninfected cell, so the result is the PCA of the infected cells presented as
the PCA of the cells.
The three failure modes are different in kind, which is what the policies are built around:
dropping a feature changes the space — a smaller question, answered exactly;
dropping a row changes the population — the same question, answered about a self-selected subset;
imputing changes the data — the original question, answered about numbers nobody measured.
NAN_AUTO, the default, picks between the first two by how the
missingness is shaped, because in this table the two shapes are far apart: a
structural NaN affects tens of percent of the rows (every uninfected cell), a
sporadic one affects a handful (one object where mahotas returned nothing). So
a feature missing in more than DEFAULT_STRUCTURAL_MISSING of the rows
is treated as structural and dropped by name; rows still carrying a NaN among
the survivors are dropped as sporadic and counted. Nothing is imputed, and both
actions appear in PCAResult.notes and in
PCAResult.report().
NAN_COMPLETE, NAN_DROP_FEATURES and NAN_MEAN are the
explicit versions of the three failure modes above, for a user who knows which
one they want. NAN_MEAN is offered and never chosen for them.
Infinities are missing under every policy. Ratio features produce ±inf from
a zero denominator, an inf survives dropna, and one of them destroys the
scaling and therefore every component. They become NaN before the policy runs
and are counted separately in PCAResult.n_infinite.
3. Constant and collinear columns¶
A constant column has SD 0, so standardising it is 0/0. Left in, it
contributes a NaN that spreads through the whole decomposition; “fixed” by
substituting a 1, it becomes a zero-variance direction that the solver is free
to return as a component with an arbitrary loading pattern. It is dropped
before the decomposition and named in PCAResult.dropped_features —
a column with one value cannot explain variation in anything.
Perfectly collinear columns are a different problem and get a different
answer: they are kept, because the feature list is the user’s statement of
what they care about and picking which of area and area_um2 to discard
is not this module’s call. What is refused is inventing components out of the
redundancy. The numerical rank is computed from the singular values and no more
than rank components are ever returned, so a table of 12 features spanning
9 real directions yields 9 components, not 12 — the last three would have ~0
explained variance and a direction chosen by floating-point noise. Groups of
features that are perfectly correlated are also detected and reported, since a
duplicated feature silently doubles its own weight in a correlation PCA.
Honest explained variance¶
PCAResult.explained_variance_ratio is over the total variance of the
analysed matrix, so it sums to 1 across all rank components rather than
across the ones that were kept. Alongside it, every result carries how much of
the original table it is actually about (PCAResult.n_rows_in versus
len(result), PCAResult.n_features_in versus its features): “PC1
explains 71%” means something different when the matrix is three features and
8% of the objects.
And the single most common way a PCA of morphology features misleads anyone is
that PC1 is just size. When every feature loads with the same sign, PC1 is
a general-magnitude axis — big objects at one end, small at the other — which
is usually a fact about segmentation rather than biology.
PCAResult.is_size_like() detects it from the loading signs and
PCAResult.headline() says it in words, unprompted.
No Qt in here¶
Pure numpy and pandas, like spacr.qt.widgets.graph_spec and
spacr.selection: usable from a notebook, testable without a display, and
free for the later screens (the feature explorer, the gate editor) to reuse
without inheriting a widget.
Exceptions¶
A PCA that cannot mean anything, with the reason in the message. |
Classes¶
Functions¶
|
The columns worth offering as PCA features, sorted. |
|
|
|
|
|
Decompose |
Module Contents¶
- exception spacr.qt.widgets.pca_model.PCAError[source]¶
Bases:
ValueErrorA PCA that cannot mean anything, with the reason in the message.
Raised rather than returned as an empty result: every one of these is a sentence the user can act on (“every row has a NaN in
pathogen_area”), and a caller that swallowed it would draw an empty scatter with no explanation. Screens catch it and show the message.Initialize self. See help(type(self)) for accurate signature.
- class spacr.qt.widgets.pca_model.PCAResult[source]¶
A decomposition, plus everything needed to read it honestly.
- Parameters:
features – the columns actually decomposed, in matrix order. Not the ones asked for — see
dropped_features.rows – positional indices into the frame
pca()was given, naming the objects that survived the NaN policy. Positional because a measurement frame carries a duplicated or reset index often enough that positions are the only safe currency, the same reasonFacetPaneluses them.scores –
(len(rows), k)— each object’s coordinate on each component, in the units of the analysed (post-scaling) matrix.loadings –
(len(features), k), unit-norm columns. The direction of each component in feature space.correlations –
(len(features), k)Pearson r between each feature and each component score. This is what the biplot arrows are: unlike a unit-norm loading, an r is a number with a meaning on its own, and for a standardised PCA the squared row sums are each feature’s communality.explained_variance – per component, in the analysed matrix’s units.
explained_variance_ratio – over the total variance of the analysed matrix — so it sums to 1 across all
rankcomponents, not across the ones returned.total_variance – the analysed matrix’s total variance, summed over all components (not only the ones returned), in its own units.
centre – per feature, the mean subtracted before the decomposition, aligned with
features.scale – per feature, the divisor applied after centring: the standard deviation (
ddof=1) under z-score scaling, else 1.rank – the numerical rank. No more components than this exist.
n_rows_in – rows in the frame
pca()was given, before the NaN policy.n_features_in – requested features that exist in the frame, before the NaN policy and constant columns removed any.
scaling – the scaling mode used, one of
SCALE_MODES.nan_policy – the missing-value policy used, one of
NAN_POLICIES.dropped_features –
{column: why}, for everything that did not reach the decomposition.dropped_rows – objects removed by the NaN policy.
n_infinite – cells that were ±inf and were treated as missing.
collinear_groups – features that are perfectly correlated with each other. Kept in the analysis, reported because each such group weighs as many times as it has members.
notes – plain-language notes on what the analysis did, e.g. rank deficiency.
- dominant(k: int = 0) Tuple[str, float][source]¶
(feature, share)— the feature with the largest|loading|on componentk, and how much of the component’s squared loading it carries.A share near 1 means the component is that feature under another name, which is worth saying before anyone reads biology into it.
- is_degenerate(k: int = 0) bool[source]¶
Whether component
kholds so little variance that its direction is not identified.Near-collinear features leave a numerically full-rank matrix whose last directions are floating-point residue: the solver returns a direction because it must, not because the data has one. A component under
DEGENERATE_RATIOof the total variance is reported with that said, so nobody reads a loading pattern out of rounding error.
- is_size_like(k: int = 0) bool[source]¶
Whether component
kis a general-magnitude axis.Every feature loading with the same sign is the classic “size” factor of morphometrics: one end big objects, the other end small ones. It is real, it is usually PC1, and it is usually a statement about segmentation and focus rather than about biology — so it is detected and said out loud rather than left for the reader to notice.
Two guards keep it from crying wolf. It needs three features, because with two “same sign” is a coin flip. And it needs the loading spread over more than one of them: a component that is a single feature trivially agrees with itself about sign, and calling that a size axis would be a statement about arithmetic rather than about the objects.
- loadings_frame() pandas.DataFrame[source]¶
One row per feature: its loading and its correlation on each component. What “export the loadings” writes.
- plane_features(kx: int = 0, ky: int = 1, count: int = 8) Tuple[int, ...][source]¶
Indices of the features best represented in the
(kx, ky)plane.Ranked by
r_x² + r_y²— how much of the feature is visible in this plane at all. Drawing the other four hundred arrows would say nothing except that the figure is full; drawing the longest ones is the standard choice because a short arrow means “this feature points somewhere you are not looking”, not “this feature does not matter”.
- scores_frame(source: pandas.DataFrame, *, components: int | None = None) pandas.DataFrame[source]¶
source’s surviving rows withPC1…PCkcolumns added.Every original column is kept, which is the whole point: it is what lets the scores plot colour by
gene, facet byplateIDand publish a real object-key selection when a cluster is brushed. The Graph Builder then draws it with no knowledge that it is a PCA.- Parameters:
source – the same frame
pca()was given; its surviving rows are taken by position.components – how many
PCcolumns to add;Noneadds all, and other values are clamped between 1 andn_components.
- sign_agreement(k: int = 0) float[source]¶
How one-sided component
k’s loadings are, in[0.5, 1].The larger share of squared loading sitting on a single sign. 1.0 means every feature moves the same way along this axis.
- top_features(k: int = 0, count: int = 5) Tuple[Tuple[str, float], ...][source]¶
The
countfeatures loading hardest on componentk, as(feature, loading)with the sign kept, strongest first.
- variance_frame() pandas.DataFrame[source]¶
One row per component — the scree plot’s data, and its CSV.
- property component_names: Tuple[str, ...][source]¶
The components’ display names, in order.
- Returns:
one name per component.
- property cumulative_ratio: numpy.ndarray[source]¶
Running total of
explained_variance_ratio.
- property n_features: int[source]¶
How many columns went into the decomposition.
The columns ACTUALLY used, which can be fewer than the spec asked for: a constant or all-null column contributes no variance and is dropped before fitting.
- Returns:
the feature count.
Fraction of the given objects the analysis is about.
- class spacr.qt.widgets.pca_model.PCASpec[source]¶
What to decompose and how. Frozen and JSON round-tripping, like
GraphSpec, so a saved analysis is something a settings file or a report can carry.- Parameters:
features – the columns to decompose. Empty means
candidate_features()of whatever frame it is run against — which is a default, not a promise: the result records the features it actually used.n_components – how many to compute. Capped at the numerical rank, always: components past the rank have no direction.
scaling –
SCALE_ZSCORE(default) orSCALE_NONE.nan_policy – one of
NAN_POLICIES.structural_missing – the
NAN_AUTOthreshold.
- Raises:
PCAError – on an unknown scaling or policy, or a non-positive component count — at the point the spec is built, not at render time.
- __post_init__() None[source]¶
De-duplicate the features and validate the decomposition settings.
- Raises:
PCAError – if
scalingornan_policyis not one this module offers, ifn_componentsis below 1, or ifstructural_missingis not a fraction of rows in[0, 1]. The component count is capped at the module’s maximum rather than refused: asking for more components than that is a request for noise, not an error.
- classmethod from_dict(payload: Mapping[str, Any]) PCASpec[source]¶
Rebuild from
to_dict(); unknown keys ignored, missing keys defaulted, so an analysis written by another build still opens.- Parameters:
payload – a mapping as
to_dict()writes it;features,n_components,scaling,nan_policyandstructural_missingare read and any other key is ignored.
- classmethod from_json(text: str) PCASpec[source]¶
Rebuild a spec from JSON text.
- Parameters:
text – the JSON text.
- Returns:
the rebuilt spec.
- to_json() str[source]¶
This spec as JSON text, keys sorted so the file is diffable.
- Returns:
the JSON text.
- with_components(n: int) PCASpec[source]¶
A copy keeping a different number of components.
- Parameters:
n – how many to keep.
- Returns:
the new spec.
- with_features(features: Sequence[str]) PCASpec[source]¶
A copy decomposing a different set of columns.
A COPY: a spec is a value, so the one a view is showing is never edited underneath it.
- Parameters:
features – the columns to decompose.
- Returns:
the new spec.
- with_nan_policy(policy: str) PCASpec[source]¶
A copy handling missing values differently.
- Parameters:
policy – the policy’s name.
- Returns:
the new spec.
- with_scaling(scaling: str) PCASpec[source]¶
A copy using a different scaling.
SCALING IS NOT COSMETIC HERE. PCA maximises variance, so unscaled columns in different units let the largest unit dominate every component – which is why this is a spec field and not a display option.
- Parameters:
scaling – the scaling’s name.
- Returns:
the new spec.
- spacr.qt.widgets.pca_model.candidate_features(frame: pandas.DataFrame) Tuple[str, ...][source]¶
The columns worth offering as PCA features, sorted.
Continuous by
spacr.qt.widgets.graph_spec.column_kinds()— which is a re-reading of the Local Data Filter’s classifier, the one column classifier in this codebase. That rule already excludes object keys (they identify rather than describe) and small-cardinality numeric codes likecell_countand a class label, which are counts and labels rather than measured quantities and have no business setting the scale of a component.A user who disagrees about a particular column puts it in
PCASpec.featuresexplicitly; nothing here refuses it.- Parameters:
frame – the measurement table whose columns are classified.
- spacr.qt.widgets.pca_model.component_index(name: str) int | None[source]¶
'PC1' -> 0, andNonefor anything that is not a PC column.- Parameters:
name – a column name such as
"PC1"; stripped and matched case-insensitively.
- spacr.qt.widgets.pca_model.component_name(index: int) str[source]¶
0 -> 'PC1'. The column name a score lands in, in one place.- Parameters:
index – the 0-based component index; converted to
int.
- spacr.qt.widgets.pca_model.pca(frame: pandas.DataFrame, spec: PCASpec | None = None) PCAResult[source]¶
Decompose
frame’s features underspec.The whole policy is in the module docstring; the short version is that features are standardised unless told otherwise, NaN is never imputed unless told to, constant columns are dropped by name, collinear ones are kept and reported, and no more components are returned than the data has independent directions.
- Parameters:
frame – the measurement table; the features are read from its columns and coerced to numbers, with ±inf treated as missing.
spec – what to decompose and how;
Noneuses a defaultPCASpec(every candidate feature, z-scored).
- Raises:
PCAError – whenever the answer would be meaningless — with the reason and the way out in the message.