spacr.regression_annotation¶
Select cells for annotation and evaluate the resulting classifier.
This module implements annotation strategies used by the Cells tab. A strategy selects cells, optionally fits a classifier, applies the fitted model to cells excluded from training, and evaluates it on a preselected holdout.
Evaluation groups observations at AnnotationRequest.group_by so cells
from the same well or other independence unit cannot be divided between
training and test sets. Holdout groups are selected before a strategy chooses
cells and are unavailable to that strategy.
Score-based selection can leak phenotype information when the score or its
inputs are also model features. score_input_columns() identifies those
columns, and LeakageReport compares fits with and without them.
Declared strategies without an implementation raise
StrategyNotImplemented.
Exceptions¶
Raised when an annotation strategy cannot use the supplied data. |
|
Raised when the selected labels cannot support model fitting. |
|
Raised when a declared annotation strategy is not implemented. |
Classes¶
Configure cell selection, model fitting, and holdout evaluation. |
|
Store one strategy's selection, predictions, and evaluation. |
|
Summarize model performance on the reserved holdout. |
|
Compare fits that include and exclude score-derived inputs. |
|
Resolved data shared by all strategies in one run. |
|
Describe one cell-selection strategy. |
Functions¶
|
The columns that COULD be features, judged on names and dtypes alone. |
|
Return measurement columns suitable for model fitting. |
|
Partition feature columns into intensity and shape views. |
|
Return the strategy keys currently supported by |
|
Return one user-facing description for each strategy. |
|
Why |
|
Resolve columns, group identifiers, labels, and the holdout. |
|
Format a group identifier for display, such as |
|
Run one strategy end to end. |
|
Identify known or likely inputs to the classification score. |
|
How many cells carry a finite value in the score column. |
|
Return the menu entry for |
|
Return all strategy keys in menu order. |
|
How many cells carry an annotation in |
|
Return a mask selecting rows whose group matches |
|
Return whether XGBoost can be imported. |
Module Contents¶
- exception spacr.regression_annotation.AnnotationStrategyError[source]¶
Bases:
ValueErrorRaised when an annotation strategy cannot use the supplied data.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.regression_annotation.NotEnoughLabels[source]¶
Bases:
AnnotationStrategyErrorRaised when the selected labels cannot support model fitting.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.regression_annotation.StrategyNotImplemented[source]¶
Bases:
AnnotationStrategyErrorRaised when a declared annotation strategy is not implemented.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.regression_annotation.AnnotationRequest[source]¶
Configure cell selection, model fitting, and holdout evaluation.
- Parameters:
frame – one row per cell, with the score column, the acquisition metadata the split level needs, and whatever measurements are to be fitted on.
score_column – the per-object classification score.
feature_columns – the columns to fit on.
Nonemeans infer them – seefeature_columns().score_inputs – Columns used to calculate the score.
Noneinfers them withscore_input_columns(); an explicit list provides the more reliable leakage specification.group_by – Independence level that a split may not cross;
'well','field','plate'or'cell'.wells – Selected guide wells. A well matches when every supplied name token is part of that well’s identity, so
'r1_c1'matches the wellplate1/r1/c1. Empty means the complete screen is eligible.label_column – Column containing existing manual annotations. If it contains at least two classes, those annotations define the holdout labels; otherwise labels are derived from the score.
positive_control_wells – Wells whose cells are known positives.
negative_control_wells – Wells whose cells are known negatives.
n_positive – Size of the positive set, matched contrast draw, and annotation queues.
holdout_fraction – Fraction of independence groups reserved before strategy-specific selection.
seed – Random seed controlling selection and splitting.
leakage – One of
LEAKAGE_MODES.model –
'auto','xgboost'or'hist_gradient_boosting'.measure – the uncertainty measure, from
spacr.active_learning.UNCERTAINTY_MEASURES.n_clusters – Cluster count for diversity sampling. Zero uses the annotation budget as the cluster count.
n_bins – Number of score strata.
confidence – Minimum self-training probability for accepting a model prediction as a pseudo-label.
rounds – Maximum number of self-training rounds.
neighbours – Maximum propagated neighbours per seed.
distance_quantile – Propagation-radius quantile of observed nearest-neighbour distances.
distance_cut – Explicit propagation radius in standardized feature space. Overrides
distance_quantilewhen provided.correlation_cut – Absolute rank-correlation threshold for treating a feature as a score input.
- validated() AnnotationRequest[source]¶
Validate common parameters and return this request.
- class spacr.regression_annotation.AnnotationResult[source]¶
Store one strategy’s selection, predictions, and evaluation.
- Variables:
strategy – the key that produced it.
title – that strategy’s name on screen.
selection – one row per chosen cell, indexed as the object table is, with
annotation_rolenaming why it was chosen.holdout – the hold-out rows, same shape.
predictions – the fitted model applied to every cell it was not fitted on – the rest of the screen and the rest of the chosen wells – or None when the strategy fitted nothing.
fit – the headline hold-out report, or None.
leakage – the same selection with and without the score’s own inputs, or None when nothing was fitted.
notes – Decisions and limitations recorded during the run.
counts – Selection and evaluation counts.
- class spacr.regression_annotation.FitReport[source]¶
Summarize model performance on the reserved holdout.
- Variables:
model – which estimator produced it.
features – the columns it was fitted on.
n_train – cells fitted on.
n_test – hold-out cells scored on.
accuracy – hold-out accuracy.
balanced_accuracy – Holdout accuracy averaged across classes.
roc_auc – hold-out area under the ROC curve, or None when the hold-out holds one class.
positive_share_train – share of the training rows labelled positive.
positive_share_test – share of the hold-out labelled positive.
label_source – where the labels came from, in words.
split_summary – the splitter’s own account of the hold-out.
- class spacr.regression_annotation.LeakageReport[source]¶
Compare fits that include and exclude score-derived inputs.
- Variables:
mode – the leakage mode the run was asked for.
dropped – Columns removed from the leakage-controlled fit.
with_score_inputs – the fit that keeps them, or None.
without_score_inputs – the fit that removes them, or None.
survival – Fraction of the inclusive fit’s lift over chance retained after score inputs are removed.
Noneif either fit is absent or the inclusive fit is at chance.
- class spacr.regression_annotation.Prepared[source]¶
Resolved data shared by all strategies in one run.
- Variables:
frame – the object table, unchanged.
groups – one independence-group id per row.
level – the level those ids are at.
features – the columns a model may be fitted on.
score_inputs – the columns the score is a function of.
honest_features –
featureswithscore_inputsremoved.labels – Binary reference label per row. Values are valid only where
knownis true.known – Rows carrying a usable reference label.
label_source – Human-readable source of the reference labels.
threshold – Score threshold defining the positive class, or
NaNwhen manual annotations define the labels.holdout – Positional indices of all rows in reserved holdout groups.
selectable – Positional indices available to strategies.
chosen – Available indices within selected guide wells.
split – Split provenance returned by the grouped splitter.
notes – Decisions and caveats generated while preparing the data.
- holdout_labels() numpy.ndarray[source]¶
Return reference labels for holdout rows.
- labelled() numpy.ndarray[source]¶
Return selectable row indices with usable reference labels.
Return the positive-label fraction among
positions.- Parameters:
positions – integer row positions whose prepared labels are read.
- class spacr.regression_annotation.Strategy[source]¶
Describe one cell-selection strategy.
- Parameters:
key – the stored value; stable across releases.
title – what the entry is called on screen.
purpose – what it is for, in one sentence.
cost – Principal limitation or trade-off, in one sentence.
implemented – Whether
run()can execute the strategy.needs – Required data components:
'score','features','labels','controls'.
- spacr.regression_annotation.candidate_feature_columns(frame: pandas.DataFrame, score_column: str = DEFAULT_SCORE_COLUMN) Tuple[str, ...][source]¶
The columns that COULD be features, judged on names and dtypes alone.
feature_columns()is the authority and it reads every value: a numeric column that does not vary is not a feature, and only the values say so. This is the cheap half of that question – one pass overframe.dtypesrather than over the rows – for a caller that has to answer “can this strategy run at all” every time a chooser changes, on a table that may hold half a million objects.It is therefore the OPTIMISTIC answer: a table this accepts can still be refused by
feature_columns(), and the run is what refuses it. A table this rejects has no measurement column under any reading.- Parameters:
frame – the object table.
score_column – the score, which is never a feature.
- Returns:
the candidate columns, in table order.
- spacr.regression_annotation.feature_columns(frame: pandas.DataFrame, score_column: str = DEFAULT_SCORE_COLUMN, explicit: Sequence[str] | None = None) Tuple[str, ...][source]¶
Return measurement columns suitable for model fitting.
Identifier columns, classifier outputs, and invariant columns are excluded to reduce plate-layout and score leakage.
- Parameters:
frame – the object table.
score_column – the score, which is never a feature.
explicit – Optional explicit feature list. Missing columns raise an error rather than being removed silently.
- Returns:
the columns, in table order.
- Raises:
AnnotationStrategyError – an explicit column is absent, or nothing at all is left to fit on.
- spacr.regression_annotation.feature_views(columns: Sequence[str]) Dict[str, Tuple[str, ...]][source]¶
Partition feature columns into intensity and shape views.
Feature families are assigned with
spacr.column_groups.classify(). Unclassified columns are distributed alternately so both views remain usable; the resulting column counts are included in strategy reports.- Parameters:
columns – the feature columns.
- Returns:
{'intensity': (...), 'shape': (...)}.
- spacr.regression_annotation.implemented_keys() Tuple[str, ...][source]¶
Return the strategy keys currently supported by
run().
Return one user-facing description for each strategy.
- spacr.regression_annotation.missing_requirement(key: Any, frame: pandas.DataFrame | None, score_column: str = DEFAULT_SCORE_COLUMN, *, label_column: str = '', positive_control_wells: Sequence[str] = (), negative_control_wells: Sequence[str] = ()) str[source]¶
Why
keycannot run on this table, or""when it can.ASKED BEFORE THE RUN, NOT AFTER IT. Every refusal in this module is raised while a strategy is executing, which is the right place for it and the wrong time for a user: choosing “Diversity sampling over clusters” on a coefficient table with no measurement columns joined to it should say so on the control, not a run later. This is the cheap pre-flight that lets a chooser grey itself with the reason – it reads dtypes and one column, never the whole matrix, so it can be asked on every change of the menu.
IT IS THE OPTIMISTIC HALF. An empty answer means nothing this can see is missing;
prepare()still reads the values and can still refuse – a score whose cells are all in one well, a feature column that turns out not to vary. What it will never do is stay silent about a table that has no score, no annotations, no measurements, or no control wells named for the strategy that needs them.- Parameters:
key – a strategy key, or a
Strategy.frame – the object rows on screen, or None.
score_column – the per-object classification score.
label_column – a column of human annotations, when one is named.
positive_control_wells – the positive control wells named.
negative_control_wells – the negative control wells named.
- Returns:
one sentence naming what is missing and what to do about it, or “” when the strategy can be run.
- Raises:
AnnotationStrategyError –
keyis not on the menu.
- spacr.regression_annotation.prepare(request: AnnotationRequest, entry: Strategy | None = None) Prepared[source]¶
Resolve columns, group identifiers, labels, and the holdout.
Complete independence groups are reserved before strategy-specific selection. Strategies cannot select rows from those groups.
Strategies that fit a model require measurement columns. Score-stratified and random sampling do not fit a model and therefore remain available for result tables that do not contain joined measurements.
- Parameters:
request – what to run.
entry – the menu entry about to be run, when it is known. It is read only for what the strategy needs;
Nonerequires everything, which is the strict answer a caller with no entry to hand should get.
- Returns:
the shared setup.
- Raises:
AnnotationStrategyError – the table cannot support a leakage-safe split, or has no features, labels or score.
- spacr.regression_annotation.readable_group(value: Any) str[source]¶
Format a group identifier for display, such as
plate1/r1/c1.- Parameters:
value – a group id from the splitter.
- Returns:
the same identity with its parts separated visibly.
- spacr.regression_annotation.run(key: Any, request: AnnotationRequest, prepared: Prepared | None = None) AnnotationResult[source]¶
Run one strategy end to end.
- Parameters:
- Returns:
what the strategy chose, fitted and measured.
- Raises:
StrategyNotImplemented – The selected entry has no implementation.
AnnotationStrategyError – the data cannot support the strategy.
- spacr.regression_annotation.score_input_columns(frame: pandas.DataFrame, score_column: str = DEFAULT_SCORE_COLUMN, features: Sequence[str] | None = None, explicit: Sequence[str] | None = None, correlation_cut: float = 0.5) Tuple[str, ...][source]¶
Identify known or likely inputs to the classification score.
Without an explicit input list, two rules are applied:
any column whose name marks it as a classifier output – the score itself, probabilities, logits, other prediction columns;
any feature whose absolute Spearman correlation with the score reaches
correlation_cut.
Pass the classifier’s feature list as
explicitwhen available to replace this correlation-based approximation.- Parameters:
frame – the object table.
score_column – the score.
features – the candidate feature columns.
explicit – the classifier’s own inputs, when they are known.
correlation_cut – the absolute rank correlation at which a feature counts as one of the score’s inputs.
- Returns:
the columns, in table order, including the score itself.
- spacr.regression_annotation.scored_cells(frame: pandas.DataFrame, score_column: str = DEFAULT_SCORE_COLUMN) int[source]¶
How many cells carry a finite value in the score column.
- spacr.regression_annotation.strategy(key: Any) Strategy[source]¶
Return the menu entry for
key.- Parameters:
key – a strategy key, or a
Strategy.- Raises:
AnnotationStrategyError – no entry has that key; the message lists the ones that do.
- spacr.regression_annotation.strategy_keys() Tuple[str, ...][source]¶
Return all strategy keys in menu order.
- spacr.regression_annotation.usable_annotations(frame: pandas.DataFrame, label_column: str) int[source]¶
How many cells carry an annotation in
label_column.Zero when the column is absent, empty, or holds one class only – which is the same answer
_reference_labels()gives, so a chooser and a run cannot disagree about whether there are labels to fit on.- Parameters:
frame – the object table.
label_column – the annotation column, or “”.
- Returns:
the number of annotated cells, or 0 when they are unusable.
- spacr.regression_annotation.wells_selected(groups: Sequence[Any], wanted: Sequence[str]) numpy.ndarray[source]¶
Return a mask selecting rows whose group matches
wanted.A well matches when every token of the name given is one of the group’s own tokens, so
'r1_c1'matchesplate1/r1/c1and'plate2_r1_c1'does not. Matching on tokens rather than on the complete string allows plate-map well names to match internal group identifiers without exposing their separator format.- Parameters:
groups – one group id per row.
wanted – the well names chosen.
- Returns:
the mask; all-True when nothing was named.