spacr.regression_annotation

Select cells for annotation and evaluate the resulting classifier.

This module implements annotation strategies used by the Cells tab. A strategy selects cells, optionally fits a classifier, applies the fitted model to cells excluded from training, and evaluates it on a preselected holdout.

Evaluation groups observations at AnnotationRequest.group_by so cells from the same well or other independence unit cannot be divided between training and test sets. Holdout groups are selected before a strategy chooses cells and are unavailable to that strategy.

Score-based selection can leak phenotype information when the score or its inputs are also model features. score_input_columns() identifies those columns, and LeakageReport compares fits with and without them. Declared strategies without an implementation raise StrategyNotImplemented.

Exceptions

AnnotationStrategyError

Raised when an annotation strategy cannot use the supplied data.

NotEnoughLabels

Raised when the selected labels cannot support model fitting.

StrategyNotImplemented

Raised when a declared annotation strategy is not implemented.

Classes

AnnotationRequest

Configure cell selection, model fitting, and holdout evaluation.

AnnotationResult

Store one strategy's selection, predictions, and evaluation.

FitReport

Summarize model performance on the reserved holdout.

LeakageReport

Compare fits that include and exclude score-derived inputs.

Prepared

Resolved data shared by all strategies in one run.

Strategy

Describe one cell-selection strategy.

Functions

candidate_feature_columns(→ Tuple[str, ...])

The columns that COULD be features, judged on names and dtypes alone.

feature_columns(→ Tuple[str, ...])

Return measurement columns suitable for model fitting.

feature_views(→ Dict[str, Tuple[str, ...]])

Partition feature columns into intensity and shape views.

implemented_keys(→ Tuple[str, ...])

Return the strategy keys currently supported by run().

menu(→ Tuple[str, ...])

Return one user-facing description for each strategy.

missing_requirement(, negative_control_wells)

Why key cannot run on this table, or "" when it can.

prepare(→ Prepared)

Resolve columns, group identifiers, labels, and the holdout.

readable_group(→ str)

Format a group identifier for display, such as plate1/r1/c1.

run(→ AnnotationResult)

Run one strategy end to end.

score_input_columns(→ Tuple[str, ...])

Identify known or likely inputs to the classification score.

scored_cells(→ int)

How many cells carry a finite value in the score column.

strategy(→ Strategy)

Return the menu entry for key.

strategy_keys(→ Tuple[str, ...])

Return all strategy keys in menu order.

usable_annotations(→ int)

How many cells carry an annotation in label_column.

wells_selected(→ numpy.ndarray)

Return a mask selecting rows whose group matches wanted.

xgboost_available(→ bool)

Return whether XGBoost can be imported.

Module Contents

exception spacr.regression_annotation.AnnotationStrategyError[source]

Bases: ValueError

Raised when an annotation strategy cannot use the supplied data.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.regression_annotation.NotEnoughLabels[source]

Bases: AnnotationStrategyError

Raised when the selected labels cannot support model fitting.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.regression_annotation.StrategyNotImplemented[source]

Bases: AnnotationStrategyError

Raised when a declared annotation strategy is not implemented.

Initialize self. See help(type(self)) for accurate signature.

class spacr.regression_annotation.AnnotationRequest[source]

Configure cell selection, model fitting, and holdout evaluation.

Parameters:
  • frame – one row per cell, with the score column, the acquisition metadata the split level needs, and whatever measurements are to be fitted on.

  • score_column – the per-object classification score.

  • feature_columns – the columns to fit on. None means infer them – see feature_columns().

  • score_inputs – Columns used to calculate the score. None infers them with score_input_columns(); an explicit list provides the more reliable leakage specification.

  • group_by – Independence level that a split may not cross; 'well', 'field', 'plate' or 'cell'.

  • wells – Selected guide wells. A well matches when every supplied name token is part of that well’s identity, so 'r1_c1' matches the well plate1/r1/c1. Empty means the complete screen is eligible.

  • label_column – Column containing existing manual annotations. If it contains at least two classes, those annotations define the holdout labels; otherwise labels are derived from the score.

  • positive_control_wells – Wells whose cells are known positives.

  • negative_control_wells – Wells whose cells are known negatives.

  • n_positive – Size of the positive set, matched contrast draw, and annotation queues.

  • holdout_fraction – Fraction of independence groups reserved before strategy-specific selection.

  • seed – Random seed controlling selection and splitting.

  • leakage – One of LEAKAGE_MODES.

  • model – 'auto', 'xgboost' or 'hist_gradient_boosting'.

  • measure – the uncertainty measure, from spacr.active_learning.UNCERTAINTY_MEASURES.

  • n_clusters – Cluster count for diversity sampling. Zero uses the annotation budget as the cluster count.

  • n_bins – Number of score strata.

  • confidence – Minimum self-training probability for accepting a model prediction as a pseudo-label.

  • rounds – Maximum number of self-training rounds.

  • neighbours – Maximum propagated neighbours per seed.

  • distance_quantile – Propagation-radius quantile of observed nearest-neighbour distances.

  • distance_cut – Explicit propagation radius in standardized feature space. Overrides distance_quantile when provided.

  • correlation_cut – Absolute rank-correlation threshold for treating a feature as a score input.

validated() → AnnotationRequest[source]

Validate common parameters and return this request.

class spacr.regression_annotation.AnnotationResult[source]

Store one strategy’s selection, predictions, and evaluation.

Variables:
  • strategy – the key that produced it.

  • title – that strategy’s name on screen.

  • selection – one row per chosen cell, indexed as the object table is, with annotation_role naming why it was chosen.

  • holdout – the hold-out rows, same shape.

  • predictions – the fitted model applied to every cell it was not fitted on – the rest of the screen and the rest of the chosen wells – or None when the strategy fitted nothing.

  • fit – the headline hold-out report, or None.

  • leakage – the same selection with and without the score’s own inputs, or None when nothing was fitted.

  • notes – Decisions and limitations recorded during the run.

  • counts – Selection and evaluation counts.

role_counts() → Dict[str, int][source]

Return cell counts by annotation role, including the holdout.

summary() → str[source]

Return a report of selections, fitting, and holdout evaluation.

write(folder: str) → Dict[str, str][source]

Write selection, holdout, prediction, and report files.

Parameters:

folder – the directory to write into; created when absent.

Returns:

{what: path} for every file written.

class spacr.regression_annotation.FitReport[source]

Summarize model performance on the reserved holdout.

Variables:
  • model – which estimator produced it.

  • features – the columns it was fitted on.

  • n_train – cells fitted on.

  • n_test – hold-out cells scored on.

  • accuracy – hold-out accuracy.

  • balanced_accuracy – Holdout accuracy averaged across classes.

  • roc_auc – hold-out area under the ROC curve, or None when the hold-out holds one class.

  • positive_share_train – share of the training rows labelled positive.

  • positive_share_test – share of the hold-out labelled positive.

  • label_source – where the labels came from, in words.

  • split_summary – the splitter’s own account of the hold-out.

summary() → str[source]

Return a one-line summary of holdout performance.

property lift: float[source]

Return performance above the 0.5 chance level.

Balanced accuracy has chance at 0.5 whatever the class balance, and so does the area under the ROC curve, so both give a lift that can be compared between two fits on different columns.

class spacr.regression_annotation.LeakageReport[source]

Compare fits that include and exclude score-derived inputs.

Variables:
  • mode – the leakage mode the run was asked for.

  • dropped – Columns removed from the leakage-controlled fit.

  • with_score_inputs – the fit that keeps them, or None.

  • without_score_inputs – the fit that removes them, or None.

  • survival – Fraction of the inclusive fit’s lift over chance retained after score inputs are removed. None if either fit is absent or the inclusive fit is at chance.

summary() → str[source]

Return both fit summaries and their leakage comparison.

class spacr.regression_annotation.Prepared[source]

Resolved data shared by all strategies in one run.

Variables:
  • frame – the object table, unchanged.

  • groups – one independence-group id per row.

  • level – the level those ids are at.

  • features – the columns a model may be fitted on.

  • score_inputs – the columns the score is a function of.

  • honest_features – features with score_inputs removed.

  • labels – Binary reference label per row. Values are valid only where known is true.

  • known – Rows carrying a usable reference label.

  • label_source – Human-readable source of the reference labels.

  • threshold – Score threshold defining the positive class, or NaN when manual annotations define the labels.

  • holdout – Positional indices of all rows in reserved holdout groups.

  • selectable – Positional indices available to strategies.

  • chosen – Available indices within selected guide wells.

  • split – Split provenance returned by the grouped splitter.

  • notes – Decisions and caveats generated while preparing the data.

holdout_labels() → numpy.ndarray[source]

Return reference labels for holdout rows.

labelled() → numpy.ndarray[source]

Return selectable row indices with usable reference labels.

positive_share(positions: Sequence[int]) → float[source]

Return the positive-label fraction among positions.

Parameters:

positions – integer row positions whose prepared labels are read.

property annotated: bool[source]

Return whether reference labels came from manual annotations.

class spacr.regression_annotation.Strategy[source]

Describe one cell-selection strategy.

Parameters:
  • key – the stored value; stable across releases.

  • title – what the entry is called on screen.

  • purpose – what it is for, in one sentence.

  • cost – Principal limitation or trade-off, in one sentence.

  • implemented – Whether run() can execute the strategy.

  • needs – Required data components: 'score', 'features', 'labels', 'controls'.

describe() → str[source]

Return the strategy purpose and limitations as one paragraph.

spacr.regression_annotation.candidate_feature_columns(frame: pandas.DataFrame, score_column: str = DEFAULT_SCORE_COLUMN) → Tuple[str, ...][source]

The columns that COULD be features, judged on names and dtypes alone.

feature_columns() is the authority and it reads every value: a numeric column that does not vary is not a feature, and only the values say so. This is the cheap half of that question – one pass over frame.dtypes rather than over the rows – for a caller that has to answer “can this strategy run at all” every time a chooser changes, on a table that may hold half a million objects.

It is therefore the OPTIMISTIC answer: a table this accepts can still be refused by feature_columns(), and the run is what refuses it. A table this rejects has no measurement column under any reading.

Parameters:
  • frame – the object table.

  • score_column – the score, which is never a feature.

Returns:

the candidate columns, in table order.

spacr.regression_annotation.feature_columns(frame: pandas.DataFrame, score_column: str = DEFAULT_SCORE_COLUMN, explicit: Sequence[str] | None = None) → Tuple[str, ...][source]

Return measurement columns suitable for model fitting.

Identifier columns, classifier outputs, and invariant columns are excluded to reduce plate-layout and score leakage.

Parameters:
  • frame – the object table.

  • score_column – the score, which is never a feature.

  • explicit – Optional explicit feature list. Missing columns raise an error rather than being removed silently.

Returns:

the columns, in table order.

Raises:

AnnotationStrategyError – an explicit column is absent, or nothing at all is left to fit on.

spacr.regression_annotation.feature_views(columns: Sequence[str]) → Dict[str, Tuple[str, ...]][source]

Partition feature columns into intensity and shape views.

Feature families are assigned with spacr.column_groups.classify(). Unclassified columns are distributed alternately so both views remain usable; the resulting column counts are included in strategy reports.

Parameters:

columns – the feature columns.

Returns:

{'intensity': (...), 'shape': (...)}.

spacr.regression_annotation.implemented_keys() → Tuple[str, ...][source]

Return the strategy keys currently supported by run().

spacr.regression_annotation.menu() → Tuple[str, ...][source]

Return one user-facing description for each strategy.

spacr.regression_annotation.missing_requirement(key: Any, frame: pandas.DataFrame | None, score_column: str = DEFAULT_SCORE_COLUMN, *, label_column: str = '', positive_control_wells: Sequence[str] = (), negative_control_wells: Sequence[str] = ()) → str[source]

Why key cannot run on this table, or "" when it can.

ASKED BEFORE THE RUN, NOT AFTER IT. Every refusal in this module is raised while a strategy is executing, which is the right place for it and the wrong time for a user: choosing “Diversity sampling over clusters” on a coefficient table with no measurement columns joined to it should say so on the control, not a run later. This is the cheap pre-flight that lets a chooser grey itself with the reason – it reads dtypes and one column, never the whole matrix, so it can be asked on every change of the menu.

IT IS THE OPTIMISTIC HALF. An empty answer means nothing this can see is missing; prepare() still reads the values and can still refuse – a score whose cells are all in one well, a feature column that turns out not to vary. What it will never do is stay silent about a table that has no score, no annotations, no measurements, or no control wells named for the strategy that needs them.

Parameters:
  • key – a strategy key, or a Strategy.

  • frame – the object rows on screen, or None.

  • score_column – the per-object classification score.

  • label_column – a column of human annotations, when one is named.

  • positive_control_wells – the positive control wells named.

  • negative_control_wells – the negative control wells named.

Returns:

one sentence naming what is missing and what to do about it, or “” when the strategy can be run.

Raises:

AnnotationStrategyError – key is not on the menu.

spacr.regression_annotation.prepare(request: AnnotationRequest, entry: Strategy | None = None) → Prepared[source]

Resolve columns, group identifiers, labels, and the holdout.

Complete independence groups are reserved before strategy-specific selection. Strategies cannot select rows from those groups.

Strategies that fit a model require measurement columns. Score-stratified and random sampling do not fit a model and therefore remain available for result tables that do not contain joined measurements.

Parameters:
  • request – what to run.

  • entry – the menu entry about to be run, when it is known. It is read only for what the strategy needs; None requires everything, which is the strict answer a caller with no entry to hand should get.

Returns:

the shared setup.

Raises:

AnnotationStrategyError – the table cannot support a leakage-safe split, or has no features, labels or score.

spacr.regression_annotation.readable_group(value: Any) → str[source]

Format a group identifier for display, such as plate1/r1/c1.

Parameters:

value – a group id from the splitter.

Returns:

the same identity with its parts separated visibly.

spacr.regression_annotation.run(key: Any, request: AnnotationRequest, prepared: Prepared | None = None) → AnnotationResult[source]

Run one strategy end to end.

Parameters:
  • key – the strategy key, or a Strategy.

  • request – what to run it on.

  • prepared – a setup built earlier by prepare(), when several strategies are being compared on one hold-out. Built here when it is not given.

Returns:

what the strategy chose, fitted and measured.

Raises:
spacr.regression_annotation.score_input_columns(frame: pandas.DataFrame, score_column: str = DEFAULT_SCORE_COLUMN, features: Sequence[str] | None = None, explicit: Sequence[str] | None = None, correlation_cut: float = 0.5) → Tuple[str, ...][source]

Identify known or likely inputs to the classification score.

Without an explicit input list, two rules are applied:

  • any column whose name marks it as a classifier output – the score itself, probabilities, logits, other prediction columns;

  • any feature whose absolute Spearman correlation with the score reaches correlation_cut.

Pass the classifier’s feature list as explicit when available to replace this correlation-based approximation.

Parameters:
  • frame – the object table.

  • score_column – the score.

  • features – the candidate feature columns.

  • explicit – the classifier’s own inputs, when they are known.

  • correlation_cut – the absolute rank correlation at which a feature counts as one of the score’s inputs.

Returns:

the columns, in table order, including the score itself.

spacr.regression_annotation.scored_cells(frame: pandas.DataFrame, score_column: str = DEFAULT_SCORE_COLUMN) → int[source]

How many cells carry a finite value in the score column.

spacr.regression_annotation.strategy(key: Any) → Strategy[source]

Return the menu entry for key.

Parameters:

key – a strategy key, or a Strategy.

Raises:

AnnotationStrategyError – no entry has that key; the message lists the ones that do.

spacr.regression_annotation.strategy_keys() → Tuple[str, ...][source]

Return all strategy keys in menu order.

spacr.regression_annotation.usable_annotations(frame: pandas.DataFrame, label_column: str) → int[source]

How many cells carry an annotation in label_column.

Zero when the column is absent, empty, or holds one class only – which is the same answer _reference_labels() gives, so a chooser and a run cannot disagree about whether there are labels to fit on.

Parameters:
  • frame – the object table.

  • label_column – the annotation column, or “”.

Returns:

the number of annotated cells, or 0 when they are unusable.

spacr.regression_annotation.wells_selected(groups: Sequence[Any], wanted: Sequence[str]) → numpy.ndarray[source]

Return a mask selecting rows whose group matches wanted.

A well matches when every token of the name given is one of the group’s own tokens, so 'r1_c1' matches plate1/r1/c1 and 'plate2_r1_c1' does not. Matching on tokens rather than on the complete string allows plate-map well names to match internal group identifiers without exposing their separator format.

Parameters:
  • groups – one group id per row.

  • wanted – the well names chosen.

Returns:

the mask; all-True when nothing was named.

spacr.regression_annotation.xgboost_available() → bool[source]

Return whether XGBoost can be imported.