spacr.classifier_quality

Measure classifier performance and correct observed class fractions.

A confusion matrix requires labelled outcomes; an unlabelled score column cannot establish sensitivity or specificity by itself. Use a labelled test split or out-of-fold predictions when available. When labels are unavailable, deconvolve() estimates two score distributions and reports whether their separation is sufficient to support the estimate.

Classifier errors bias the observed positive fraction rather than merely widening its uncertainty. The correction helpers expose that bias across prevalence levels and refuse the Rogan–Gladen correction when sensitivity and specificity do not identify it.

Classes

Confusion

Binary confusion counts and their operating characteristics.

Functions

best_threshold(→ Confusion)

Select an operating point for annotation.

confusion(→ Confusion)

The confusion matrix from a labelled split.

deconvolve(→ Dict[str, float])

Estimate two class distributions from unlabelled scores.

discover_test_splits(→ Dict[str, str])

Find classifier test outputs under a screen directory.

from_test_split(→ Dict[str, float])

Sensitivity and specificity from a written test split.

inflation_by_prevalence() → List[Dict[str, float]])

Quantify classifier-induced fraction inflation across prevalences.

measure_screen(→ Dict[str, Dict[str, float]])

Measure sensitivity and specificity for each plate in a screen.

operating_points(→ List[Confusion])

Return confusion matrices across quantile-based score thresholds.

rogan_gladen(→ Dict[str, float])

Apply the Rogan--Gladen correction to an observed positive share.

sensitivity_by_prevalence(→ List[Dict[str, float]])

Measure classifier performance across well-prevalence bands.

training_wells() → numpy.ndarray)

Return the cells belonging to classifier-training columns.

Module Contents

class spacr.classifier_quality.Confusion[source]

Binary confusion counts and their operating characteristics.

Parameters:
  • true_positive – number of positively labelled cells whose score meets the threshold.

  • false_positive – number of negatively labelled cells whose score meets the threshold.

  • true_negative – number of negatively labelled cells whose score falls below the threshold.

  • false_negative – number of positively labelled cells whose score falls below the threshold.

  • threshold – score cutoff used to classify cells as positive.

summary() → str[source]

Return sensitivity, specificity, accuracy, and prevalence text.

property accuracy: float[source]

Return the correctly classified share of all cells.

Accuracy should be interpreted with sensitivity and specificity when class prevalence is imbalanced.

property prevalence: float[source]

Return the positively labelled share of all cells.

property sensitivity: float[source]

Return the true-positive rate among positive cells.

property specificity: float[source]

Return the true-negative rate among negative cells.

property usable: bool[source]

Whether a Rogan-Gladen correction can be made from this.

se + sp <= 1 means the classifier carries no information at the chosen threshold, and the correction divides by se + sp - 1.

spacr.classifier_quality.best_threshold(scores: Sequence[float], labels: Sequence[bool], *, criterion: str = 'youden') → Confusion[source]

Select an operating point for annotation.

Parameters:
  • scores – classifier score for each labelled cell.

  • labels – true for cells belonging to the positive class, aligned one-to-one with scores.

  • criterion – 'youden' maximises se + sp - 1, which is the denominator of the Rogan–Gladen correction and therefore favours stable prevalence correction.

spacr.classifier_quality.confusion(scores: Sequence[float], labels: Sequence[bool], threshold: float = 0.5) → Confusion[source]

The confusion matrix from a labelled split.

Parameters:
  • scores – the classifier’s score per cell.

  • labels – True where the cell really is the positive class.

  • threshold – score at or above which the call is positive.

spacr.classifier_quality.deconvolve(scores: Sequence[float], *, seed: int = 0) → Dict[str, float][source]

Estimate two class distributions from unlabelled scores.

Parameters:

scores – unlabelled classifier scores to model as a two-component mixture; non-finite values are ignored.

A two-component Gaussian mixture estimates prevalence, sensitivity, specificity, and a midpoint threshold. separation is the distance between component means in pooled standard deviations; the result marks estimates trustworthy only when separation is at least two. This is a model-based fallback and is weaker evidence than a labelled test split.

spacr.classifier_quality.discover_test_splits(root: str, *, pattern: str = '*test_*.csv') → Dict[str, str][source]

Find classifier test outputs under a screen directory.

Parameters:
  • root – the screen directory holding one folder per plate.

  • pattern – how the training code named them. The default matches both shapes spaCR writes – *_test_acc.csv and *_test_result.csv.

Returns:

{plate folder name: path}, with the newest matching file selected when a plate contains several.

spacr.classifier_quality.from_test_split(path: str, *, threshold: float | None = None) → Dict[str, float][source]

Sensitivity and specificity from a written test split.

Parameters:
  • path – a per-cell *_test_acc.csv or a summary *_test_result.csv.

  • threshold – for a per-cell file, the score to call positive at. None takes the Youden point. Ignored for a summary file, which has already chosen one.

Returns:

sensitivity, specificity, threshold, accuracy, sample count, and per_cell indicating which file shape was read. Only per-cell files support recalculation at a different threshold.

spacr.classifier_quality.inflation_by_prevalence(sensitivity: float, specificity: float, *, prevalences: Sequence[float] = (0.5, 0.3, 0.2, 0.1, 0.05, 0.02, 0.01)) → List[Dict[str, float]][source]

Quantify classifier-induced fraction inflation across prevalences.

For each true prevalence, the function reports the expected observed prevalence, the Rogan–Gladen-corrected value, and their ratio. False positives can dominate rare classes because they are applied to the much larger negative population.

spacr.classifier_quality.measure_screen(root: str, *, pattern: str = '*test_*.csv', threshold: float | None = None) → Dict[str, Dict[str, float]][source]

Measure sensitivity and specificity for each plate in a screen.

Parameters:

root – screen directory containing one result folder per plate.

Plates are evaluated separately because their classifiers and selected thresholds can differ. The return value maps plate-folder names to the metrics produced by from_test_split().

spacr.classifier_quality.operating_points(scores: Sequence[float], labels: Sequence[bool], *, steps: int = 50) → List[Confusion][source]

Return confusion matrices across quantile-based score thresholds.

Parameters:
  • scores – classifier score for each labelled cell.

  • labels – true for cells belonging to the positive class, aligned one-to-one with scores.

The sequence exposes the sensitivity-specificity trade-off rather than evaluating only the conventional threshold of 0.5.

spacr.classifier_quality.rogan_gladen(observed: float, sensitivity: float, specificity: float, *, n: int | None = None) → Dict[str, float][source]

Apply the Rogan–Gladen correction to an observed positive share.

p_true = (p_observed - (1 - sp)) / (se + sp - 1)

Parameters:
  • observed – observed share called positive, conventionally between zero and one.

  • sensitivity – true-positive rate at the chosen classifier threshold.

  • specificity – true-negative rate at the chosen classifier threshold.

  • n – the number of cells, if the standard error is wanted. The correction inflates variance by 1 / (se + sp - 1)^2.

The result includes the unclipped denominator, a clipping indicator, and variance inflation. Correction is unusable when se + sp is one.

spacr.classifier_quality.sensitivity_by_prevalence(scores: Sequence[float], labels: Sequence[bool], wells: Sequence[str], *, threshold: float = 0.5, bins: int = 4) → List[Dict[str, float]][source]

Measure classifier performance across well-prevalence bands.

Parameters:
  • scores – classifier score for each labelled cell.

  • labels – true for cells belonging to the positive class.

  • wells – well label for each score and truth value.

Returns one row per populated band with prevalence, sensitivity, specificity, accuracy, and cell count. Dependence on prevalence can reveal that a classifier is using well context rather than only cell phenotype.

spacr.classifier_quality.training_wells(wells: Sequence[str], *, columns: Sequence[int] = (1, 2)) → numpy.ndarray[source]

Return the cells belonging to classifier-training columns.

Training wells must be excluded from performance calibration; otherwise the calibration measures in-sample fit. Both r1_c2 and c2 well labels are accepted. Unparseable labels are retained for validation rather than being silently classified as training data.

Parameters:
  • wells (sequence of str) – One well label per cell.

  • columns (sequence of int, default=(1, 2)) – One-based column numbers used for classifier training.

Returns:

numpy.ndarray – Boolean mask with one element per input well.

Nested helpers

best_threshold.value(point: Confusion) → float

Return Youden’s J, ranking a non-finite result last.

spacr/classifier_quality.py:168

deconvolve.above(mean, spread)

Estimate the Gaussian probability above the fitted midpoint.

A zero-width component is treated as a definite side of the cut rather than passed to a division by zero.

spacr/classifier_quality.py:290