spacr.annotation_validation

Evaluate guide-annotation strategies against measurable reference cases.

Real pooled screens do not provide cell-level guide ground truth. Validation therefore combines four complementary checks: simulated screens with known assignments, guide-to-well permutations that should reduce performance to chance, held-out controls, and order-sensitivity analysis for sequential methods.

Results report coverage and precision separately. This avoids treating a method that labels every cell unreliably as equivalent to one that abstains on uncertain cells and labels the remainder accurately.

Classes

Screen

Simulated screen with known cell-level guide assignments.

Verdict

How an annotation did against a known truth.

Functions

baseline_chance(→ List[str])

Sample each cell's guide from its well's sequencing fractions.

baseline_majority(→ List[str])

Assign every cell to the largest-fraction guide in its well.

benchmark(→ Dict[str, Dict[str, object]])

Evaluate annotation strategies on simulated and permuted screens.

calibration(→ List[Tuple[float, float, int]])

Summarize confidence calibration among non-abstaining calls.

count_agreement(→ Dict[str, object])

Compare per-well guide calls with sequencing fractions.

default_scenarios(→ Dict[str, Screen])

Return simulated screens that isolate major annotation confounds.

mixed_ratio_calibration(→ Dict[str, object])

Compare imaging-derived and sequencing-reported control mixtures.

mixture_proportion(→ float)

Estimate a well's positive-control share from cell features.

order_sensitivity(→ Dict[str, float])

Measure sensitivity of a sequential method to ranking order.

permuted(→ Screen)

Return a screen with guide-to-well assignments permuted.

score_annotation(→ Verdict)

Compare guide calls with known cell-level assignments.

synthesise(→ Screen)

Build a simulated screen with known guide assignments.

Module Contents

class spacr.annotation_validation.Screen[source]

Simulated screen with known cell-level guide assignments.

Parameters:
  • features – (n_cells, n_features) simulated cell-measurement matrix.

  • scores – classifier score for each cell after applying the scenario’s configured classifier error.

  • wells – well identifier for each simulated cell.

  • truth – known guide carried by each simulated cell.

  • fractions – biased and thresholded {well: {guide: fraction}} values exposed to the annotation method.

  • true_fractions – corresponding unbiased well-guide fractions retained as simulation truth.

  • guides – every guide represented by the simulated screen, in generator order.

  • meta – simulation parameters and auxiliary facts needed to interpret the scenario.

__len__() → int[source]

Return the simulated-cell count represented by the feature rows.

class spacr.annotation_validation.Verdict[source]

How an annotation did against a known truth.

Parameters:
  • coverage – fraction of cells receiving a non-abstaining annotation.

  • precision – fraction of annotated cells assigned to the correct guide.

  • recall – fraction of all evaluated cells assigned to the correct guide.

  • per_guide – guide names mapped to (precision, recall) values.

  • confusion – counts keyed by (true guide, called guide).

  • n – total number of truth labels evaluated.

summary() → str[source]

Format coverage, called-cell precision, and all-cell recall.

spacr.annotation_validation.baseline_chance(screen: Screen, *, seed: int = 0) → List[str][source]

Sample each cell’s guide from its well’s sequencing fractions.

Parameters:

screen – screen providing each cell’s well and the well-level guide fractions.

Together with baseline_majority(), this baseline brackets the performance available from sequencing counts without cell measurements.

spacr.annotation_validation.baseline_majority(screen: Screen) → List[str][source]

Assign every cell to the largest-fraction guide in its well.

Parameters:

screen – screen providing each cell’s well and the well-level guide fractions.

This sequencing-only baseline measures how much apparent performance is available from the guide fractions without cell-level features.

spacr.annotation_validation.benchmark(strategies: Mapping[str, Callable[[Screen], Sequence[str]]], scenarios: Mapping[str, Screen] | None = None, *, null_seed: int = 7) → Dict[str, Dict[str, object]][source]

Evaluate annotation strategies on simulated and permuted screens.

Parameters:
  • strategies – {name: screen -> one guide per cell}.

  • scenarios – {name: Screen}. Defaults to default_scenarios().

Returns:

{scenario: {strategy: {"real": Verdict, "null": Verdict}}}.

Each scenario is paired with its own guide-to-well permutation. The reported gain is the difference between real and null precision. The sequencing-only majority and chance baselines are included automatically.

spacr.annotation_validation.calibration(truth: Sequence[str], called: Sequence[str], confidence: Sequence[float], *, bins: int = 5, abstain: str = 'Non_annotated') → List[Tuple[float, float, int]][source]

Summarize confidence calibration among non-abstaining calls.

Parameters:
  • truth – reference guide name for every evaluated cell.

  • called – guide call or abstention label for every evaluated cell.

  • confidence – reported confidence corresponding to each call.

Returns one (mean confidence, observed accuracy, count) tuple per populated confidence bin. This distinguishes accurate probability estimates from labels that happen to have high aggregate precision.

spacr.annotation_validation.count_agreement(called: Sequence[str], wells: Sequence[str], reported: Mapping[str, float], guide: str) → Dict[str, object][source]

Compare per-well guide calls with sequencing fractions.

Returns reported and called shares for each usable well together with median and worst absolute errors. Count agreement evaluates aggregate calibration; it does not establish that individual cells received the correct guide.

spacr.annotation_validation.default_scenarios(seed: int = 0) → Dict[str, Screen][source]

Return simulated screens that isolate major annotation confounds.

Separate scenarios cover no feature signal, incomplete penetrance, fraction inflation, classifier error, crowded wells, and their combined realistic case.

spacr.annotation_validation.mixed_ratio_calibration(features: numpy.ndarray, wells: Sequence[str], reported: Mapping[str, float], *, pure_pc_wells: Sequence[str] | None = None, pure_nc_wells: Sequence[str] | None = None, pure_low: float = 0.05, pure_high: float = 0.95) → Dict[str, object][source]

Compare imaging-derived and sequencing-reported control mixtures.

Parameters:
  • features – (n_cells, n_features) over the control wells.

  • wells – one well label per cell.

  • reported – {well: PC fraction sequencing reported}.

  • pure_low – at or below this, a well is taken as pure NC.

  • pure_high – at or above this, a well is taken as pure PC.

Returns:

slope, intercept, the per-well estimates, and what was used.

A slope near one indicates agreement between cellular and sequencing fractions. Values below one indicate that sequencing reports a larger positive-control fraction than imaging estimates.

The median pairwise slope limits the influence of one-sided contamination: a hit sharing a control well can add phenotype-positive cells, whereas a least-squares slope would be pulled toward that contaminated well.

spacr.annotation_validation.mixture_proportion(features: numpy.ndarray, positive: numpy.ndarray, negative: numpy.ndarray) → float[source]

Estimate a well’s positive-control share from cell features.

Parameters:
  • features – (n_cells, n_features) for one mixed well.

  • positive – the pure-PC reference cells.

  • negative – the pure-NC reference cells.

Returns:

the estimated proportion, clipped to [0, 1].

Projects the well’s mean onto the line between the two reference means:

mu_w = pi * mu_PC + (1 - pi) * mu_NC

solved by least squares. The method requires no per-cell labels in the mixed well and clips the estimate to [0, 1]. It uses feature means rather than fitting a high-dimensional mixture density.

spacr.annotation_validation.order_sensitivity(run: Callable[[Sequence[Tuple[str, float]]], Sequence[str]], ranking: Sequence[Tuple[str, float]], *, repeats: int = 3, seed: int = 0) → Dict[str, float][source]

Measure sensitivity of a sequential method to ranking order.

Parameters:
  • run – takes a ranking and returns one guide name per cell.

  • ranking – the honest order.

Returns:

{"changed": share of cells that differ, "repeats": n}.

Each repeat randomly swaps adjacent ranking entries, preserving the broad confidence order while perturbing ties and near-ties. The returned mean and worst changed-cell shares quantify order dependence.

spacr.annotation_validation.permuted(screen: Screen, *, seed: int = 0) → Screen[source]

Return a screen with guide-to-well assignments permuted.

Parameters:

screen – internally consistent screen whose well-level guide assignments will be shuffled.

The permutation preserves cells, features, classifier scores, well sizes, and the distribution of guide fractions while breaking the sequencing-to- imaging relationship. Truth and fractions are permuted together so the returned object remains internally consistent for null benchmarking.

spacr.annotation_validation.score_annotation(truth: Sequence[str], called: Sequence[str], *, abstain: str = 'Non_annotated', guides: Sequence[str] | None = None) → Verdict[source]

Compare guide calls with known cell-level assignments.

Parameters:
  • truth – reference guide name for every evaluated cell.

  • called – guide call or abstention label for every evaluated cell.

Coverage is the fraction of cells that receive a non-abstaining call. Precision is calculated only among called cells, while recall is the fraction of all cells called correctly. The result also contains per-guide precision and recall plus the complete confusion counts.

spacr.annotation_validation.synthesise(*, wells: int = 24, guides_per_well: int = 4, cells_per_well: int = 60, features: int = 6, effect: float = 2.0, penetrance: float = 1.0, fraction_bias: float = 1.0, fraction_threshold: float = 0.0, classifier_accuracy: float = 1.0, guides: int = 8, seed: int = 0) → Screen[source]

Build a simulated screen with known guide assignments.

Parameters:
  • effect – how far a guide’s cells sit from the origin in feature space, in standard deviations. At zero the simulated guides are indistinguishable in feature space.

  • penetrance – the share of a guide’s cells that actually show its phenotype. Remaining cells are drawn around the origin.

  • fraction_bias – multiply the reported fractions by this, then renormalise. 1.8 reproduces what normalised_share documents on a screen after thresholding.

  • fraction_threshold – drop guides below this share BEFORE renormalising.

  • classifier_accuracy – the score is flipped toward the wrong tail this often. Values closer to 1 produce better-separated classes.

Returns:

the Screen.