spacr.sudoku

Infer cell-level guide assignments across wells.

Pooled screens provide guide counts for each well, not a guide identity for each cell. This module combines those count constraints with similarity in measurement space. It selects high-confidence anchor cells, propagates their labels over a k-nearest-neighbour graph, and projects the result onto the per-well guide fractions reported by sequencing.

Assignments can therefore borrow evidence for the same perturbation across wells while retaining an explicit abstention state. The propagated mass, competing-label evidence, and total anchor reach remain available separately so ambiguous or unsupported calls can be inspected instead of forced.

Notes

The propagation and class-mass normalization follow the label-propagation framework described by Zhu and Ghahramani (2003) and Zhou et al. (2004). The implementation has no trained graph-model weights; each result can be traced to its anchors and the sequencing constraint.

Classes

SudokuResult

Cell-level guide assignments and their separate evidence components.

Functions

anchors_for(→ numpy.ndarray)

Indices of the cells taken as near-certain examples of guide.

constrain_to_fractions(→ numpy.ndarray)

Project the propagated mass onto the counts sequencing implies.

propagate(→ numpy.ndarray)

Propagate guide-anchor mass through a cell-similarity graph.

similarity_graph(features, *[, neighbours, mutual])

Construct a symmetric k-nearest-neighbour cell-affinity graph.

sudoku(→ SudokuResult)

Assign guides to cells while retaining competing evidence separately.

sudoku_all(→ SudokuResult)

Assign guides sequentially in descending confidence order.

Module Contents

class spacr.sudoku.SudokuResult[source]

Cell-level guide assignments and their separate evidence components.

Parameters:
  • guides – assigned guide for each input cell, using ABSTAIN when no guide is called; empty when there are no cells or candidate guides.

  • affirm – (n_cells, n_guides) normalized support propagated from each guide’s anchors.

  • eliminate – (n_cells, n_guides) competing-guide evidence, computed as one minus affirm.

  • reach – (n_cells,) propagated support relative to the median positive reach; sequential runs retain each cell’s maximum across processed rounds.

  • posterior – (n_cells, n_guides) guide probabilities after applying the well-level fraction constraint.

  • names – guide names in the matrix-column order used by affirm, eliminate, and posterior.

  • report – assignment counts, thresholds, anchor diagnostics, and warnings recorded by the run.

called() → int[source]

How many cells were annotated.

property abstained: numpy.ndarray[source]

Boolean mask of the cells no guide was named for.

spacr.sudoku.anchors_for(guide: str, wells: Sequence[str], fractions: Mapping[str, Mapping[str, float]], scores: numpy.ndarray, *, quantile: float = 0.9, min_fraction: float = 0.5, max_per_well: int = 50) → numpy.ndarray[source]

Indices of the cells taken as near-certain examples of guide.

Parameters:
  • guide – guide whose high-fraction wells and high-scoring cells are being selected as anchors.

  • wells – one well label per cell.

  • fractions – {well: {guide: fraction}}.

  • scores – the classification score per cell.

  • quantile – within an anchor well, the score quantile above which a cell is taken.

  • min_fraction – minimum sequencing fraction for a well to contribute anchors. This limits anchor selection to wells in which a high-scoring cell is plausibly associated with the guide.

  • max_per_well – maximum anchors contributed by one well.

Returns:

cell indices, possibly empty.

Anchor selection uses the classifier score. To avoid circular inference, sudoku() excludes that score from graph features by default; cell morphology then determines propagation beyond the anchors.

spacr.sudoku.constrain_to_fractions(mass: numpy.ndarray, wells: Sequence[str], names: Sequence[str], fractions: Mapping[str, Mapping[str, float]], *, iterations: int = 200, tolerance: float = 1e-09) → numpy.ndarray[source]

Project the propagated mass onto the counts sequencing implies.

Parameters:
  • mass – graph-propagated cell-by-guide evidence matrix.

  • wells – one well identifier for each row of mass.

  • names – guide identifiers in the column order of mass.

  • fractions – sequencing fractions mapped by well and then guide.

Within each well, scale the guide columns so each sums to pi_g * N_w and renormalise the rows to 1, alternately. This is iterative proportional fitting – the same fixed point spacr.guide_attribution.posterior() uses, applied here to graph-propagated evidence instead of a one-dimensional likelihood.

The row constraint prevents the graph from assigning every cell in a well to the guide with the greatest anchor mass when sequencing supports only a limited fraction for that guide.

spacr.sudoku.propagate(graph, seeds: numpy.ndarray, *, alpha: float = 0.9, iterations: int = 100, tolerance: float = 0.0001, dtype=np.float32) → numpy.ndarray[source]

Propagate guide-anchor mass through a cell-similarity graph.

Parameters:
  • graph – Sparse affinity matrix returned by similarity_graph(), with shape (n_cells, n_cells).

  • seeds – Anchor weights with shape (n_cells, n_guides). Rows for unanchored cells should contain zeros.

  • alpha – Relative weight assigned to neighbouring cells. Values near one favour graph propagation; alpha=1 removes the seed term and is therefore unsuitable for guide assignment.

  • iterations – Maximum number of propagation updates.

  • tolerance – Stop when the largest element-wise update is no greater than this value.

  • dtype – Floating-point type used for the seed and normalized graph matrices. The float32 default reduces memory and runtime for large screens; use float64 when additional numerical precision is required.

Returns:

Unnormalized propagated mass with shape (n_cells, n_guides).

The update follows local-and-global consistency, F <- alpha S F + (1 - alpha) Y, where S = D^-1/2 W D^-1/2. The result is intentionally not row-normalized: sudoku() uses the total received mass to distinguish cells with weak support from confident assignments.

spacr.sudoku.similarity_graph(features: numpy.ndarray, *, neighbours: int = 15, mutual: bool = True)[source]

Construct a symmetric k-nearest-neighbour cell-affinity graph.

Parameters:
  • features – (n_cells, n_features), standardised here.

  • neighbours – number of nearest neighbours considered per cell.

  • mutual – retain an edge only when both cells identify each other as neighbours. This allows isolated outliers to have zero reach rather than receiving forced connections.

Returns:

sparse CSR affinity matrix with a zero diagonal.

Edge weights use a heat kernel with a local scale defined by each cell’s distance to its kth neighbour. Local scaling supports populations with different sampling densities.

spacr.sudoku.sudoku(features: numpy.ndarray, scores: numpy.ndarray, wells: Sequence[str], fractions: Mapping[str, Mapping[str, float]], guides: Sequence[str], *, anchors: Mapping[str, Sequence[int]] | None = None, neighbours: int = 15, alpha: float = 0.9, decision: float = DEFAULT_DECISION, reach_floor: float = DEFAULT_REACH_FLOOR, anchor_quantile: float = 0.9, anchor_min_fraction: float = 0.5, use_score_as_feature: bool = False, mutual: bool = True) → SudokuResult[source]

Assign guides to cells while retaining competing evidence separately.

Parameters:
  • features – (n_cells, n_features) cell measurements.

  • scores – classification score per cell, used to choose anchors, and by default not used as a graph feature.

  • wells – one well label per cell.

  • fractions – {well: {guide: fraction}}.

  • guides – guide identifiers to consider for assignment.

  • anchors – optional explicit anchor indices per guide, overriding anchors_for().

  • use_score_as_feature – include the classifier score in graph features. Disabled by default because the score also selects anchors; enabling it introduces circular evidence and is recorded in the result report.

Returns:

the SudokuResult.

Separate support and competing-label evidence distinguish four outcomes:

  • high support, low competition: confident assignment;

  • low support, high competition: confident exclusion;

  • high support, high competition: ambiguous between guides;

  • low support, low competition: unsupported by the anchor populations.

spacr.sudoku.sudoku_all(features: numpy.ndarray, scores: numpy.ndarray, wells: Sequence[str], fractions: Mapping[str, Mapping[str, float]], ranking: Sequence[Tuple[str, float]], *, decision: float = DEFAULT_DECISION, max_guides: int = 50, **kwargs) → SudokuResult[source]

Assign guides sequentially in descending confidence order.

Parameters:
  • features – cell-by-feature matrix used to propagate anchor support.

  • scores – classification score for every cell, aligned to features.

  • wells – well identifier for every cell.

  • fractions – sequencing fractions as {well: {guide: fraction}}.

  • ranking – [(guide, confidence)] in descending processing order. The caller defines confidence, for example by combining effect size and statistical significance.

  • max_guides – maximum number of ranked guides to process.

Returns:

one SudokuResult over all cells.

Each round applies sudoku() to unclaimed cells and removes accepted assignments. Processing stops when a round assigns no cells. Because this greedy procedure is order-sensitive, claimed_by_round is retained in the report; spacr.annotation_validation evaluates sensitivity to ranking order.