spacr.sudoku¶
Infer cell-level guide assignments across wells.
Pooled screens provide guide counts for each well, not a guide identity for each cell. This module combines those count constraints with similarity in measurement space. It selects high-confidence anchor cells, propagates their labels over a k-nearest-neighbour graph, and projects the result onto the per-well guide fractions reported by sequencing.
Assignments can therefore borrow evidence for the same perturbation across wells while retaining an explicit abstention state. The propagated mass, competing-label evidence, and total anchor reach remain available separately so ambiguous or unsupported calls can be inspected instead of forced.
Notes
The propagation and class-mass normalization follow the label-propagation framework described by Zhu and Ghahramani (2003) and Zhou et al. (2004). The implementation has no trained graph-model weights; each result can be traced to its anchors and the sequencing constraint.
Classes¶
Cell-level guide assignments and their separate evidence components. |
Functions¶
|
Indices of the cells taken as near-certain examples of |
|
Project the propagated mass onto the counts sequencing implies. |
|
Propagate guide-anchor mass through a cell-similarity graph. |
|
Construct a symmetric k-nearest-neighbour cell-affinity graph. |
|
Assign guides to cells while retaining competing evidence separately. |
|
Assign guides sequentially in descending confidence order. |
Module Contents¶
- class spacr.sudoku.SudokuResult[source]¶
Cell-level guide assignments and their separate evidence components.
- Parameters:
guides – assigned guide for each input cell, using
ABSTAINwhen no guide is called; empty when there are no cells or candidate guides.affirm –
(n_cells, n_guides)normalized support propagated from each guide’s anchors.eliminate –
(n_cells, n_guides)competing-guide evidence, computed as one minusaffirm.reach –
(n_cells,)propagated support relative to the median positive reach; sequential runs retain each cell’s maximum across processed rounds.posterior –
(n_cells, n_guides)guide probabilities after applying the well-level fraction constraint.names – guide names in the matrix-column order used by
affirm,eliminate, andposterior.report – assignment counts, thresholds, anchor diagnostics, and warnings recorded by the run.
- property abstained: numpy.ndarray[source]¶
Boolean mask of the cells no guide was named for.
- spacr.sudoku.anchors_for(guide: str, wells: Sequence[str], fractions: Mapping[str, Mapping[str, float]], scores: numpy.ndarray, *, quantile: float = 0.9, min_fraction: float = 0.5, max_per_well: int = 50) numpy.ndarray[source]¶
Indices of the cells taken as near-certain examples of
guide.- Parameters:
guide – guide whose high-fraction wells and high-scoring cells are being selected as anchors.
wells – one well label per cell.
fractions –
{well: {guide: fraction}}.scores – the classification score per cell.
quantile – within an anchor well, the score quantile above which a cell is taken.
min_fraction – minimum sequencing fraction for a well to contribute anchors. This limits anchor selection to wells in which a high-scoring cell is plausibly associated with the guide.
max_per_well – maximum anchors contributed by one well.
- Returns:
cell indices, possibly empty.
Anchor selection uses the classifier score. To avoid circular inference,
sudoku()excludes that score from graph features by default; cell morphology then determines propagation beyond the anchors.
- spacr.sudoku.constrain_to_fractions(mass: numpy.ndarray, wells: Sequence[str], names: Sequence[str], fractions: Mapping[str, Mapping[str, float]], *, iterations: int = 200, tolerance: float = 1e-09) numpy.ndarray[source]¶
Project the propagated mass onto the counts sequencing implies.
- Parameters:
mass – graph-propagated cell-by-guide evidence matrix.
wells – one well identifier for each row of
mass.names – guide identifiers in the column order of
mass.fractions – sequencing fractions mapped by well and then guide.
Within each well, scale the guide columns so each sums to
pi_g * N_wand renormalise the rows to 1, alternately. This is iterative proportional fitting – the same fixed pointspacr.guide_attribution.posterior()uses, applied here to graph-propagated evidence instead of a one-dimensional likelihood.The row constraint prevents the graph from assigning every cell in a well to the guide with the greatest anchor mass when sequencing supports only a limited fraction for that guide.
- spacr.sudoku.propagate(graph, seeds: numpy.ndarray, *, alpha: float = 0.9, iterations: int = 100, tolerance: float = 0.0001, dtype=np.float32) numpy.ndarray[source]¶
Propagate guide-anchor mass through a cell-similarity graph.
- Parameters:
graph – Sparse affinity matrix returned by
similarity_graph(), with shape(n_cells, n_cells).seeds – Anchor weights with shape
(n_cells, n_guides). Rows for unanchored cells should contain zeros.alpha – Relative weight assigned to neighbouring cells. Values near one favour graph propagation;
alpha=1removes the seed term and is therefore unsuitable for guide assignment.iterations – Maximum number of propagation updates.
tolerance – Stop when the largest element-wise update is no greater than this value.
dtype – Floating-point type used for the seed and normalized graph matrices. The
float32default reduces memory and runtime for large screens; usefloat64when additional numerical precision is required.
- Returns:
Unnormalized propagated mass with shape
(n_cells, n_guides).
The update follows local-and-global consistency,
F <- alpha S F + (1 - alpha) Y, whereS = D^-1/2 W D^-1/2. The result is intentionally not row-normalized:sudoku()uses the total received mass to distinguish cells with weak support from confident assignments.
- spacr.sudoku.similarity_graph(features: numpy.ndarray, *, neighbours: int = 15, mutual: bool = True)[source]¶
Construct a symmetric k-nearest-neighbour cell-affinity graph.
- Parameters:
features –
(n_cells, n_features), standardised here.neighbours – number of nearest neighbours considered per cell.
mutual – retain an edge only when both cells identify each other as neighbours. This allows isolated outliers to have zero reach rather than receiving forced connections.
- Returns:
sparse CSR affinity matrix with a zero diagonal.
Edge weights use a heat kernel with a local scale defined by each cell’s distance to its kth neighbour. Local scaling supports populations with different sampling densities.
- spacr.sudoku.sudoku(features: numpy.ndarray, scores: numpy.ndarray, wells: Sequence[str], fractions: Mapping[str, Mapping[str, float]], guides: Sequence[str], *, anchors: Mapping[str, Sequence[int]] | None = None, neighbours: int = 15, alpha: float = 0.9, decision: float = DEFAULT_DECISION, reach_floor: float = DEFAULT_REACH_FLOOR, anchor_quantile: float = 0.9, anchor_min_fraction: float = 0.5, use_score_as_feature: bool = False, mutual: bool = True) SudokuResult[source]¶
Assign guides to cells while retaining competing evidence separately.
- Parameters:
features –
(n_cells, n_features)cell measurements.scores – classification score per cell, used to choose anchors, and by default not used as a graph feature.
wells – one well label per cell.
fractions –
{well: {guide: fraction}}.guides – guide identifiers to consider for assignment.
anchors – optional explicit anchor indices per guide, overriding
anchors_for().use_score_as_feature – include the classifier score in graph features. Disabled by default because the score also selects anchors; enabling it introduces circular evidence and is recorded in the result report.
- Returns:
the
SudokuResult.
Separate support and competing-label evidence distinguish four outcomes:
high support, low competition: confident assignment;
low support, high competition: confident exclusion;
high support, high competition: ambiguous between guides;
low support, low competition: unsupported by the anchor populations.
- spacr.sudoku.sudoku_all(features: numpy.ndarray, scores: numpy.ndarray, wells: Sequence[str], fractions: Mapping[str, Mapping[str, float]], ranking: Sequence[Tuple[str, float]], *, decision: float = DEFAULT_DECISION, max_guides: int = 50, **kwargs) SudokuResult[source]¶
Assign guides sequentially in descending confidence order.
- Parameters:
features – cell-by-feature matrix used to propagate anchor support.
scores – classification score for every cell, aligned to
features.wells – well identifier for every cell.
fractions – sequencing fractions as
{well: {guide: fraction}}.ranking –
[(guide, confidence)]in descending processing order. The caller defines confidence, for example by combining effect size and statistical significance.max_guides – maximum number of ranked guides to process.
- Returns:
one
SudokuResultover all cells.
Each round applies
sudoku()to unclaimed cells and removes accepted assignments. Processing stops when a round assigns no cells. Because this greedy procedure is order-sensitive,claimed_by_roundis retained in the report;spacr.annotation_validationevaluates sensitivity to ranking order.