spacr.guide_attribution¶
A guide and a probability for every cell.
NOT spacr.attribution, which is the Grad-CAM and saliency library – “what
a trained classifier attends to”. This is “which guide is in this cell”, a
different question with an unfortunately similar name.
A pooled screen never observes which cell carried which guide: sequencing gives a FRACTION per well. But a guide with a measurable effect on the classification score leaves a trace in the scores themselves, and that is enough to attribute cells probabilistically.
prior pi_g the guide’s read fraction in the well, normalised likelihood f_g(s) how likely this score is if the cell carries g posterior r_ig ∝ pi_g * f_g(s_i)
AND THE CONSTRAINT THAT MAKES IT WORK:
sum_i r_ig = N * pi_g for every guide
Without it, two guides pointing the same way both claim the same top cells – “if i have 2 grnas both in the same direction and want to attribute cells to both my arithmatic wont work well because both pick the top hits which would be wrong”. With it, the total mass attributed to each guide is pinned to its read fraction, so they SPLIT the top cells in proportion instead. It is enforced by iterative proportional fitting, which is the same fixed point as fitting a mixture whose mixing proportions are known and fixed.
WHAT THIS IS NOT. It is not a genotype. Every name in this module says “attributed” for that reason: a reader who takes the answer for an observation has been misled by the column name alone.
TWO GUIDES WITH THE SAME EFFECT ARE UNIDENTIFIABLE, and the method says so rather than inventing a split: their posteriors come back at exactly pi_1 : pi_2 for every cell, which is flat, ambiguous and correct.
Classes¶
Guide assignment for each cell with its observed sequencing counts. |
|
What one cell was attributed, and how sure that is. |
|
Whether a guide can produce a real assignment, BEFORE any is made. |
Functions¶
|
Give every cell exactly one guide, with the counts sequencing implies. |
|
|
|
One |
|
How many INDEPENDENT measurements |
|
The guides' fractions, scaled to sum to 1. |
|
|
|
|
|
Can |
Module Contents¶
- class spacr.guide_attribution.Assignment[source]¶
Guide assignment for each cell with its observed sequencing counts.
- Variables:
guides – assigned guide name for each input cell, in input order.
cost – total negative log likelihood of the assignment; lower is better.
degeneracy – mean cost of changing a cell to its best alternative guide, in nats; values near zero indicate interchangeable solutions.
counts – number of cells assigned to each guide.
- class spacr.guide_attribution.Attribution[source]¶
What one cell was attributed, and how sure that is.
- Variables:
guide – assigned guide name, or
AMBIGUOUSwhen the posterior did not reach the call threshold.probability – largest posterior probability reached for the cell.
ambiguous – whether
probabilitywas below the call threshold.entropy – Shannon entropy of the cell’s guide posterior, in bits.
runner_up – second-highest guide name, or an empty string when no second guide exists.
- property called: bool[source]¶
Whether this cell was assigned a guide at all.
AMBIGUOUS IS NOT CALLED. A cell whose reads support two guides has no single perturbation, and counting it as called would attribute a phenotype to whichever guide happened to be listed first.
- Returns:
True when unambiguously assigned.
- class spacr.guide_attribution.Preflight[source]¶
Whether a guide can produce a real assignment, BEFORE any is made.
- Variables:
guide – guide whose achievable posterior was evaluated.
wells – number of wells containing a positive fraction of the guide.
callable_wells – wells in which the guide can reach
threshold.best – highest posterior ceiling reached across the guide’s wells.
threshold – posterior probability required for a guide call.
- spacr.guide_attribution.assign_well(scores: Sequence[float], fractions: Mapping[str, float], effects: Mapping[str, float], *, centre: float = 0.0, scale: float = 1.0, likelihood: str = 'lognormal', seed: int = 0) Assignment[source]¶
Give every cell exactly one guide, with the counts sequencing implies.
- Parameters:
scores – one classification or measurement score per cell, in the desired output order.
fractions – measured guide fractions for the well; usable positive values are normalized before integer cell counts are apportioned.
effects – fitted score effect keyed by guide; absent guides use a zero effect.
- Returns:
an
Assignment.guidesis one name per cell, in the order the scores came.
The counts are
round(N * pi_g)adjusted to sum to N exactly – a rounding that left a cell unassigned would break the one rule that makes this an assignment at all. The largest remainders take the slack, which is the standard apportionment and is deterministic.
- spacr.guide_attribution.attributable(effect: float, scale: float, prior: float, *, threshold: float = DEFAULT_THRESHOLD, others: Sequence[Tuple[float, float]] | None = None, span: float = 4.0, centre: float = 0.0, likelihood: str = 'lognormal', grid: int = 513) Tuple[bool, float][source]¶
(can_it, best_possible)– can this guide ever be called?Determine whether a guide can reach the assignment threshold within the supplied score range.
- Parameters:
effect – the guide’s fitted shift on the selected likelihood scale.
scale – the positive spread of scores used to express the plausible range and likelihood width; non-positive values resolve to one.
prior – the guide’s sequencing fraction in the well, clipped to the closed interval from zero to one.
others – the competition, as
(effect, prior)pairs. Omit it and the rest of the well is treated as one competitor with no effect, which provides the most permissive comparison.span – how far out, in units of
scale, a score may plausibly lie. The posterior ceiling is maximized across this finite range.
The calculation evaluates competing guides at the same candidate score. Opposite-signed effects can therefore separate more strongly than a flat competitor. Without a finite
spanthe plain-Bayes posterior may tend to one, so no useful finite ceiling would exist.
- spacr.guide_attribution.attribute_well(scores: Sequence[float], fractions: Mapping[str, float], effects: Mapping[str, float], *, threshold: float = DEFAULT_THRESHOLD, centre: float = 0.0, scale: float = 1.0, likelihood: str = 'lognormal', seed: int = 0) Tuple[Attribution, ...][source]¶
One
Attributionper cell, in the order the scores came.- Parameters:
scores – one classification or measurement score per cell, in the output order.
fractions – sequencing fractions keyed by guide for this well. The usable positive fractions are normalised before attribution.
effects – fitted score effect keyed by guide; a missing guide is treated as having zero effect.
threshold – “if it gets a grna above 0.55 then it gets annotated with that gene and the probability score”, and below it the cell is tagged
AMBIGUOUS– carrying the highest probability any guide reached anyway, because that number is the useful part of a refusal.seed – an exact tie is broken AT RANDOM, and the generator is seeded: spaCR runs are reproducible, and a coin flip that landed differently on a re-run would annotate the same screen two ways.
- spacr.guide_attribution.effective_dimension(matrix: numpy.ndarray) float[source]¶
How many INDEPENDENT measurements
matrix’s columns amount to.The participation ratio of the correlation matrix’s eigenvalues,
(sum lambda)^2 / sum lambda^2– 785 for 785 orthogonal columns, and 1 for 785 copies of one column. The same statistic the sweep uses to count a guide’s effective wells, applied to the other axis.- Parameters:
matrix – cells x measurements, already finite.
- Returns:
the effective count, between 1 and the number of columns.
- spacr.guide_attribution.normalise_fractions(fractions: Mapping[str, float]) Dict[str, float][source]¶
The guides’ fractions, scaled to sum to 1.
- Parameters:
fractions – guide names mapped to their measured fraction in one well. Missing, non-finite, and non-positive values are discarded.
“first calculate what all the fractions in that well add up to if not 1 then normalize to 1”. A well whose fractions sum to nothing usable comes back empty rather than uniform: inventing a prior where there is no sequencing is the one thing this must not do.
- spacr.guide_attribution.posterior(scores: Sequence[float], priors: Mapping[str, float], effects: Mapping[str, float], *, centre: float = 0.0, scale: float = 1.0, likelihood: str = 'lognormal', iterations: int = 200, tolerance: float = 1e-09) Tuple[numpy.ndarray, Tuple[str, ...]][source]¶
(r, guides)– each cell’s probability of carrying each guide.- Parameters:
scores – one finite classification or measurement score per cell.
priors – normalized guide fractions; insertion order defines the returned matrix columns.
effects – fitted score shifts keyed by guide. A missing guide uses a zero shift.
- Returns:
rwith one row per cell and one column per guide, rows summing to 1, and COLUMNS summing ton_cells * prior– the constraint that stops two same-direction guides both claiming the same cells.
Iterative proportional fitting: scale the columns to the target masses, renormalise the rows to 1, repeat. Both operations are monotone in the right sense and the fixed point satisfies both constraints at once. It is the same answer as a mixture fitted with its mixing proportions held at the sequencing fractions.
- spacr.guide_attribution.posterior_multivariate(measurements: numpy.ndarray, priors: Mapping[str, float], effects: Mapping[str, Sequence[float]], *, centres: Sequence[float] | None = None, scales: Sequence[float] | None = None, likelihood: str = 'lognormal', correct_for_correlation: bool = True, iterations: int = 200, tolerance: float = 1e-09) Tuple[numpy.ndarray, Tuple[str, ...], Dict[str, float]][source]¶
posterior(), but reading EVERY measurement instead of one score.Option C. Each guide carries a vector of effects – one per measurement – and a cell’s evidence is the summed log-density across them, scaled by the design effect (see the note above). The same iterative proportional fitting then applies, so the guide masses still match the sequencing.
- Parameters:
measurements – cells x measurements.
priors – normalised sequencing fraction keyed by guide. Its key order defines the columns in the returned posterior matrix.
effects –
{guide: [effect per measurement]}, in the columns’ order. A guide with no entry is flat, which is the honest prior for a guide nothing was fitted for.correct_for_correlation – scale the evidence by
effective_dimension / n_measurements. Leave it on unless the columns really are independent; see the note above for why.
- Returns:
(r, guides, diagnostics). The diagnostics carryn_measurements,effective_dimensionand thescale_factorapplied, because a reader has to be able to see how much the correction did.
- spacr.guide_attribution.preflight(guide, fractions_by_well, effects, *, scale, centre=0.0, threshold=DEFAULT_THRESHOLD, likelihood='lognormal', span=4.0) Preflight[source]¶
Can
guideever be called, and in how many of its wells?“WHICH GUIDES CAN BE ATTRIBUTED AT ALL, reported BEFORE anything is assigned. A guide whose effect is small against the spread of scores can never reach the threshold – not with more cells, not with a better fit.” This is that report, and until it had a caller it was a library function nobody saw.
- Parameters:
guide – the guide being asked about.
fractions_by_well –
{well: {guide: fraction}}– every guide in each well, because a ceiling is a comparison and comparing a guide against nothing returns its prior.effects – each guide’s fitted effect.
scale – the spread of the scores, per plate.
- Returns:
a
Preflight.
A guide is counted callable in a well when
attributable()says so THERE – the answer differs between wells because the prior does, and a guide at 0.6 of one well and 0.02 of another is not the same question.