spacr.fraction_calibration¶
Estimate fraction_threshold from control-well concordance.
fraction_threshold removes a gRNA from a well when its share of reads is
below the selected limit. This changes the number and normalized abundance of
guides assigned to each well and therefore affects downstream coefficients.
For each candidate threshold, this module compares the positive-control read
fraction from sequencing with the positive phenotype fraction from imaging.
Candidates are scored by the median absolute well-level disagreement,
median(abs(imaging - sequencing)). The fit uses the median of pairwise
slopes to reduce sensitivity to one-sided contamination from screen hits.
Slope, intercept, residual, disagreement and guides per well are retained for
review at every candidate.
The selected threshold establishes concordance within control wells; it does not estimate performance in wells containing a complex guide library.
Functions¶
|
Compare calibration fits using raw and normalised guide fractions. |
|
The sweep in one line, for a log or a run summary. |
|
What sequencing says the positive control's share of each well is. |
|
Choose |
|
Each gRNA's share of its well, under one candidate threshold. |
Module Contents¶
- spacr.fraction_calibration.compare_normalisations(counts: pandas.DataFrame, features: numpy.ndarray, wells: Sequence[str], *, threshold: float = 0.02, **kwargs: Any) Dict[str, Any][source]¶
Compare calibration fits using raw and normalised guide fractions.
Both fits use the same control wells and threshold. The comparison therefore isolates the effect of guide-fraction normalisation. Individual slopes combine penetrance and fraction bias and should not be interpreted as absolute calibrations; their ratio cancels the shared penetrance term.
- Parameters:
counts – Control-well read counts, with one row per gRNA and well.
features – Feature matrix of shape
(n_cells, n_features)for the same wells.wells – One well identifier per cell.
threshold – Feature threshold used in both fits.
kwargs – Additional arguments for
sweep_fraction_threshold().normaliseandcandidatesare controlled by this function.
- Returns:
The raw and normalised fits, their slope ratio, and the normalisation with the smaller median absolute disagreement.
- spacr.fraction_calibration.describe(result: Mapping[str, Any]) str[source]¶
The sweep in one line, for a log or a run summary.
- Parameters:
result – calibration result returned by
sweep_fraction_threshold().
What sequencing says the positive control’s share of each well is.
- Parameters:
fractions – per-well guide fractions, typically returned by
well_fractions().positive_guide – guide identifier whose share is reported.
- Returns:
{well: fraction}, zero for a well the control did not survive the threshold in – which is a measurement, not a gap: the threshold decided that guide was not there.
- spacr.fraction_calibration.sweep_fraction_threshold(counts: pandas.DataFrame, features: numpy.ndarray, wells: Sequence[str], *, positive_guide: str, pure_pc_wells: Sequence[str], pure_nc_wells: Sequence[str], candidates: Sequence[float] = DEFAULT_THRESHOLD_CANDIDATES, normalise: bool = True, sensitivity: float | None = None, specificity: float | None = None, minimum_wells: int = MINIMUM_CALIBRATION_WELLS, training_columns: Sequence[int] = (1, 2), well_column: str = 'prc', guide_column: str = 'grna', count_column: str = 'count') Dict[str, Any][source]¶
Choose
fraction_thresholdby fitting imaging on sequencing at each.- Parameters:
counts – the control wells’ read counts, one row per gRNA per well.
features –
(n_cells, n_features)over those same wells.wells – one well label per cell.
positive_guide – the gRNA whose share is being calibrated.
pure_pc_wells – wells designated as entirely positive control in the plate design. They are not inferred from the read fraction being calibrated.
pure_nc_wells – wells that are entirely negative control, likewise.
candidates – the thresholds to try.
normalise – sweep with
normalise_fractionon or off.sensitivity – classifier sensitivity used for Rogan–Gladen rescaling of the fitted slope. No cross-experiment default is assumed.
specificity – classifier specificity used with
sensitivity; aggregate accuracy is not an equivalent substitute.minimum_wells – refuse to report below this many usable wells.
training_columns – plate columns the classifier was trained on.
- Returns:
the chosen threshold, the reason, and every candidate’s fit.
If no candidate meets the evidence requirement,
chosenisNoneandreasondescribes the limiting condition.
- spacr.fraction_calibration.well_fractions(counts: pandas.DataFrame, *, threshold: float = 0.0, normalise: bool = True, well_column: str = 'prc', guide_column: str = 'grna', count_column: str = 'count') pandas.DataFrame[source]¶
Each gRNA’s share of its well, under one candidate threshold.
- Parameters:
counts – one row per gRNA per well, with a read count.
threshold – drop a gRNA whose share is below this.
normalise – rescale the survivors of one well to sum to one, which is what the pipeline’s
normalise_fractiondoes. Off, a share is measured against every read the well produced, including the reads the threshold discarded.
- Returns:
the surviving rows with a
fractioncolumn.
The two settings interact and that is the point of sweeping them together: normalising raises every surviving share, and by more the more the threshold removed.