spacr.fraction_calibration

Estimate fraction_threshold from control-well concordance.

fraction_threshold removes a gRNA from a well when its share of reads is below the selected limit. This changes the number and normalized abundance of guides assigned to each well and therefore affects downstream coefficients.

For each candidate threshold, this module compares the positive-control read fraction from sequencing with the positive phenotype fraction from imaging. Candidates are scored by the median absolute well-level disagreement, median(abs(imaging - sequencing)). The fit uses the median of pairwise slopes to reduce sensitivity to one-sided contamination from screen hits. Slope, intercept, residual, disagreement and guides per well are retained for review at every candidate.

The selected threshold establishes concordance within control wells; it does not estimate performance in wells containing a complex guide library.

Functions

compare_normalisations(→ Dict[str, Any])

Compare calibration fits using raw and normalised guide fractions.

describe(→ str)

The sweep in one line, for a log or a run summary.

reported_control_share(→ Dict[str, float])

What sequencing says the positive control's share of each well is.

sweep_fraction_threshold(, well_column, guide_column, ...)

Choose fraction_threshold by fitting imaging on sequencing at each.

well_fractions(→ pandas.DataFrame)

Each gRNA's share of its well, under one candidate threshold.

Module Contents

spacr.fraction_calibration.compare_normalisations(counts: pandas.DataFrame, features: numpy.ndarray, wells: Sequence[str], *, threshold: float = 0.02, **kwargs: Any) → Dict[str, Any][source]

Compare calibration fits using raw and normalised guide fractions.

Both fits use the same control wells and threshold. The comparison therefore isolates the effect of guide-fraction normalisation. Individual slopes combine penetrance and fraction bias and should not be interpreted as absolute calibrations; their ratio cancels the shared penetrance term.

Parameters:
  • counts – Control-well read counts, with one row per gRNA and well.

  • features – Feature matrix of shape (n_cells, n_features) for the same wells.

  • wells – One well identifier per cell.

  • threshold – Feature threshold used in both fits.

  • kwargs – Additional arguments for sweep_fraction_threshold(). normalise and candidates are controlled by this function.

Returns:

The raw and normalised fits, their slope ratio, and the normalisation with the smaller median absolute disagreement.

spacr.fraction_calibration.describe(result: Mapping[str, Any]) → str[source]

The sweep in one line, for a log or a run summary.

Parameters:

result – calibration result returned by sweep_fraction_threshold().

spacr.fraction_calibration.reported_control_share(fractions: pandas.DataFrame, positive_guide: str, *, well_column: str = 'prc', guide_column: str = 'grna') → Dict[str, float][source]

What sequencing says the positive control’s share of each well is.

Parameters:
  • fractions – per-well guide fractions, typically returned by well_fractions().

  • positive_guide – guide identifier whose share is reported.

Returns:

{well: fraction}, zero for a well the control did not survive the threshold in – which is a measurement, not a gap: the threshold decided that guide was not there.

spacr.fraction_calibration.sweep_fraction_threshold(counts: pandas.DataFrame, features: numpy.ndarray, wells: Sequence[str], *, positive_guide: str, pure_pc_wells: Sequence[str], pure_nc_wells: Sequence[str], candidates: Sequence[float] = DEFAULT_THRESHOLD_CANDIDATES, normalise: bool = True, sensitivity: float | None = None, specificity: float | None = None, minimum_wells: int = MINIMUM_CALIBRATION_WELLS, training_columns: Sequence[int] = (1, 2), well_column: str = 'prc', guide_column: str = 'grna', count_column: str = 'count') → Dict[str, Any][source]

Choose fraction_threshold by fitting imaging on sequencing at each.

Parameters:
  • counts – the control wells’ read counts, one row per gRNA per well.

  • features – (n_cells, n_features) over those same wells.

  • wells – one well label per cell.

  • positive_guide – the gRNA whose share is being calibrated.

  • pure_pc_wells – wells designated as entirely positive control in the plate design. They are not inferred from the read fraction being calibrated.

  • pure_nc_wells – wells that are entirely negative control, likewise.

  • candidates – the thresholds to try.

  • normalise – sweep with normalise_fraction on or off.

  • sensitivity – classifier sensitivity used for Rogan–Gladen rescaling of the fitted slope. No cross-experiment default is assumed.

  • specificity – classifier specificity used with sensitivity; aggregate accuracy is not an equivalent substitute.

  • minimum_wells – refuse to report below this many usable wells.

  • training_columns – plate columns the classifier was trained on.

Returns:

the chosen threshold, the reason, and every candidate’s fit.

If no candidate meets the evidence requirement, chosen is None and reason describes the limiting condition.

spacr.fraction_calibration.well_fractions(counts: pandas.DataFrame, *, threshold: float = 0.0, normalise: bool = True, well_column: str = 'prc', guide_column: str = 'grna', count_column: str = 'count') → pandas.DataFrame[source]

Each gRNA’s share of its well, under one candidate threshold.

Parameters:
  • counts – one row per gRNA per well, with a read count.

  • threshold – drop a gRNA whose share is below this.

  • normalise – rescale the survivors of one well to sum to one, which is what the pipeline’s normalise_fraction does. Off, a share is measured against every read the well produced, including the reads the threshold discarded.

Returns:

the surviving rows with a fraction column.

The two settings interact and that is the point of sweeping them together: normalising raises every surviving share, and by more the more the threshold removed.