spacr.hit_attribution

Resolve a well-level CRISPR hit back to candidate single cells.

Sequencing says which guide reads were present in a well. It does not say which cell carried a guide, and a read fraction is not automatically an infection fraction. This module keeps that distinction explicit:

  • build_hit_cell_frame() makes the honest review queue: target-well cells ranked in the hit’s phenotype direction.

  • fit_hit_attribution() adds a cross-fitted two-component hierarchical mixture. Guide fraction changes a learned well-level prior; it is never imposed as the mean cell probability.

  • write_hit_attribution() records versioned probabilities in their own tables. Hand annotations are untouched until an explicit promotion call.

The output is named hit_like_probability, not guide identity. Only a cell-resolved barcode or arrayed perturbation turns that inference into ground truth.

Exceptions

HitAttributionError

The requested cell attribution is ambiguous or not identifiable.

InsufficientDesignError

There are too few independent wells/plates to cross-fit honestly.

Classes

HitAttributionResult

Cross-fitted hit-like probabilities and independent-unit evidence.

HitInvestigationResult

Portable result bundle used by the GUI and database persistence.

HitRunContext

The exact regression result a cell investigation came from.

Functions

build_hit_cell_frame(→ pandas.DataFrame)

Join exact cells to well guide fractions and build a review ranking.

crossfit_candidate_probabilities(...)

Cross-fit a conservative morphology classifier from bag labels.

fit_hit_attribution(→ HitAttributionResult)

Estimate cross-fitted target-hit-like probabilities.

promote_calls(→ str)

Promote stored calls while recording every previous annotation value.

promote_hit_calls(→ str)

Explicitly promote positive calls to a fresh, reversible annotation.

quantify_candidate_enrichment(...)

Quantify candidate prevalence at the well experimental unit.

quantify_hit_enrichment(→ Dict[str, Any])

Per-well enrichment, well bootstrap CI, and a within-plate null.

revert_promotion(→ int)

Restore exactly the values replaced by promote_calls().

store_attribution(→ int)

Store the GUI investigation bundle under its immutable run context.

undo_hit_promotion(→ int)

Clear values written by one promotion; preserve the audit record.

write_hit_attribution(→ str)

Persist a versioned attribution without touching annotation columns.

Module Contents

exception spacr.hit_attribution.HitAttributionError[source]

Bases: ValueError

The requested cell attribution is ambiguous or not identifiable.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.hit_attribution.InsufficientDesignError[source]

Bases: HitAttributionError

There are too few independent wells/plates to cross-fit honestly.

Initialize self. See help(type(self)) for accurate signature.

class spacr.hit_attribution.HitAttributionResult[source]

Cross-fitted hit-like probabilities and independent-unit evidence.

Variables:
  • cells – cell-level frame with cross-fitted probabilities, calls, uncertainties, and fold assignments.

  • wells – well-level probability, prevalence, score, and guide-fraction summaries.

  • guide_evidence – per-guide dose-response and probability contrasts.

  • threshold_sensitivity – well-level prevalence contrasts across the evaluated probability thresholds.

  • validation – well-resampling and permutation validation statistics.

  • feature_columns – morphology features used to fit the mixture.

  • well_columns – columns that form the well identity.

  • object_columns – columns that form the stable cell-object identity.

  • target_gene – gene whose hit-like morphology was attributed.

  • target_guides – guides treated as evidence for target_gene.

  • score_column – input column holding the original phenotype score.

  • direction – phenotype direction, "positive" or "negative".

  • threshold – probability cutoff used for hit_like_call.

  • split_level – held-out grouping level used for cross-fitting.

  • random_seed – seed used for validation resampling.

  • source_regression_run – source regression run identifier persisted with the attribution so the morphology evidence remains traceable.

  • warnings – circularity or design caveats that must accompany the probabilities and summary rather than being lost after fitting.

summary() → str[source]

The effect, its bootstrap interval, and what qualifies it.

THE INTERVAL IS PART OF THE HEADLINE, not a detail underneath it: a prevalence difference quoted without one invites a reader to treat a noisy estimate as a finding.

Returns:

a one-line summary.

class spacr.hit_attribution.HitInvestigationResult[source]

Portable result bundle used by the GUI and database persistence.

Parameters:
  • attribution_run_id – unique identifier under which this investigation is persisted.

  • context – source regression hit and run-provenance contract.

  • cells – cell-level candidate probabilities, uncertainty, calls, and held-out fold assignments.

  • wells – well-level candidate-prevalence summary.

  • enrichment – well-level effect estimates, confidence intervals, and resampling statistics.

  • feature_columns – morphology-feature columns used by the cross-fitted classifier.

  • split_level – grouping level held out during cross-fitting, normally "plate" or "well".

  • warnings – design and fit caveats retained for display and persistence.

class spacr.hit_attribution.HitRunContext[source]

The exact regression result a cell investigation came from.

Parameters:
  • regression_results_folder – folder containing the regression result from which this investigation was launched.

  • regression_run_sha256 – SHA-256 digest identifying that source regression run.

  • gene – selected hit-gene identifier.

  • phenotype – regression phenotype for which the gene was selected.

  • effect – regression effect estimate for the selected gene and phenotype.

  • guides – guide identifiers assigned to the selected hit gene.

  • fdr – multiple-testing-adjusted significance of the selected hit, or nan when unavailable.

  • direction – phenotype direction used to rank candidate cells, normally "positive" or "negative".

spacr.hit_attribution.build_hit_cell_frame(cells: pandas.DataFrame, guide_fractions: pandas.DataFrame, *, target_guides: Sequence[str], score_column: str, direction: str = 'positive', guide_column: str = 'grna', fraction_column: str = 'fraction', well_columns: Sequence[str] = WELL_COLUMNS, object_columns: Sequence[str] = OBJECT_COLUMNS) → pandas.DataFrame[source]

Join exact cells to well guide fractions and build a review ranking.

The returned candidate_percentile ranks cells only within their well. It is deliberately not named a probability or infection call.

Parameters:
  • cells – cell-level measurements with the phenotype score and stable well and object identifiers.

  • guide_fractions – one sequencing-fraction row per guide and well.

  • target_guides – guide identifiers assigned to the hit gene.

  • score_column – column in cells containing the phenotype score.

spacr.hit_attribution.crossfit_candidate_probabilities(frame: pandas.DataFrame, *, feature_columns: Sequence[str] | None = None, target_column: str = 'target_well', prefer_plate: bool = True, random_seed: int = 0, n_splits: int = 5, threshold: float = 0.5) → Tuple[pandas.DataFrame, List[str], str, List[str]][source]

Cross-fit a conservative morphology classifier from bag labels.

This is the non-parametric alternative to fit_hit_attribution()’s hierarchical mixture. It predicts candidate_probability and never calls it infection probability. Model outputs, guide fractions, object identifiers and annotations are excluded from the default feature set.

Parameters:

frame – cell-level candidate frame with target-well labels and plate, row, and column identifiers.

spacr.hit_attribution.fit_hit_attribution(frame: pandas.DataFrame, *, target_gene: str, feature_columns: Sequence[str] | None = None, include_original_score: bool = False, threshold: float = 0.8, split_by: str = 'auto', random_seed: int = 0, n_bootstrap: int = 1000, n_permutations: int = 1000, n_pipeline_permutations: int = 0, source_regression_run: str = '') → HitAttributionResult[source]

Estimate cross-fitted target-hit-like probabilities.

frame must be the output of build_hit_cell_frame(). Every cell is predicted by a mixture fitted without its well, or without its plate when at least three plates make that split identifiable.

Parameters:
  • frame – ranked cell frame produced by build_hit_cell_frame().

  • target_gene – gene name to attach to the attribution result.

spacr.hit_attribution.promote_calls(db_path: str, attribution_run_id: str, annotation_column: str) → str[source]

Promote stored calls while recording every previous annotation value.

Parameters:
  • db_path – path to the SQLite measurements database to update.

  • attribution_run_id – stored investigation whose calls to promote.

  • annotation_column – png_list column to create or update.

spacr.hit_attribution.promote_hit_calls(db_path: str, result: HitAttributionResult, *, run_id: str, annotation_column: str, positive_value: Any = 1) → str[source]

Explicitly promote positive calls to a fresh, reversible annotation.

Parameters:
  • db_path – path to the SQLite measurements database to update.

  • result – attribution result containing the positive object calls.

  • run_id – attribution-run identifier to record in the audit.

  • annotation_column – fresh png_list column to receive the calls.

spacr.hit_attribution.quantify_candidate_enrichment(scored: pandas.DataFrame, *, target_column: str = 'target_well', bootstrap_iterations: int = 1000, permutation_iterations: int = 1000, random_seed: int = 0) → Tuple[pandas.DataFrame, Dict[str, Any]][source]

Quantify candidate prevalence at the well experimental unit.

Parameters:

scored – cell-level frame containing cross-fitted candidate probabilities and well identifiers.

spacr.hit_attribution.quantify_hit_enrichment(wells: pandas.DataFrame, *, random_seed: int = 0, n_bootstrap: int = 1000, n_permutations: int = 1000) → Dict[str, Any][source]

Per-well enrichment, well bootstrap CI, and a within-plate null.

Parameters:

wells – well-level frame containing target-guide fractions, hit-like prevalence, and mean hit-like probabilities.

spacr.hit_attribution.revert_promotion(db_path: str, promotion_id: str) → int[source]

Restore exactly the values replaced by promote_calls().

Reverting something that was never promoted is a NO-OP returning 0, not an error – including on a database where nothing has ever been promoted at all. The audit table is created by promote_calls(), so until one has run it does not exist, and reaching for it raised a raw sqlite3.OperationalError: no such table out of a spaCR API. An undo that crashes when there is nothing to undo makes a stray click look like a corrupt database.

Parameters:
  • db_path – path to the SQLite measurements database to update.

  • promotion_id – identifier returned by promote_calls().

spacr.hit_attribution.store_attribution(db_path: str, result: HitInvestigationResult) → int[source]

Store the GUI investigation bundle under its immutable run context.

Parameters:
  • db_path – path to the existing SQLite measurements database.

  • result – investigation bundle to persist.

spacr.hit_attribution.undo_hit_promotion(db_path: str, promotion_id: str) → int[source]

Clear values written by one promotion; preserve the audit record.

Parameters:
  • db_path – path to the SQLite measurements database to update.

  • promotion_id – identifier returned by promote_hit_calls().

spacr.hit_attribution.write_hit_attribution(db_path: str, result: HitAttributionResult, *, run_id: str | None = None) → str[source]

Persist a versioned attribution without touching annotation columns.

Parameters:
  • db_path – path to the existing SQLite measurements database.

  • result – attribution result whose run metadata and cell scores will be stored.