spacr.guide_concordance

Does a gene’s signal come from its guides agreeing, or from one guide?

WHY THIS EXISTS

Ranking a screen by gene-level p-value silently mixes two different things:

  • a gene whose guides all move the same way, none of them individually striking – the signature of a real, moderate effect; and

  • a gene with ONE surviving guide, whose “gene-level” term is arithmetically identical to that guide’s own term and carries no independent evidence.

On the TSG101 screen the second kind sits at the top of the list. Gene 244480 has three guides in the library, exactly one survives the read-fraction filter, and its gene p-value (1.6e-12) IS that guide’s p-value – ranked above EAF1 and GRA14. EAF1 is the opposite case: three guides fitted, not one of them significant on its own (p = 0.51, 0.14, 0.27), and a gene-level p of 4.6e-08.

A hit list that does not separate these is not wrong so much as unreadable, and the distinction is invisible in the volcano because both are one dot.

Functions

annotate_results(→ pandas.DataFrame)

results with the guide-support columns joined on.

concordance_report(→ str)

A few lines a human can read, for the console after a run.

flag_single_guide_hits(→ pandas.DataFrame)

Hits whose evidence is one guide. Ranked as the table ranks them.

guide_support(→ pandas.DataFrame)

One row per gene: how many guides back it, and how well they agree.

Module Contents

spacr.guide_concordance.annotate_results(results: pandas.DataFrame, alpha: float = 0.05) → pandas.DataFrame[source]

results with the guide-support columns joined on.

Parameters:

results – result rows whose features are to be annotated with per-gene guide support.

Returned as a copy: a diagnostic must not quietly rewrite the table the caller is about to save.

spacr.guide_concordance.concordance_report(results: pandas.DataFrame, alpha: float = 0.05, top: int = 15, controls: dict | None = None) → str[source]

A few lines a human can read, for the console after a run.

Parameters:
  • results – guide-level coefficient table accepted by guide_support(), including feature names, effects, and P values.

  • controls – {gene_id: role}, e.g. {"239740": "positive", "233460": "negative"}. A negative control appearing in the hit list is the most useful line in the report and is easy to miss when it is just another six-digit number.

spacr.guide_concordance.flag_single_guide_hits(results: pandas.DataFrame, alpha: float = 0.05, p_column: str = 'p_value') → pandas.DataFrame[source]

Hits whose evidence is one guide. Ranked as the table ranks them.

Parameters:

results – result rows from which per-gene guide support is computed.

These are not necessarily false – a single guide can be the only one that cut – but they are a different claim from a gene whose guides agree, and a hit list that presents them identically invites the reader to treat them the same.

spacr.guide_concordance.guide_support(results: pandas.DataFrame, alpha: float = 0.05) → pandas.DataFrame[source]

One row per gene: how many guides back it, and how well they agree.

Parameters:
  • results – a regression coefficient table with feature, coefficient and p_value.

  • alpha – what counts as an individually significant guide.

Returns:

a frame indexed by gene with

n_guides

guide terms actually fitted, i.e. guides that survived filtration.

n_guides_significant

how many reached alpha on their own.

n_same_direction

how many share the sign of the gene’s mean effect. Guides that disagree in DIRECTION are the strongest argument that a hit is noise, and no p-value threshold reveals that.

concordance

n_same_direction / n_guides. 1.0 means every guide points the same way.

single_guide

True when one guide carries the whole gene. Such a gene’s gene-level p is that guide’s p and is not independent evidence.

gene_p

the gene-level term’s p-value, when there is one.