spacr.guide_concordance¶
Does a gene’s signal come from its guides agreeing, or from one guide?
WHY THIS EXISTS
Ranking a screen by gene-level p-value silently mixes two different things:
a gene whose guides all move the same way, none of them individually striking – the signature of a real, moderate effect; and
a gene with ONE surviving guide, whose “gene-level” term is arithmetically identical to that guide’s own term and carries no independent evidence.
On the TSG101 screen the second kind sits at the top of the list. Gene 244480 has three guides in the library, exactly one survives the read-fraction filter, and its gene p-value (1.6e-12) IS that guide’s p-value – ranked above EAF1 and GRA14. EAF1 is the opposite case: three guides fitted, not one of them significant on its own (p = 0.51, 0.14, 0.27), and a gene-level p of 4.6e-08.
A hit list that does not separate these is not wrong so much as unreadable, and the distinction is invisible in the volcano because both are one dot.
Functions¶
|
|
|
A few lines a human can read, for the console after a run. |
|
Hits whose evidence is one guide. Ranked as the table ranks them. |
|
One row per gene: how many guides back it, and how well they agree. |
Module Contents¶
- spacr.guide_concordance.annotate_results(results: pandas.DataFrame, alpha: float = 0.05) pandas.DataFrame[source]¶
resultswith the guide-support columns joined on.- Parameters:
results – result rows whose features are to be annotated with per-gene guide support.
Returned as a copy: a diagnostic must not quietly rewrite the table the caller is about to save.
- spacr.guide_concordance.concordance_report(results: pandas.DataFrame, alpha: float = 0.05, top: int = 15, controls: dict | None = None) str[source]¶
A few lines a human can read, for the console after a run.
- Parameters:
results – guide-level coefficient table accepted by
guide_support(), including feature names, effects, and P values.controls –
{gene_id: role}, e.g.{"239740": "positive", "233460": "negative"}. A negative control appearing in the hit list is the most useful line in the report and is easy to miss when it is just another six-digit number.
- spacr.guide_concordance.flag_single_guide_hits(results: pandas.DataFrame, alpha: float = 0.05, p_column: str = 'p_value') pandas.DataFrame[source]¶
Hits whose evidence is one guide. Ranked as the table ranks them.
- Parameters:
results – result rows from which per-gene guide support is computed.
These are not necessarily false – a single guide can be the only one that cut – but they are a different claim from a gene whose guides agree, and a hit list that presents them identically invites the reader to treat them the same.
- spacr.guide_concordance.guide_support(results: pandas.DataFrame, alpha: float = 0.05) pandas.DataFrame[source]¶
One row per gene: how many guides back it, and how well they agree.
- Parameters:
results – a regression coefficient table with
feature,coefficientandp_value.alpha – what counts as an individually significant guide.
- Returns:
a frame indexed by gene with
n_guidesguide terms actually fitted, i.e. guides that survived filtration.
n_guides_significanthow many reached
alphaon their own.n_same_directionhow many share the sign of the gene’s mean effect. Guides that disagree in DIRECTION are the strongest argument that a hit is noise, and no p-value threshold reveals that.
concordancen_same_direction / n_guides. 1.0 means every guide points the same way.single_guideTrue when one guide carries the whole gene. Such a gene’s gene-level p is that guide’s p and is not independent evidence.
gene_pthe gene-level term’s p-value, when there is one.