spacr.rra

Aggregate guide-level scores into gene-level calls using alpha-RRA.

RRA ranks guides and tests whether a gene’s guides occupy unusually strong positions. It therefore remains usable when a gene-fraction regression design is singular.

Statistic

Within each direction, rank every guide by score. A gene with k guides occupies normalized ranks r_1 <= ... <= r_k in (0, 1]. Under the global null, r_i follows Beta(i, k - i + 1). The statistic is

rho = min_i Beta(i, k - i + 1).cdf(r_i)

where the minimum is restricted to ranks in the top alpha fraction. Permutation p values are estimated by reassigning guides among genes with the same guide count. Depletion and enrichment are evaluated separately.

Functions

describe(→ str)

Return the formula and model guidance for the model tab's text box.

rank_aggregate(scores, groups, *[, alpha, direction, ...])

Aggregate per-guide scores to per-gene calls by rank.

Module Contents

spacr.rra.describe(alpha: float = DEFAULT_ALPHA) → str[source]

Return the formula and model guidance for the model tab’s text box.

Parameters:

alpha – top-rank cutoff to interpolate into the explanation; this formatter does not validate the supplied value.

Returns:

alpha-RRA formula, interpretation, and CRISPR-screen guidance.

spacr.rra.rank_aggregate(scores, groups, *, alpha: float = DEFAULT_ALPHA, direction: str = 'both', n_permutations: int = DEFAULT_PERMUTATIONS, seed: int = 0, correction: str = 'fdr_bh', fdr_alpha: float = 0.05)[source]

Aggregate per-guide scores to per-gene calls by rank.

Parameters:
  • scores – one score per guide – a coefficient, a log fold change, anything where more negative means more depleted.

  • groups – the gene each guide belongs to, same length.

  • alpha – the top fraction of the ranking to aggregate over.

  • direction – "neg", "pos" or "both".

  • n_permutations – draws per distinct guide count.

  • seed – the permutation seed, so a re-run of the same screen gives the same P values. A screen whose hit list moves between runs of the same data is a screen nobody can act on.

  • correction – any method spacr.multiple_testing accepts.

  • fdr_alpha – the level the correction targets.

Returns:

a DataFrame, one row per gene, sorted by the strongest direction’s P value.

Raises:

ValueError – scores and groups differ in length, alpha is outside (0, 1], or direction is not one of DIRECTIONS.

Guides with non-finite scores are omitted rather than ranked. This keeps an unestimated guide from being interpreted as strongly depleted.