spacr.measurement_scan

Scan measurements for gene-associated effects.

The scan fits the same guide-level model to each measurement and reports both within-measurement and across-scan multiplicity corrections. A Simes global- null p value summarizes each measurement before the across-scan correction; the default correction is Benjamini-Hochberg.

Effect sizes are coefficients divided by the model’s residual standard deviation. ScanResult.effective_n_tests reports the Li and Ji estimate for correlated measurements, which is used when across_scan_method='bonferroni_effective'.

Exceptions

ScanRefused

The scan would not have meant what its output claimed.

Classes

MeasurementEffect

One scanned measurement: its best gene, and both corrections.

ScanResult

Every measurement the scan looked at, and how it was corrected.

Functions

effective_number_of_tests(→ float)

Li & Ji (2005) effective number of independent tests.

scan_measurements(, within_run_method, ...)

Fit every measurement against the guides and rank them by effect size.

simes_p_value(→ float)

Simes' global-null P value for one family of tests.

Module Contents

exception spacr.measurement_scan.ScanRefused[source]

Bases: ValueError

The scan would not have meant what its output claimed.

Initialize self. See help(type(self)) for accurate signature.

class spacr.measurement_scan.MeasurementEffect[source]

One scanned measurement: its best gene, and both corrections.

Parameters:
  • measurement – Name of the numeric response column that was scanned.

  • n_wells – Number of wells with finite values used for this measurement’s fit.

  • n_genes – Number of non-baseline gene terms in this measurement’s fitted design.

  • top_gene – Gene term with the largest absolute standardized effect.

  • effect_size – Signed coefficient of top_gene divided by this measurement’s residual standard deviation.

  • coefficient – Signed top_gene coefficient in the measurement’s original units.

  • p_value – Raw two-sided P value for top_gene.

  • within_run_q – Smallest adjusted P value across all gene terms for this measurement; it need not belong to top_gene.

  • within_run_hits – Number of gene terms rejected by the within-measurement correction at the requested alpha.

  • measurement_p – Simes global-null P value combining this measurement’s raw gene-term P values for the across-scan family.

  • across_scan_q – This measurement’s value after correction across all successfully scanned measurements.

  • survives_within_run – Whether at least one gene term survived the within-measurement correction.

  • survives_across_scan – Whether this measurement survived the across-scan correction.

class spacr.measurement_scan.ScanResult[source]

Every measurement the scan looked at, and how it was corrected.

Parameters:
  • rows – Successfully fitted measurement effects; measurements that could not be fitted are recorded in skipped instead.

  • skipped – Mapping from each unscanned measurement name to its refusal or fitting reason.

  • block_columns – Blocking columns actually retained in the fitted design after absent and single-valued candidates were removed.

  • gene_column – Input-frame column containing the gene assignment.

  • control_genes – Requested control-gene labels recorded as the effect-baseline request; absent labels are retained here even though the design falls back to an observed baseline.

  • within_run_method – Canonical multiplicity method applied across gene terms separately within each measurement.

  • across_scan_method – Canonical multiplicity method applied across measurements’ Simes global-null P values, including "bonferroni_effective" for the Li–Ji divisor.

  • alpha – Significance level used for both within-measurement and across-scan decisions.

  • effective_n_tests – Li–Ji estimate computed from successfully scanned measurement columns, or nan when no estimate is available.

  • genes_dropped – Mapping from each non-control gene removed for insufficient well replication to its observed well count.

describe() → str[source]

What a user has to be told before they read the table.

frame()[source]

The result table, ranked by effect size, largest first.

surviving() → Tuple[MeasurementEffect, ...][source]

Return measurements significant after across-scan correction.

property n_measurements_scanned: int[source]

Return the number of measurements fitted successfully.

spacr.measurement_scan.effective_number_of_tests(matrix) → float[source]

Li & Ji (2005) effective number of independent tests.

M_eff = sum_i f(|lambda_i|) over the eigenvalues of the correlation matrix, with f(x) = I(x >= 1) + (x - floor(x)): an eigenvalue of 3 counts as one independent test, an eigenvalue of 0 counts as none.

Why it is reported. Twelve columns that are four things measured three ways are not twelve tests, and correcting as though they were is the mirror image of not correcting at all – it hides real effects instead of inventing them. This number is what lets a reader see which situation they are in, and it is the divisor 'bonferroni_effective' uses.

Parameters:

matrix – (n_wells, n_measurements) array. Columns with no variance are dropped; NaNs are handled pairwise-complete.

Returns:

the estimate, or nan when it cannot be computed.

spacr.measurement_scan.scan_measurements(frame, *, measurements: Iterable[str] | None = None, gene_column: str = 'gene', guide_column: str | None = 'grna', block_columns: Sequence[str] = DEFAULT_BLOCK_COLUMNS, control_genes: Sequence[str] = (), within_run_method: str = 'fdr_bh', across_scan_method: str = 'fdr_bh', min_wells_per_gene: int = MIN_WELLS_PER_GENE, alpha: float = 0.05) → ScanResult[source]

Fit every measurement against the guides and rank them by effect size.

Parameters:
  • frame – one row per well, carrying the gene assignment, any blocking factors, and the measurement columns to scan.

  • measurements – the columns to scan. Default: every numeric column that is not identity or design (see _DESIGN_COLUMNS).

  • gene_column – the independent variable.

  • guide_column – excluded from the scanned columns when present.

  • block_columns – fixed effects fitted alongside the genes and not reported. Default DEFAULT_BLOCK_COLUMNS – the screen. Absent or single-valued blocks are dropped.

  • control_genes – the baseline the effects are measured from.

  • within_run_method – correction across the genes of ONE measurement, i.e. what a single-measurement analysis would have reported.

  • across_scan_method – correction across the measurements. Any spacr.multiple_testing method, plus 'bonferroni_effective' – Bonferroni divided by effective_number_of_tests() rather than by the column count.

  • min_wells_per_gene – genes with fewer wells than this are dropped before fitting, and named in ScanResult.genes_dropped. See MIN_WELLS_PER_GENE – a gene in one well has nothing corroborating it, and on real data keeping them made the across-scan correction fire on 65% of scans over permuted labels instead of 5%. Pass 1 to keep them, knowing that.

  • alpha – the level both corrections target.

Returns:

the ScanResult.

Raises:

ScanRefused – when there is no gene column, fewer than two genes, or nothing numeric to scan.

spacr.measurement_scan.simes_p_value(p_values) → float[source]

Simes’ global-null P value for one family of tests.

min_i (m * p_(i) / i) over the sorted P values – the probability of seeing a family this extreme when none of its members is real. Identical to the minimum Benjamini-Hochberg adjusted P, and valid under the positive dependence that gene terms in one design actually have.

This is what one measurement contributes to the across-scan stage. The raw minimum of 23 gene tests is not a P value and would smuggle the within-measurement multiplicity past the second correction.

Parameters:

p_values – raw P values for one family of tests; non-finite values are ignored.