spacr.measurement_scan¶
Scan measurements for gene-associated effects.
The scan fits the same guide-level model to each measurement and reports both within-measurement and across-scan multiplicity corrections. A Simes global- null p value summarizes each measurement before the across-scan correction; the default correction is Benjamini-Hochberg.
Effect sizes are coefficients divided by the model’s residual standard
deviation. ScanResult.effective_n_tests reports the Li and Ji estimate
for correlated measurements, which is used when
across_scan_method='bonferroni_effective'.
Exceptions¶
The scan would not have meant what its output claimed. |
Classes¶
One scanned measurement: its best gene, and both corrections. |
|
Every measurement the scan looked at, and how it was corrected. |
Functions¶
|
Li & Ji (2005) effective number of independent tests. |
|
Fit every measurement against the guides and rank them by effect size. |
|
Simes' global-null P value for one family of tests. |
Module Contents¶
- exception spacr.measurement_scan.ScanRefused[source]¶
Bases:
ValueErrorThe scan would not have meant what its output claimed.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.measurement_scan.MeasurementEffect[source]¶
One scanned measurement: its best gene, and both corrections.
- Parameters:
measurement – Name of the numeric response column that was scanned.
n_wells – Number of wells with finite values used for this measurement’s fit.
n_genes – Number of non-baseline gene terms in this measurement’s fitted design.
top_gene – Gene term with the largest absolute standardized effect.
effect_size – Signed coefficient of
top_genedivided by this measurement’s residual standard deviation.coefficient – Signed
top_genecoefficient in the measurement’s original units.p_value – Raw two-sided P value for
top_gene.within_run_q – Smallest adjusted P value across all gene terms for this measurement; it need not belong to
top_gene.within_run_hits – Number of gene terms rejected by the within-measurement correction at the requested alpha.
measurement_p – Simes global-null P value combining this measurement’s raw gene-term P values for the across-scan family.
across_scan_q – This measurement’s value after correction across all successfully scanned measurements.
survives_within_run – Whether at least one gene term survived the within-measurement correction.
survives_across_scan – Whether this measurement survived the across-scan correction.
- class spacr.measurement_scan.ScanResult[source]¶
Every measurement the scan looked at, and how it was corrected.
- Parameters:
rows – Successfully fitted measurement effects; measurements that could not be fitted are recorded in
skippedinstead.skipped – Mapping from each unscanned measurement name to its refusal or fitting reason.
block_columns – Blocking columns actually retained in the fitted design after absent and single-valued candidates were removed.
gene_column – Input-frame column containing the gene assignment.
control_genes – Requested control-gene labels recorded as the effect-baseline request; absent labels are retained here even though the design falls back to an observed baseline.
within_run_method – Canonical multiplicity method applied across gene terms separately within each measurement.
across_scan_method – Canonical multiplicity method applied across measurements’ Simes global-null P values, including
"bonferroni_effective"for the Li–Ji divisor.alpha – Significance level used for both within-measurement and across-scan decisions.
effective_n_tests – Li–Ji estimate computed from successfully scanned measurement columns, or
nanwhen no estimate is available.genes_dropped – Mapping from each non-control gene removed for insufficient well replication to its observed well count.
- surviving() Tuple[MeasurementEffect, ...][source]¶
Return measurements significant after across-scan correction.
- spacr.measurement_scan.effective_number_of_tests(matrix) float[source]¶
Li & Ji (2005) effective number of independent tests.
M_eff = sum_i f(|lambda_i|)over the eigenvalues of the correlation matrix, withf(x) = I(x >= 1) + (x - floor(x)): an eigenvalue of 3 counts as one independent test, an eigenvalue of 0 counts as none.Why it is reported. Twelve columns that are four things measured three ways are not twelve tests, and correcting as though they were is the mirror image of not correcting at all – it hides real effects instead of inventing them. This number is what lets a reader see which situation they are in, and it is the divisor
'bonferroni_effective'uses.- Parameters:
matrix –
(n_wells, n_measurements)array. Columns with no variance are dropped; NaNs are handled pairwise-complete.- Returns:
the estimate, or
nanwhen it cannot be computed.
- spacr.measurement_scan.scan_measurements(frame, *, measurements: Iterable[str] | None = None, gene_column: str = 'gene', guide_column: str | None = 'grna', block_columns: Sequence[str] = DEFAULT_BLOCK_COLUMNS, control_genes: Sequence[str] = (), within_run_method: str = 'fdr_bh', across_scan_method: str = 'fdr_bh', min_wells_per_gene: int = MIN_WELLS_PER_GENE, alpha: float = 0.05) ScanResult[source]¶
Fit every measurement against the guides and rank them by effect size.
- Parameters:
frame – one row per well, carrying the gene assignment, any blocking factors, and the measurement columns to scan.
measurements – the columns to scan. Default: every numeric column that is not identity or design (see
_DESIGN_COLUMNS).gene_column – the independent variable.
guide_column – excluded from the scanned columns when present.
block_columns – fixed effects fitted alongside the genes and not reported. Default
DEFAULT_BLOCK_COLUMNS– the screen. Absent or single-valued blocks are dropped.control_genes – the baseline the effects are measured from.
within_run_method – correction across the genes of ONE measurement, i.e. what a single-measurement analysis would have reported.
across_scan_method – correction across the measurements. Any
spacr.multiple_testingmethod, plus'bonferroni_effective'– Bonferroni divided byeffective_number_of_tests()rather than by the column count.min_wells_per_gene – genes with fewer wells than this are dropped before fitting, and named in
ScanResult.genes_dropped. SeeMIN_WELLS_PER_GENE– a gene in one well has nothing corroborating it, and on real data keeping them made the across-scan correction fire on 65% of scans over permuted labels instead of 5%. Pass1to keep them, knowing that.alpha – the level both corrections target.
- Returns:
the
ScanResult.- Raises:
ScanRefused – when there is no gene column, fewer than two genes, or nothing numeric to scan.
- spacr.measurement_scan.simes_p_value(p_values) float[source]¶
Simes’ global-null P value for one family of tests.
min_i (m * p_(i) / i)over the sorted P values – the probability of seeing a family this extreme when none of its members is real. Identical to the minimum Benjamini-Hochberg adjusted P, and valid under the positive dependence that gene terms in one design actually have.This is what one measurement contributes to the across-scan stage. The raw minimum of 23 gene tests is not a P value and would smuggle the within-measurement multiplicity past the second correction.
- Parameters:
p_values – raw P values for one family of tests; non-finite values are ignored.