spacr.multiple_testing¶
Multiple-testing corrections for spaCR, with one name for every method.
Every correction spaCR offers lives here, so the GUI dropdown, the CLI, the
settings validator and spacr.guide_permutation cannot drift apart: the
dropdown is built from METHODS, the validator checks against the same
table, and the analysis calls adjust_p_values().
Two properties matter for a pooled screen and are easy to get wrong:
Missing P values must not join the family. A guide that could not be tested is not a test. Counting it inflates the family size and makes every real discovery less significant. NaNs pass through untouched here.
The family is the displayed set. The caller decides what a family is (spaCR corrects separately per classifier outcome and per minimum-support threshold); this module only corrects whatever vector it is handed.
Beyond the statsmodels inventory this module implements Storey’s q-value, which is standard in screening work and absent from statsmodels. It estimates the proportion of true nulls (pi0) rather than assuming it is 1, so it is uniformly less conservative than Benjamini-Hochberg while targeting the same false-discovery rate.
Classes¶
One correction: its canonical key, label, family and controlled rate. |
Functions¶
|
Return |
|
Return the canonical key for |
|
The largest RAW P value |
|
Estimate the proportion of true null hypotheses (Storey's pi0). |
|
Estimate the local false-discovery rate for a family of p values. |
|
Canonical keys in dropdown order. |
|
Human-readable label for |
|
Return Storey q-values for a vector of P values. |
Module Contents¶
- class spacr.multiple_testing.MethodSpec[source]¶
One correction: its canonical key, label, family and controlled rate.
- Parameters:
key – canonical settings and command-line identifier for the correction.
label – human-readable method name shown in selectors.
controls – error rate controlled by the method:
"FDR","FWER", or"nothing".statsmodels_name – method name passed to
statsmodels, orNonefor locally implemented or absent correction.summary – concise explanation of the method and its assumptions.
- spacr.multiple_testing.adjust_p_values(p_values, method='fdr_bh', alpha=0.05)[source]¶
Return
(adjusted, rejected)for one multiple-testing family.- Parameters:
p_values – P values for every test in the family. Non-finite entries stay non-finite, are never rejected, and do not count toward the family size.
method – any key, statsmodels name or alias accepted by
canonical_method().alpha – the level the correction targets, strictly inside (0, 1).
- Returns:
adjusted(same shape as the input) andrejected, a boolean array.
- spacr.multiple_testing.canonical_method(method) str[source]¶
Return the canonical key for
method.Accepts the canonical keys, the statsmodels spellings and the common aliases.
Nonemaps to'none'. RaisesValueErrorwith the full inventory for anything else, rather than silently falling back to a correction the user did not ask for.- Parameters:
method – canonical key, statsmodels spelling, alias, or
None.
- spacr.multiple_testing.critical_p_value(p_values, method='fdr_bh', alpha=0.05)[source]¶
The largest RAW P value
methodcalls atalpha, orNone.THE NUMBER THAT LETS A CONTINUOUS AXIS SHOW A DISCRETE CALL. A volcano drawn against the raw P value cannot put the adjusted P on its y-axis without inheriting the step function BH’s cumulative minimum creates – but it does not have to. Every correction here is monotone in the raw P within a family, so the set it calls is always a lower set: there is a rank
ksuch that every test withp <= p_(k)is called and every test above it is not. One horizontal line at-log10(p_(k))therefore divides the plot EXACTLY as the correction does, on an axis with no steps anywhere.For Benjamini-Hochberg this is the textbook identity
q_(i) <= alpha if and only if p_(i) <= alpha * i / n
taken at the largest
ithat satisfies it. Computed here from the correction’s own rejection call rather than from that formula, so the line cannot disagree with the colours beside it and the same one function answers for all thirteen methods rather than for one.THE LINE IS NOT
alpha. Drawing it at-log10(0.05)is the mistake this replaces: that is the uncorrected threshold, it is far higher than the corrected one, and a reader measuring against it reads far too much of the screen as called.NoneWHEN NOTHING IS CALLED, and it is not a failure. There is then nokand no line, and saying so is a finding; drawing one atalphainstead would claim a threshold the procedure never reached.- Parameters:
p_values – the raw P values of ONE multiple-testing family.
method – any spelling
canonical_method()accepts.alpha – the level the correction targets.
- Returns:
the raw-P threshold as a float, or
None.
- spacr.multiple_testing.estimate_pi0(p_values, *, lambdas: Sequence[float] | None = None) float[source]¶
Estimate the proportion of true null hypotheses (Storey’s pi0).
Uses the smoothed bootstrap-free spline-free estimator: pi0(lambda) is computed over a grid and the estimate is taken at the largest lambda whose value is stable, then clipped to (0, 1]. With few tests the grid collapses and the estimator returns 1.0, which makes the q-values fall back to Benjamini-Hochberg – the conservative answer, not an error.
- Parameters:
p_values – raw P values; non-finite entries are excluded.
- spacr.multiple_testing.local_fdr(p_values, *, pi0: float | None = None)[source]¶
Estimate the local false-discovery rate for a family of p values.
- Parameters:
p_values (array-like) – Raw p values from one testing family.
pi0 (float, optional) – Proportion of true null hypotheses. When omitted, it is estimated as
w + (1 - w) * afrom the fitted beta-uniform mixture.
- Returns:
numpy.ndarray – Local false-discovery rates in
[0, 1]with the same shape as the input. NaNs are preserved and excluded from the fit.
Notes
The local FDR estimates the posterior probability that an individual test belongs to the null component:
lfdr(p) = pi0 * f0(p) / f(p), f0(p) = 1, p ∈ [0, 1]
fis a beta-uniform mixture whose alternative component isBeta(a, 1). Unlike a q value, which summarizes a tail of discoveries, the local FDR describes one test. Families smaller thanLOCAL_FDR_MIN_TESTSreturn1for every finite value because a density cannot be estimated reliably from so few observations.
- spacr.multiple_testing.method_label(method) str[source]¶
Human-readable label for
method.- Parameters:
method – any correction spelling accepted by
canonical_method().
- spacr.multiple_testing.storey_qvalue(p_values, *, pi0: float | None = None)[source]¶
Return Storey q-values for a vector of P values.
q[i]is the minimum positive-FDR at which testiis called significant. The result is monotone in the P value, as the definition requires. NaNs are preserved.- Parameters:
p_values – raw P-value array whose shape and NaNs are preserved.