spacr.figures.stats¶
Pick the right test from the data, and show the working.
The choice is mechanical, so the software makes it:
groups variance distribution test 2 equal ~normal Student’s t (two-sided) 2 unequal ~normal Welch’s t 2 any not normal Mann-Whitney U >2 equal ~normal one-way ANOVA >2 unequal ~normal Welch’s ANOVA >2 any not normal Kruskal-Wallis
TWO THINGS THAT MAKE THIS EASY TO GET QUIETLY WRONG, and neither raises an error when you get it wrong – they hand back a confident number instead.
THE ASSUMPTION TESTS ARE THEMSELVES TESTS. On n = 3 Levene has almost no
power, so “p = 0.7, variances are equal” actually means “we could not tell”.
Reading that as licence to use Student’s t is how a screen reports a
difference that is not there. Below MIN_N_FOR_ASSUMPTIONS this module
records the check as UNINFORMATIVE and selects Welch’s t-test or
Mann-Whitney, which costs a little power when the assumption did hold and
protects the result when it did not. That asymmetry is the whole argument:
one direction loses a bit of sensitivity, the other publishes a false
positive.
THE UNIT OF REPLICATION. spaCR measures thousands of cells across a
handful of wells. A test across CELLS when the replicate is the WELL is
pseudoreplication and will return p < 1e-10 on pure noise, because n is
inflated by a factor of a thousand. Every result here states what n counted,
and compare() takes a unit so a caller can aggregate first.
A p-value alone is not reportable. Every result carries the test by name, n per group, an effect size, and the assumption checks with their own numbers.
THIS IS THE ONE ENGINE THAT CHOOSES A TEST. spacr.sp_stats used to
choose its own and disagreed with this one on three of five inputs, always by
taking the parametric branch where the checks had no power to refuse it. It is
now a translation layer onto compare(),
check_normality() and check_equal_variance() that keeps its older
signatures and result keys. Change the choices here and both entry points move
together; tests/test_one_engine_decides_which_test_applies.py fails if they
ever come apart again.
Classes¶
One assumption check, and whether it could see anything. |
|
One test, everything needed to report it, and how it was chosen. |
Functions¶
|
Levene, MEDIAN-centred. |
|
Shapiro-Wilk per group, and whether it could see anything. |
|
Choose and run the right test for these groups. |
|
The asterisks for a p-value, or |
|
Every comparison as one frame, corrected across them. |
Module Contents¶
- class spacr.figures.stats.Assumption[source]¶
One assumption check, and whether it could see anything.
- Parameters:
name – name of the assumption test.
statistic – test statistic, or
nanwhen it could not be computed.p_value – test p-value, or
nanwhen it could not be computed.informative – whether the check had enough usable data to interpret.
verdict – plain-language conclusion, including inconclusive cases.
passed – decision made by the check’s own rule; callers must not re-derive it from
p_value.
- class spacr.figures.stats.Comparison[source]¶
One test, everything needed to report it, and how it was chosen.
- Variables:
test – name of the selected statistical test.
statistic – statistic returned by that test.
p_value – unadjusted p-value returned by that test.
groups – group labels in the order tested.
n – usable observation counts for those groups, in matching order.
unit – independent unit represented by one observation, such as a well, cell, or guide; this prevents reporting row count as replication.
effect_size – estimated magnitude of the group difference on the scale named by
effect_name.effect_name – statistic used for
effect_size, such as Cohen’s d.ci – lower and upper confidence bounds for the reported effect, or
Nonewhen the selected method cannot provide them.assumptions – diagnostic checks that selected or qualified this test.
reason – plain-language explanation of why this test was selected.
correction – multiple-testing method applied to obtain
p_adjusted.p_adjusted – corrected p-value, or
nanwhen no correction applies.
- property marks: str[source]¶
The significance stars for this comparison.
FROM THE ADJUSTED P WHEN THERE IS ONE. Starring an unadjusted p in a figure that made many comparisons is how a multiple-testing problem becomes a claim; the raw value is used only when no adjustment was made.
- Returns:
the stars, or an empty string.
- spacr.figures.stats.check_equal_variance(groups: Sequence[numpy.ndarray]) Assumption[source]¶
Levene, MEDIAN-centred.
The median-centred Brown-Forsythe form is less sensitive to non-normal data than the mean-centred form. This function is called before the normality verdict is known.
- Parameters:
groups – numeric sample array for each group being compared.
- spacr.figures.stats.check_normality(groups: Sequence[numpy.ndarray]) Assumption[source]¶
Shapiro-Wilk per group, and whether it could see anything.
- Parameters:
groups – numeric sample array for each group being compared.
- spacr.figures.stats.compare(groups: Mapping[str, Sequence], *, unit: str = 'observation', paired: bool = False, force: str | None = None) Comparison[source]¶
Choose and run the right test for these groups.
- Parameters:
groups –
{label: values}. Two or more.unit – what ONE observation is – ‘well’, ‘cell’, ‘guide’. Stated in the result, because a test across cells when the replicate is the well is pseudoreplication and returns p < 1e-10 on noise.
paired – the groups are matched (the same wells before and after).
force – a test name to use instead of the chosen one.
- Returns:
a
Comparison.- Raises:
ValueError – with fewer than two groups, or a group too small to test. Refused rather than returned as NaN: a comparison that could not be made is not a comparison with an unknown answer.
- spacr.figures.stats.stars(p) str[source]¶
The asterisks for a p-value, or
n.s.written out.Non-significant comparisons are SHOWN rather than omitted, which is what the published figures do – a missing bracket reads as a comparison nobody made.
- Parameters:
p – p-value to translate into the reporting convention.
- spacr.figures.stats.table(comparisons: Sequence[Comparison], *, correction: str = 'fdr_bh')[source]¶
Every comparison as one frame, corrected across them.
Correcting ACROSS the comparisons is the part a hand-written stats table always forgets: six pairwise tests at 0.05 is a 26% chance of at least one false positive, and the individual p-values give no hint of it.
- Parameters:
comparisons – completed comparison results to tabulate and correct as one family.