spacr.figures.stats

Pick the right test from the data, and show the working.

The choice is mechanical, so the software makes it:

groups variance distribution test 2 equal ~normal Student’s t (two-sided) 2 unequal ~normal Welch’s t 2 any not normal Mann-Whitney U >2 equal ~normal one-way ANOVA >2 unequal ~normal Welch’s ANOVA >2 any not normal Kruskal-Wallis

TWO THINGS THAT MAKE THIS EASY TO GET QUIETLY WRONG, and neither raises an error when you get it wrong – they hand back a confident number instead.

THE ASSUMPTION TESTS ARE THEMSELVES TESTS. On n = 3 Levene has almost no power, so “p = 0.7, variances are equal” actually means “we could not tell”. Reading that as licence to use Student’s t is how a screen reports a difference that is not there. Below MIN_N_FOR_ASSUMPTIONS this module records the check as UNINFORMATIVE and selects Welch’s t-test or Mann-Whitney, which costs a little power when the assumption did hold and protects the result when it did not. That asymmetry is the whole argument: one direction loses a bit of sensitivity, the other publishes a false positive.

THE UNIT OF REPLICATION. spaCR measures thousands of cells across a handful of wells. A test across CELLS when the replicate is the WELL is pseudoreplication and will return p < 1e-10 on pure noise, because n is inflated by a factor of a thousand. Every result here states what n counted, and compare() takes a unit so a caller can aggregate first.

A p-value alone is not reportable. Every result carries the test by name, n per group, an effect size, and the assumption checks with their own numbers.

THIS IS THE ONE ENGINE THAT CHOOSES A TEST. spacr.sp_stats used to choose its own and disagreed with this one on three of five inputs, always by taking the parametric branch where the checks had no power to refuse it. It is now a translation layer onto compare(), check_normality() and check_equal_variance() that keeps its older signatures and result keys. Change the choices here and both entry points move together; tests/test_one_engine_decides_which_test_applies.py fails if they ever come apart again.

Classes

Assumption

One assumption check, and whether it could see anything.

Comparison

One test, everything needed to report it, and how it was chosen.

Functions

check_equal_variance(→ Assumption)

Levene, MEDIAN-centred.

check_normality(→ Assumption)

Shapiro-Wilk per group, and whether it could see anything.

compare(→ Comparison)

Choose and run the right test for these groups.

stars(→ str)

The asterisks for a p-value, or n.s. written out.

table(comparisons, *[, correction])

Every comparison as one frame, corrected across them.

Module Contents

class spacr.figures.stats.Assumption[source]

One assumption check, and whether it could see anything.

Parameters:
  • name – name of the assumption test.

  • statistic – test statistic, or nan when it could not be computed.

  • p_value – test p-value, or nan when it could not be computed.

  • informative – whether the check had enough usable data to interpret.

  • verdict – plain-language conclusion, including inconclusive cases.

  • passed – decision made by the check’s own rule; callers must not re-derive it from p_value.

class spacr.figures.stats.Comparison[source]

One test, everything needed to report it, and how it was chosen.

Variables:
  • test – name of the selected statistical test.

  • statistic – statistic returned by that test.

  • p_value – unadjusted p-value returned by that test.

  • groups – group labels in the order tested.

  • n – usable observation counts for those groups, in matching order.

  • unit – independent unit represented by one observation, such as a well, cell, or guide; this prevents reporting row count as replication.

  • effect_size – estimated magnitude of the group difference on the scale named by effect_name.

  • effect_name – statistic used for effect_size, such as Cohen’s d.

  • ci – lower and upper confidence bounds for the reported effect, or None when the selected method cannot provide them.

  • assumptions – diagnostic checks that selected or qualified this test.

  • reason – plain-language explanation of why this test was selected.

  • correction – multiple-testing method applied to obtain p_adjusted.

  • p_adjusted – corrected p-value, or nan when no correction applies.

sentence() → str[source]

The legend line: test, n, convention. Never a bare p.

property marks: str[source]

The significance stars for this comparison.

FROM THE ADJUSTED P WHEN THERE IS ONE. Starring an unadjusted p in a figure that made many comparisons is how a multiple-testing problem becomes a claim; the raw value is used only when no adjustment was made.

Returns:

the stars, or an empty string.

spacr.figures.stats.check_equal_variance(groups: Sequence[numpy.ndarray]) → Assumption[source]

Levene, MEDIAN-centred.

The median-centred Brown-Forsythe form is less sensitive to non-normal data than the mean-centred form. This function is called before the normality verdict is known.

Parameters:

groups – numeric sample array for each group being compared.

spacr.figures.stats.check_normality(groups: Sequence[numpy.ndarray]) → Assumption[source]

Shapiro-Wilk per group, and whether it could see anything.

Parameters:

groups – numeric sample array for each group being compared.

spacr.figures.stats.compare(groups: Mapping[str, Sequence], *, unit: str = 'observation', paired: bool = False, force: str | None = None) → Comparison[source]

Choose and run the right test for these groups.

Parameters:
  • groups – {label: values}. Two or more.

  • unit – what ONE observation is – ‘well’, ‘cell’, ‘guide’. Stated in the result, because a test across cells when the replicate is the well is pseudoreplication and returns p < 1e-10 on noise.

  • paired – the groups are matched (the same wells before and after).

  • force – a test name to use instead of the chosen one.

Returns:

a Comparison.

Raises:

ValueError – with fewer than two groups, or a group too small to test. Refused rather than returned as NaN: a comparison that could not be made is not a comparison with an unknown answer.

spacr.figures.stats.stars(p) → str[source]

The asterisks for a p-value, or n.s. written out.

Non-significant comparisons are SHOWN rather than omitted, which is what the published figures do – a missing bracket reads as a comparison nobody made.

Parameters:

p – p-value to translate into the reporting convention.

spacr.figures.stats.table(comparisons: Sequence[Comparison], *, correction: str = 'fdr_bh')[source]

Every comparison as one frame, corrected across them.

Correcting ACROSS the comparisons is the part a hand-written stats table always forgets: six pairwise tests at 0.05 is a 26% chance of at least one false positive, and the individual p-values give no hint of it.

Parameters:

comparisons – completed comparison results to tabulate and correct as one family.