spacr.permutation_qc

Evaluate residual exchangeability for blocked permutation tests.

Permutation inference assumes that residuals can be exchanged within each block. This module reports serial autocorrelation, positional gradients, and block-level diagnostics that may invalidate that assumption. These checks complement residual-versus-fitted and Q-Q plots: normality and constant variance do not establish exchangeability across plate positions.

Functions

autocorrelation(→ float)

Durbin-Watson on residuals in the order given.

block_residual_report(→ Dict[str, Any])

Everything the permutation's assumption needs, per block and overall.

exchangeability_verdict(→ Dict[str, Any])

Is the within-block shuffle defensible, and if not, what to change?

plot_residual_by_position(residuals, blocks, positions)

Draw the permuted residual against plate position, one row per block.

position_effect(→ Dict[str, float])

How much of the residual one position column explains.

write_permutation_qc() → Dict[str, Any])

Write the permutation run's QC into <destination>/regression_qc/.

Module Contents

spacr.permutation_qc.autocorrelation(residuals: Sequence[float]) → float[source]

Durbin-Watson on residuals in the order given.

Parameters:

residuals – ordered residual values to test for serial dependence.

Returns:

the Durbin-Watson statistic, or nan with fewer than two finite residuals or a nonpositive sum of squared residuals.

2 is no autocorrelation, 0 is perfect positive, 4 perfect negative. Written out rather than imported so this module does not pull statsmodels for one line – and so the ORDER is explicit: it is the order the rows arrive in, which for a well table is plate reading order.

spacr.permutation_qc.block_residual_report(residuals: Sequence[float], blocks: Sequence[Any], positions: Mapping[str, Sequence[Any]] | None = None) → Dict[str, Any][source]

Everything the permutation’s assumption needs, per block and overall.

Parameters:
  • residuals – the phenotype residuals that WILL BE PERMUTED – not a model’s residuals, which is the distinction this whole module exists for.

  • blocks – the block each residual belongs to; the shuffle happens inside these.

  • positions – {'rowID': [...], 'columnID': [...]} or similar.

Returns:

pooled sample and block counts, pooled and per-block Durbin-Watson diagnostics, per-block means and standard deviations, and one position_effect() result per named position column.

Raises:

ValueError – if blocks or any named position column does not contain exactly one value for each residual.

PER BLOCK AND NOT ONLY POOLED. The shuffle is within-block, so a pooled statistic can look healthy while one plate is badly structured – and that one plate is where the false positives come from.

spacr.permutation_qc.exchangeability_verdict(report: Mapping[str, Any]) → Dict[str, Any][source]

Is the within-block shuffle defensible, and if not, what to change?

Parameters:

report – diagnostics returned by block_residual_report().

Returns:

{'ok': bool, 'findings': [...], 'remedy': str}.

THE REMEDY IS NAMED, NOT IMPLIED. “Durbin-Watson 1.22” is a number; “add rowID to guide_nuisance_columns” is something the reader can do, and it is the same setting the run already has.

spacr.permutation_qc.plot_residual_by_position(residuals: Sequence[float], blocks: Sequence[Any], positions: Mapping[str, Sequence[Any]], report: Mapping[str, Any] | None = None, verdict: Mapping[str, Any] | None = None, removed: Sequence[str] = (), title: str = '')[source]

Draw the permuted residual against plate position, one row per block.

Parameters:
  • residuals – the residuals that WERE PERMUTED.

  • blocks – the block of each residual; the shuffle happens inside it.

  • positions – {'rowID': [...], 'columnID': [...]}, aligned to residuals.

  • report – block_residual_report() for the same residuals; when given, each block’s Durbin-Watson and each column’s pooled p-value are drawn on the panel they describe.

  • verdict – exchangeability_verdict() of report; its remedy is written under the panels, so the setting that removes a gradient is named where the gradient is seen.

  • removed – nuisance columns already removed before residualisation, named on the figure so a run with position removed can be told from one without.

  • title – the figure’s title, normally the outcome column.

Returns:

a matplotlib.figure.Figure.

Raises:

ValueError – if positions is empty or a column is misaligned.

A GRADIENT ACROSS THE PLATE IS WHAT THIS SHOWS AND NO OTHER PANEL DOES. Residuals-vs-fitted shows spread and Q-Q shows shape; neither shows that row A sits above row P inside one plate, which is exactly what makes the wells in that plate not swappable.

spacr.permutation_qc.position_effect(residuals: Sequence[float], positions: Sequence[Any]) → Dict[str, float][source]

How much of the residual one position column explains.

Parameters:
  • residuals – residual values whose positional structure is measured.

  • positions – position level aligned to each residual.

Returns:

a mapping containing eta_squared and levels. With at least three finite observations it also contains p_value; when two or more levels vary, it includes omega_squared, worst_level, and worst_departure, the signed departure of that level’s mean from the grand residual mean.

Raises:

ValueError – if residuals and positions do not contain exactly one value for each other.

THE NUMBER THAT MATTERS FOR EXCHANGEABILITY. If rows differ from each other more than cells within a row do, then shuffling across rows inside a plate is shuffling things that are not alike, and the permutation null is wider than the truth.

spacr.permutation_qc.write_permutation_qc(destination: str, outcome: str, residuals: Sequence[float], blocks: Sequence[Any], positions: Mapping[str, Sequence[Any]], report: Mapping[str, Any], verdict: Mapping[str, Any], removed: Sequence[str] = ()) → Dict[str, Any][source]

Write the permutation run’s QC into <destination>/regression_qc/.

The same folder a parametric run writes, so a permutation run’s diagnostics are found where every other run’s are.

Parameters:
  • destination – the run’s results folder.

  • outcome – the phenotype column the residuals belong to; it names the files, so a multi-outcome run keeps one set per outcome.

  • residuals – the residuals that WERE PERMUTED.

  • blocks – the block of each residual.

  • positions – the position columns, aligned to residuals.

  • report – block_residual_report() of the same residuals.

  • verdict – exchangeability_verdict() of report.

  • removed – nuisance columns removed before residualisation.

Returns:

{'dir', 'report', 'figure', 'figure_error'}. The JSON report is always written; figure is None and figure_error says why when there is nothing to draw against. The panel is written by spacr.plot.save_figure(), so its format and resolution follow the figure preferences as the parametric panels beside it do, and figure names the file on disk.

Nested helpers

plot_residual_by_position._distance(block)

How far a block’s Durbin-Watson sits from 2, or -1 without one.

Parameters:

block – a block label from labels.

Returns:

|DW - 2|, so the most autocorrelated blocks sort first; -1.0 when the report has no finite statistic.

spacr/permutation_qc.py:319