spacr.batch_correction

Plate/batch-effect correction for tabular microscopy measurements.

The implementation is dependency-light (pandas/numpy), preserves the original row/index order, and never changes metadata columns. It is shared by Image UMAP, Classify (ML), UMAP hyperparameter search, and regression.

Attributes

METHODS

Supported correction methods.

NO_COVARIATE

Explicit declaration that no biological signal needs protecting.

Classes

BatchCorrectionReport

Diagnostics for one correction operation.

Functions

correct_batch_effects(→ Tuple[pandas.DataFrame, ...)

Normalize numeric features within acquisition batches.

correct_from_metadata(→ Tuple[pandas.DataFrame, ...)

Correct a feature frame using named columns from an aligned metadata frame.

correction_kwargs(→ Dict[str, Any])

Translate shared GUI settings into correction-call keyword arguments.

write_report(→ pathlib.Path)

Write a correction report as stable JSON and return its path.

Module Contents

class spacr.batch_correction.BatchCorrectionReport[source]

Diagnostics for one correction operation.

Parameters:
  • method – Normalized correction method requested for the operation; it remains recorded when a one-batch operation becomes a warned no-op.

  • batch_column – Human-readable metadata-column name used to identify the supplied batch labels.

  • batches – Sorted distinct batch labels after conversion to strings.

  • features – Numeric feature-column names returned by the operation, whether corrected or left unchanged by a no-op.

  • rows – Number of input feature rows considered.

  • controls – Total rows matching the reference controls used by control_center; zero for other methods.

  • centroid_spread_before – Mean across-feature standard deviation of batch centroids before correction, or None when unavailable.

  • centroid_spread_after – The same batch-centroid diagnostic after correction or a no-op, or None when unavailable.

  • covariate_columns – Source biological-covariate columns supplied to ComBat.

  • covariate_terms – Design-matrix terms expanded from those covariates when ComBat was fitted; empty when no fit was performed.

  • covariate_spread_before – Batch-centroid-style spread across categorical covariate groups before correction, or None for no or continuous covariates.

  • covariate_spread_after – The same categorical-covariate spread after correction or a no-op, or None when unavailable.

  • warnings – Explicit no-op, fallback, unchanged-batch, constant-feature, or ComBat limitation messages.

to_dict() → Dict[str, Any][source]

Return a JSON-serializable report.

spacr.batch_correction.correct_batch_effects(features: pandas.DataFrame, batch: pandas.Series, *, method: str = 'none', batch_column: str = 'plateID', control: pandas.Series | None = None, control_values: Any = None, covariate: Any = None, combat_mean_only: bool = False, combat_empirical_bayes: bool = True, min_samples: int = 3, missing_control: str = 'error') → Tuple[pandas.DataFrame, BatchCorrectionReport][source]

Normalize numeric features within acquisition batches.

center removes per-batch mean shifts while preserving the global mean. zscore aligns per-batch means and variances to the global distribution. robust_zscore does the same with median/MAD and is less sensitive to heavy-tailed single-cell measurements. control_center estimates only a location shift from negative/reference controls in every batch, preserving treatment dispersion and usually best preserving biology.

combat is the empirical-Bayes method of Johnson, Li & Rabinovich (2007). Unlike the four above it fits a model: every feature is regressed on batch indicators and on the biology named by covariate, and only the batch part of that fit is removed. The per-batch location and scale are then shrunk toward the distribution of the same parameter across all features, which is what makes it usable on a plate with few wells where a per-feature estimate would be noise.

That covariate is not optional and has no default. A batch effect estimated without it absorbs any contrast that happens to differ between plates – which, in a screen where treatments are laid out plate by plate, is the treatment effect. The correction then reports a cleaner batch diagnostic and a dead result. Pass the biology to keep, or pass NO_COVARIATE to record that there is none.

Parameters:
  • features – numeric feature DataFrame; metadata must not be included.

  • batch – batch/plate label aligned to features.index.

  • method – one of METHODS.

  • batch_column – human-readable source column for diagnostics.

  • control – optional aligned series used by control_center.

  • control_values – scalar/list values selecting reference controls.

  • covariate – required by combat and ignored by every other method – a Series or DataFrame of biology to preserve, aligned to features.index, or NO_COVARIATE to declare there is none.

  • combat_mean_only – correct only the additive batch shift and leave each batch’s dispersion alone. The right choice when plates differ in offset but the assay’s noise model is stable, and when rescaling a variance would manufacture significance.

  • combat_empirical_bayes – False uses raw per-feature batch estimates with no shrinkage. Only sensible with many rows per batch; it exists so a test can show the shrinkage is doing something.

  • min_samples – minimum rows (or controls) required per batch.

  • missing_control – "error" or "skip" for a batch lacking enough reference-control rows.

Returns:

(corrected_features, report).

Raises:

ValueError – for unknown methods, misaligned metadata, non-numeric features, missing batches, insufficient required controls, a missing combat covariate, or a covariate confounded with batch.

spacr.batch_correction.correct_from_metadata(features: pandas.DataFrame, metadata: pandas.DataFrame, *, batch_correction: str = 'none', batch_column: str = 'plateID', batch_control_column: str | None = None, batch_control_values: Any = None, batch_covariate_column: Any = None, batch_combat_mean_only: bool = False, batch_min_samples: int = 3, batch_missing_control: str = 'error') → Tuple[pandas.DataFrame, BatchCorrectionReport][source]

Correct a feature frame using named columns from an aligned metadata frame.

This adapter is the shared boundary used by UMAP, ML, and regression. It validates column names once and keeps metadata out of the numeric feature matrix.

Parameters:
  • features – numeric feature DataFrame.

  • metadata – DataFrame containing batch and optional control columns.

  • batch_correction – correction method accepted by correct_batch_effects().

  • batch_column – metadata column identifying batches.

  • batch_control_column – optional reference-control metadata column.

  • batch_control_values – values selecting reference controls.

  • batch_covariate_column – metadata column(s) naming the biology that combat must preserve – one name, a comma-separated string, or a list. The literal "none" declares that there is none. Required by combat and ignored by every other method; leaving it blank makes combat refuse to run rather than quietly delete the contrast.

  • batch_combat_mean_only – correct only the additive shift.

  • batch_min_samples – minimum samples/reference controls per batch.

  • batch_missing_control – "error" or "skip".

Returns:

corrected features and diagnostics report.

Raises:

ValueError – when metadata cannot be aligned or required columns are absent.

spacr.batch_correction.correction_kwargs(settings: Mapping[str, Any], *, default_control_column: str | None = None, default_control_values: Any = None) → Dict[str, Any][source]

Translate shared GUI settings into correction-call keyword arguments.

Emits exactly six keys – batch_correction, batch_column, batch_control_column, batch_control_values, batch_min_samples and batch_missing_control – with the same defaults as correct_from_metadata(). The combat-only keys batch_covariate_column and batch_combat_mean_only are deliberately left out so the result stays safe to splat into signatures that do not accept them; a caller using batch_correction="combat" must pass batch_covariate_column alongside this mapping or correct_from_metadata() raises. batch_combat_mean_only stays optional and defaults to False.

Parameters:
  • settings – settings mapping the batch keys are read from.

  • default_control_column – control column used when batch_control_column is absent or an empty string; a whitespace-only name is kept verbatim.

  • default_control_values – control values used when batch_control_values is absent or blank.

Returns:

keyword arguments for correct_from_metadata().

spacr.batch_correction.write_report(report: BatchCorrectionReport, path: Any) → pathlib.Path[source]

Write a correction report as stable JSON and return its path.

Parameters:
  • report – completed batch-correction report to serialize.

  • path – destination JSON path to replace atomically.

spacr.batch_correction.METHODS = ('none', 'center', 'zscore', 'robust_zscore', 'control_center', 'combat')[source]

Supported correction methods.

spacr.batch_correction.NO_COVARIATE = 'no_covariate'[source]

Explicit declaration that no biological signal needs protecting.

ComBat estimates the batch effect from whatever variation is left after the design matrix has absorbed the biology. If the biology is not in that design, it is part of “whatever is left” and gets removed along with the plate effect. That failure is silent: the corrected table looks cleaner, the batch diagnostic improves, and the treatment effect is gone.

So correct_batch_effects() refuses to run method="combat" until the caller has answered the question. Pass the covariate to keep, or pass this constant to state on the record that there is nothing to keep – which is only true when every batch holds the same mixture of conditions, or when the output feeds an unsupervised embedding with no contrast to protect.