spacr.parameter_sweep

Workflow inputs and outputs

Parameter Sweep

Run explicitly selected regression settings and compare their diagnostics. Keep the input data and evaluation question fixed.

Open: Regression → Parameter Sweep.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Object classification scores — Saved score CSVs and, when merged, measurements/measurements.db, table png_list. Relevant tables, depending on the route: png_list. Relevant columns, depending on the route: pred, cv_predictions, ml_pred, predictions.

  • Guide counts per well — Map Barcodes run folder: unique_combinations.csv and annotated_reads.h5. Well identity requires the corresponding barcode references. Relevant columns, depending on the route: count.

Outputs

  • Run/model comparison — Comparison tables and figures from compatible saved runs or masks; agreement is not ground-truth accuracy.

  • Regression results and hits — Selected run results folder: coefficient/result CSVs, hit tables, settings and diagnostic figures.

API reference.

Module tutorial.

Parameter Sweep: vary regression settings and see which change the answer.

A pooled screen has no single correct analysis. It has a model family, an aggregation rule, a unit of analysis, a set of nuisance effects, a multiple-testing correction, and three or four filtration cutoffs – and a hit list that is stable across all of them is a very different claim from one that appears under exactly one combination.

This module makes that comparison a command rather than a week. It builds a trial list from a search space, runs each trial into its own folder, and returns one tidy row per trial: the settings, whether it ran, how many wells and guides survived, how many hits it called, where the named controls landed, and how long it took.

Two things make it fast enough to be worth having:

  • Preparation is shared. Loading 226k score rows and 642k count rows, aggregating per well, thresholding and joining costs about twenty seconds and depends only on the FILTRATION settings. Every trial sharing those reuses one prepared frame, so a sweep over models and corrections pays that cost once per filtration cell rather than once per trial.

  • A failed trial is a result. Many combinations are illegal by construction – quantile refuses alpha, the penalised families refuse cov_type, random_row_column_effects replaces the backend entirely. Those are recorded with their reason and the sweep continues, because “this combination is not allowed” is information about the design space, not a crash.

The search space is declared as data (DEFAULT_SWEEP_SPACE), so adding an axis is adding a key.

Classes

SweepSpace

Define variable and fixed settings for a parameter sweep.

Functions

be_polite(→ None)

Lower the current worker's CPU, I/O, and OOM-kill priority.

build_trials(→ list[dict])

Enumerate accepted trials from a sweep space.

containment_available(→ bool)

Return whether a user scope accepts the trial memory limits.

containment_note(→ str)

Return a user-facing summary of trial resource containment.

correction_rows(→ list[dict])

One row per correction, from ONE fit.

free_memory_gb(→ float)

Memory the kernel says is actually available, in GB.

memory_is_low(→ bool)

True when the machine or spaCR's own process tree crossed its limit.

rank_trials(→ pandas.DataFrame)

Order trials by recovery of a named control role.

recommended_workers(*[, measured_gib, requested])

Choose a worker count from available memory and CPU capacity.

rerun_trial(→ dict)

Re-run one trial and hand back its settings, output and FIGURES.

run_sweep(→ pandas.DataFrame)

Run parameter-sweep trials sequentially.

run_sweep_parallel(→ pandas.DataFrame)

Run parameter-sweep trials concurrently in spawned processes.

run_trial_contained(→ dict)

Run one trial in a fresh child process and return a status row.

settings_for_trial(→ dict)

The full settings dict that produced row.

summarise_sweep() → dict)

Summarize sweep completion, controls, and hit-count sensitivity.

Module Contents

class spacr.parameter_sweep.SweepSpace[source]

Define variable and fixed settings for a parameter sweep.

Parameters:
  • axes (dict of str to list, optional) – Settings to vary. Each trial receives one value from every list. Defaults to a copy of DEFAULT_SWEEP_SPACE.

  • fixed (dict, optional) – Settings copied into every trial after the Cartesian product is formed and before filters are evaluated.

  • filters (list of callable, optional) – Predicates called with a complete trial dictionary. A predicate returns None to accept a trial or a reason string to reject it. An empty list selects the built-in compatibility filters when trials are built.

size() → int[source]

Return the raw Cartesian-product size before filtering.

spacr.parameter_sweep.be_polite() → None[source]

Lower the current worker’s CPU, I/O, and OOM-kill priority.

The adjustments are best-effort and platform dependent. On Linux, the worker uses the lowest CPU and I/O priorities and sets oom_score_adj=800 so it is preferred over interactive applications when the system is under memory pressure.

spacr.parameter_sweep.build_trials(space: SweepSpace, *, mode: str = 'grid', max_trials: int = 5000, seed: int = 0) → list[dict][source]

Enumerate accepted trials from a sweep space.

Parameters:
  • space (SweepSpace) – Variable axes, fixed settings, and optional rejection predicates.

  • mode ({'grid', 'random'}, default='grid') – 'grid' visits the Cartesian product in axis order. 'random' shuffles that product without replacement before applying the limit.

  • max_trials (int, default=5000) – Maximum number of accepted trials to return. Rejected combinations do not count toward this limit.

  • seed (int, default=0) – Seed used to shuffle combinations in 'random' mode. It has no effect in 'grid' mode.

Returns:

list of dict – Complete setting dictionaries with one-based trial_id values. Fixed settings are present when filters are evaluated.

Raises:

ValueError – If mode is not 'grid' or 'random'.

Notes

Random mode materializes the full Cartesian product before shuffling it. Use narrower axes when the unfiltered product is very large.

spacr.parameter_sweep.containment_available() → bool[source]

Return whether a user scope accepts the trial memory limits.

Finding systemd-run is insufficient: hosted runners and containers may ship the executable without a usable user manager or delegated memory controller. A small no-op scope verifies the properties used for a real trial without allocating meaningful memory.

spacr.parameter_sweep.containment_note() → str[source]

Return a user-facing summary of trial resource containment.

Returns:

str – Active kernel limits when containment is available, or a warning that explains the remaining safeguards and safer alternatives when it is unavailable.

spacr.parameter_sweep.correction_rows(output: Mapping[str, Any], methods: Sequence[str], alpha: float = 0.05) → list[dict][source]

One row per correction, from ONE fit.

A multiple-testing correction is applied to p-values that already exist; it does not change the model, the design, or a single coefficient. Sweeping it as an axis therefore refits the identical regression once per method – thirteen fits to obtain thirteen numbers that all come from the first one.

On this screen that is the difference between ~24 hours and ~2, for exactly the same answers, which is why it is worth doing rather than clever.

Returns:

[{'multiple_testing_method': m, 'n_below_alpha': n, ...}, ...]

spacr.parameter_sweep.free_memory_gb() → float[source]

Memory the kernel says is actually available, in GB.

MemAvailable, not MemFree: free memory on a working machine is close to zero because the page cache holds the rest, and scheduling a trial against that number would refuse every trial on a healthy box. MemAvailable is the kernel’s own estimate of what a new allocation could get without swapping.

Returns:

available memory in GB, or inf where the file does not exist. Infinity rather than zero deliberately – this is a safety check, and a check that cannot read the machine must not become a check that blocks every run on it.

spacr.parameter_sweep.memory_is_low(floor_gib: float = 8.0, spacr_ceiling_gib: float | None = None) → bool[source]

True when the machine or spaCR’s own process tree crossed its limit.

Checked between submissions, not only at the start: the other things on the machine – an editor, the spaCR GUI, another analysis – grow while the sweep runs, and the sweep must yield to them rather than race them.

spacr_ceiling_gib is optional because the machine-wide free-memory floor remains the default safety policy. When a caller supplies a run budget, the active resource sampler’s latest process-tree total makes the guard attributable: it can distinguish “the machine is busy” from “spaCR is the reason” instead of treating both as the same number.

spacr.parameter_sweep.rank_trials(results: pandas.DataFrame, *, role: str = 'positive') → pandas.DataFrame[source]

Order trials by recovery of a named control role.

Parameters:
  • results (pandas.DataFrame) – Sweep results containing <role>_control_percentile and optionally status.

  • role ({'positive', 'negative'}, default='positive') – Control role whose percentile determines the ordering.

Returns:

pandas.DataFrame – Copy ordered by increasing control percentile, with missing values and failed trials last. The input object is returned unchanged if the percentile column is absent or contains no finite values.

Notes

Percentile is used instead of raw rank so trials that fit different numbers of coefficients remain comparable. Trials that did not recover the control remain in the table rather than being dropped.

spacr.parameter_sweep.recommended_workers(*, measured_gib=None, requested=None)[source]

Choose a worker count from available memory and CPU capacity.

Parameters:
  • measured_gib (float or None, optional) – Peak resident memory for one representative trial, in GiB. None uses ASSUMED_TRIAL_GIB.

  • requested (int or None, optional) – Preferred maximum worker count. The result is still limited by memory, available CPU cores, and MAX_WORKERS.

Returns:

  • workers (int) – Recommended number of worker processes, always at least one.

  • reason (str) – Explanation of the memory estimate and any reduction from the requested count, suitable for logs. The Qt screen renders the same budget with localized templates.

Notes

The calculation budgets MEMORY_BUDGET_FRACTION of currently available memory. If memory cannot be measured, at most two workers are recommended.

spacr.parameter_sweep.rerun_trial(base_settings: Mapping[str, Any], row: Mapping[str, Any], *, destination: str | None = None) → dict[source]

Re-run one trial and hand back its settings, output and FIGURES.

The figures are live matplotlib Figures, not paths: a saved page cannot be restyled, and the point of clicking a row is to look at that condition properly – change the thresholds, recolour it, fix the legend – rather than to be shown a picture of it.

Only figures this call created are returned. A screen that already has figures open must not have them swept up and re-attributed to a trial they did not come from.

spacr.parameter_sweep.run_sweep(base_settings: Mapping[str, Any], destination, space: SweepSpace | None = None, *, mode: str = 'grid', max_trials: int = 5000, seed: int = 0, controls: Mapping[str, str] | None = None, progress_every: int = 10, learn_from_failures: int = 2, corrections: Sequence[str] | None = None, contained: bool = True, qc: bool = False, memory_floor_gb: float = FREE_MEMORY_FLOOR_GB, runner: Callable | None = None) → pandas.DataFrame[source]

Run parameter-sweep trials sequentially.

Parameters:
  • base_settings (mapping) – Regression settings shared by every trial, including the score and count inputs.

  • destination (path-like) – Directory for trial folders, sweep_trials.json, and sweep_results.csv.

  • space (SweepSpace or None, optional) – Search space. None uses DEFAULT_SWEEP_SPACE and the built-in compatibility filters.

  • mode ({'grid', 'random'}, default='grid') – Trial enumeration order passed to build_trials().

  • max_trials (int, default=5000) – Maximum number of accepted trials.

  • seed (int, default=0) – Random-order seed used when mode='random'.

  • controls (mapping or None, optional) – Mapping from control aliases to identifiers. Control recovery metrics are added to each successful result row.

  • progress_every (int, default=10) – Print and attempt an incremental CSV write after this many trials. Use zero to disable periodic progress messages; a CSV is still written after each completed in-process trial.

  • learn_from_failures (int, default=2) – Skip later trials with the same model, inference, analysis-unit, and penalty signature after this many matching failures. Use zero to run every trial.

  • corrections (sequence of str or None, optional) – Multiple-testing methods to apply to the p-values from one fitted model, producing one row per method. This option applies to the in-process path only; contained children return summary rows rather than coefficient frames.

  • contained (bool, default=True) – Run each trial through run_trial_contained(). Hard resource limits are conditional on a usable systemd user scope; otherwise the child is uncapped but uses reduced priority and thread limits.

  • qc (bool, default=False) – Generate the full regression diagnostic figure suite for every trial.

  • memory_floor_gb (float, default=FREE_MEMORY_FLOOR_GB) – Stop starting contained trials when available memory falls below this threshold, in GB.

  • runner (callable or None, optional) – In-process regression callable. None uses contained child trials when contained=True and spacr.ml.perform_regression() otherwise. Injected callables bypass child containment and the memory floor.

Returns:

pandas.DataFrame – Trial settings, status, timing, fit metrics, and control metrics. The frame is also written to sweep_results.csv as trials finish.

spacr.parameter_sweep.run_sweep_parallel(base_settings: Mapping[str, Any], destination, space: SweepSpace | None = None, *, mode: str = 'random', max_trials: int = 1000, seed: int = 0, controls: Mapping[str, str] | None = None, n_jobs: int = 8, contained: bool = True, qc: bool = False, progress_every: int = 25) → pandas.DataFrame[source]

Run parameter-sweep trials concurrently in spawned processes.

Parameters:
  • base_settings (mapping) – Regression settings shared by every trial, including the score and count inputs.

  • destination (path-like) – Directory for trial folders, sweep_trials.json, and the incremental sweep_results.csv table.

  • space (SweepSpace or None, optional) – Search space. None uses DEFAULT_SWEEP_SPACE and the built-in compatibility filters.

  • mode ({'grid', 'random'}, default='random') – Trial enumeration order passed to build_trials().

  • max_trials (int, default=1000) – Maximum number of accepted trials.

  • seed (int, default=0) – Random-order seed used when mode='random'.

  • controls (mapping or None, optional) – Mapping from control aliases to identifiers. Control recovery metrics are added to each result row.

  • n_jobs (int, default=8) – Requested pool size. recommended_workers() may reduce it based on memory and CPU capacity.

  • contained (bool, default=True) – Run each fit in a separate child through run_trial_contained(). Hard memory, swap, task, and CPU limits are enforced only when a usable systemd user scope is available; otherwise the child is uncapped but uses reduced priority and thread limits. Set to False only when trial resource use is known.

  • qc (bool, default=False) – Generate the full regression diagnostic figure suite for every trial.

  • progress_every (int, default=25) – Print progress after this many completed trials. Use zero to disable periodic progress messages.

Returns:

pandas.DataFrame – One status row per trial, sorted by trial_id. The same rows are written incrementally to sweep_results.csv.

Raises:

RuntimeError – If called from a worker process. Scripts must call this function under an if __name__ == '__main__': guard.

Notes

The pool uses the spawn start method because fitted models may import torch and OpenMP runtimes. Only a bounded number of jobs are submitted at once, allowing new submissions to pause when available memory is low. Initial processing workers start ten seconds apart. After every primary trial finishes, explicit resource-overload failures receive one serial retry in their existing trial directories. Original result/error files are retained separately, and the final table has one row per trial with primary failure details and overload_retry metadata. Invalid inputs, cancellation and unexplained native worker exits are not retried.

spacr.parameter_sweep.run_trial_contained(settings: Mapping[str, Any], *, trial_id=None, controls: Mapping[str, str] | None = None, timeout: float = 1800.0, memory_max: str = TRIAL_MEMORY_MAX, cpu_quota: str = TRIAL_CPU_QUOTA) → dict[source]

Run one trial in a fresh child process and return a status row.

When containment_available() is true, a systemd user scope enforces memory_max, disables swap, and applies cpu_quota. Otherwise the child still uses reduced priority and thread limits, but it has no hard memory ceiling; a warning is printed before it starts.

Parameters:
  • settings (mapping) – Regression settings for the child process.

  • trial_id (optional) – Identifier copied into the returned row.

  • controls (mapping, optional) – Named control identifiers used to summarize the result.

  • timeout (float, default=1800) – Maximum child runtime in seconds.

  • memory_max (str, default=TRIAL_MEMORY_MAX) – Systemd memory limit used when kernel containment is available.

  • cpu_quota (str, default=TRIAL_CPU_QUOTA) – Systemd CPU quota used when kernel containment is available.

Returns:

dict – The child’s result, or a row whose status is "killed", "timeout", or "failed".

spacr.parameter_sweep.settings_for_trial(base_settings: Mapping[str, Any], row: Mapping[str, Any], *, destination: str | None = None) → dict[source]

The full settings dict that produced row.

A sweep row is not just a record of what happened – it carries every setting the trial was given, which is what lets a user click a row and get that exact regression back rather than an approximation of it.

Values arrive as strings when the row came from the CSV rather than from memory, so they are parsed back to the types spaCR expects. A setting that will not parse is passed through unchanged: a string that was always a string must survive the round trip.

Parameters:
  • base_settings – the inputs the sweep ran on (score/count CSVs and the response column), which are not recorded per trial.

  • row – one row of the sweep results table.

  • destination – where to write this run’s output. Defaults to the folder the trial originally used.

spacr.parameter_sweep.summarise_sweep(results: pandas.DataFrame, *, controls: Sequence[str] = ('gra14', 'eaf1')) → dict[source]

Summarize sweep completion, controls, and hit-count sensitivity.

Parameters:
Returns:

dict – Trial counts, elapsed minutes, failure categories, control recovery, and hit-count ranges. Median hit counts are also grouped by correction, regression family, analysis unit, and inference mode when those columns are present. An empty input returns {'trials': 0}.

Notes

The summary emphasizes consistency across defensible analysis choices. Hit counts alone should not be used to select a model or correction.

Nested helpers

_default_filters.aggregation_belongs_to_wells(trial)

Reject aggregation variants that a per-cell analysis ignores.

spacr/parameter_sweep.py:157

_default_filters.mixed_replaces_the_backend(trial)

Reject random row/column effects beside a non-mixed backend.

spacr/parameter_sweep.py:149

_default_filters.penalty_belongs_to_penalised_families(trial)

Reject non-default alpha values for families that never read them.

spacr/parameter_sweep.py:195

_default_filters.permutation_at_cell_level_exhausts_memory(trial)

Reject the per-cell permutation fit measured to require 57 GiB.

spacr/parameter_sweep.py:179

_default_filters.permutation_has_no_row_column_terms(trial)

Reject row/column effects that the plate-blocked permutation omits.

spacr/parameter_sweep.py:187

_default_filters.permutation_ignores_the_family(trial)

Reject backend variants unused by nonparametric inference.

spacr/parameter_sweep.py:172

_default_filters.quantile_needs_its_own_unit(trial)

Reject well-level quantile fits, which require per-object rows.

spacr/parameter_sweep.py:164

build_trials.accept(trial)

Return the first filter’s rejection reason, or None.

spacr/parameter_sweep.py:251

run_sweep_parallel._fill()

Top the pool up, but never past what memory currently allows.

Submitting all 5,000 futures at once means the pool decides when to start each trial and nothing can intervene. Keeping only n_jobs in flight is what lets the memory floor below actually stop the sweep growing while the user’s editor is running.

spacr/parameter_sweep.py:1031