spacr.validate

Pre-flight validation of a spaCR settings dict.

Every crash this module exists to prevent was, in practice, a twenty-line check away: an organelle_channel one past the end of a three-channel plate, a src that points at the plate folder instead of plate/merged, a cell_mask_dim beyond the last plane of the merged array, an integer that came back from a settings CSV as the string "4". Each of those costs a full GPU run to discover.

The module is deliberately dependency-light: it imports nothing from spaCR except spacr.settings (which itself imports only os and ast), and it touches numpy only lazily, to read the header of a single .npy file. No torch, no cellpose, no image decoding. Importing and running it costs well under a second, which is the whole point — it is what dry_run uses to answer “would this run work?” before anything is allocated, loaded or written.

Public API

Problem

One thing that is wrong, with the fix.

validate_settings(settings, app_key)

Returns a list of Problem, errors and warnings mixed.

format_report(problems, settings, app_key)

Human-readable report; errors first, each with its fix line.

describe_plan(settings, app_key)

“Here is what would actually happen” summary.

Rules are derived from the code that consumes the settings, not invented: each check names its source in a comment.

Classes

Problem

One thing wrong with a settings dict.

Functions

coerce_expected_types(→ Dict[str, Any])

Return settings with text-written numbers as their declared type.

describe_plan(→ str)

Summarise what the run would actually do, without doing it.

describe_resources(→ str)

Project what the run would COST, without running it.

format_report(→ str)

Render validate_settings() output for a terminal.

run_preflight(→ List[Problem])

Validate, print the report and the plan, and hand back the problems.

validate_settings(→ List[Problem])

Check a settings dict against the data it points at.

Module Contents

class spacr.validate.Problem[source]

One thing wrong with a settings dict.

Parameters:
  • severity – "error" (the run would fail or silently produce wrong output) or "warning" (suspicious, but runnable).

  • setting – the settings key at fault, or "" when the problem is about the dataset rather than a single key.

  • message – what is wrong, phrased in the user’s terms.

  • fix – what to actually do about it.

__str__() → str[source]

The problem and its fix, on two lines.

The setting’s name leads when there is one, so a reader scanning a list of problems sees WHICH setting each belongs to before the message.

property is_error: bool[source]

True when this problem would break or corrupt the run.

spacr.validate.coerce_expected_types(settings: Dict[str, Any], app: str = '') → Dict[str, Any][source]

Return settings with text-written numbers as their declared type.

A settings CSV round-trip makes every value a string, and so does a number typed into a GUI field. expected_types is the contract those values are meant to satisfy, so converting them to it is restoring what the settings already claim to be – not reinterpreting them.

Doing it HERE, once, at the boundary, rather than at each point of use, is what stops the next consumer from being the one that crashes: mask generation died inside Cellpose on diameter > 0 with cell_diameter='60.0', and the same file had already been reported, three times, as an error the user was told to fix by hand – for a value that was perfectly well-formed.

CONSERVATIVE BY CONSTRUCTION. Only bool, int and float are converted, only from a string, only when the conversion is exact, and never when str is itself an accepted type for the key. Anything that does not convert cleanly is left exactly as it was, for validate_settings() to report.

Parameters:
  • settings – the settings mapping.

  • app – the pipeline, for the same per-app type overrides _check_types() honours.

Returns:

a new dict; the input is not modified.

spacr.validate.describe_plan(settings: Dict[str, Any], app_key: str = '') → str[source]

Summarise what the run would actually do, without doing it.

Reports the app and the function behind it, the resolved source folder, how many files were found, which objects would be segmented or measured with which channels and diameters, where output would land, and roughly how many images would be processed.

Parameters:
  • settings – the settings dict about to be handed to a pipeline.

  • app_key – which pipeline, as for validate_settings().

Returns:

the plan as a single string, no trailing newline.

spacr.validate.describe_resources(settings: Dict[str, Any], app_key: str = '') → str[source]

Project what the run would COST, without running it.

The second half of the dry-run card. describe_plan() says what would happen; this says whether this machine can see it through, which is the question that actually stops a run at three in the morning.

Parameters:
  • settings – the settings dict about to be handed to a pipeline.

  • app_key – which pipeline, as for validate_settings().

Returns:

the card as a single string, no trailing newline.

WHAT IS MEASURED, NOT GUESSED: the bytes already on disk, the shape and dtype of one real array, the free memory, and the free space on the volume that would be written to. Every one is read off the machine.

WHAT IS DERIVED, AND SAID TO BE: memory is reported as a FLOOR – n_jobs workers each holding one field is the least the run can use, before a single working copy. A floor is defensible and still catches the case that matters, which is eight workers holding a 900 MB field each on a 16 GB laptop. The measured peak is what spacr.benchmark.benchmark() reports, and this card says so rather than inventing a multiplier for it.

WHAT IS REFUSED: the size of the PNG crop tree. It is objects times crop modes, and the object count is the thing the run exists to discover. The card names it as unbounded rather than producing a number that would be wrong by an order of magnitude either way.

spacr.validate.format_report(problems: Sequence[Problem], settings: Dict[str, Any] | None = None, app_key: str = '') → str[source]

Render validate_settings() output for a terminal.

Errors come first as a group, then warnings, each entry followed by its fix line. A clean run says so explicitly.

Parameters:
  • problems – what validate_settings() returned.

  • settings – the settings dict, used only for the header line.

  • app_key – the app the check was run for, used only for the header.

Returns:

the report as a single string, no trailing newline.

spacr.validate.run_preflight(settings: Dict[str, Any], app_key: str, printer=print, trailer: str = DRY_RUN_TRAILER) → List[Problem][source]

Validate, print the report and the plan, and hand back the problems.

This is what the dry_run branch of every pipeline entry point calls, so the wording is identical wherever it is triggered from.

Parameters:
  • settings – the settings dict that would have been run.

  • app_key – which pipeline, as for validate_settings().

  • printer – where the text goes; defaults to print.

  • trailer – the closing line. The default names the in-pipeline dry_run flag, which is the wrong advice for a caller reached some other way – spacr-run --dry-run passes its own wording. Pass an empty string to suppress it entirely.

Returns:

the list of Problem found.

spacr.validate.validate_settings(settings: Dict[str, Any], app_key: str) → List[Problem][source]

Check a settings dict against the data it points at.

Image data are checked through headers and directory listings. An enabled PSF also loads its size-limited kernel to validate calibration and values; no image processing or GPU inference is run.

Parameters:
  • settings – the settings dict about to be handed to a pipeline.

  • app_key – which pipeline, e.g. 'mask', 'measure', 'classify', 'umap', 'map_barcodes'. The settings_type strings used by the GUI all work, as do a few aliases ('sequencing', 'measure_crop', …).

Returns:

list of Problem; empty means nothing was found.

Nested helpers

_check_required_paths._require_file(key: str, purpose: str, fix: str) → None

Refuse a missing or unset path, saying what it was needed FOR.

Both the purpose and the fix are carried into the message: “not set” on its own tells a user what happened and not what to do about it.

spacr/validate.py:1477

_describe_regression_plan.paths(value)

Normalize an optional scalar or sequence source into a list without inventing entries.

spacr/validate.py:2041