spacr.validate¶
Pre-flight validation of a spaCR settings dict.
Every crash this module exists to prevent was, in practice, a twenty-line
check away: an organelle_channel one past the end of a three-channel
plate, a src that points at the plate folder instead of plate/merged,
a cell_mask_dim beyond the last plane of the merged array, an integer
that came back from a settings CSV as the string "4". Each of those costs
a full GPU run to discover.
The module is deliberately dependency-light: it imports nothing from spaCR
except spacr.settings (which itself imports only os and ast),
and it touches numpy only lazily, to read the header of a single
.npy file. No torch, no cellpose, no image decoding. Importing and
running it costs well under a second, which is the whole point — it is what
dry_run uses to answer “would this run work?” before anything is
allocated, loaded or written.
Public API¶
ProblemOne thing that is wrong, with the fix.
validate_settings(settings, app_key)Returns a list of
Problem, errors and warnings mixed.format_report(problems, settings, app_key)Human-readable report; errors first, each with its fix line.
describe_plan(settings, app_key)“Here is what would actually happen” summary.
Rules are derived from the code that consumes the settings, not invented: each check names its source in a comment.
Classes¶
One thing wrong with a settings dict. |
Functions¶
|
Return |
|
Summarise what the run would actually do, without doing it. |
|
Project what the run would COST, without running it. |
|
Render |
|
Validate, print the report and the plan, and hand back the problems. |
|
Check a settings dict against the data it points at. |
Module Contents¶
- class spacr.validate.Problem[source]¶
One thing wrong with a settings dict.
- Parameters:
severity –
"error"(the run would fail or silently produce wrong output) or"warning"(suspicious, but runnable).setting – the settings key at fault, or
""when the problem is about the dataset rather than a single key.message – what is wrong, phrased in the user’s terms.
fix – what to actually do about it.
- spacr.validate.coerce_expected_types(settings: Dict[str, Any], app: str = '') Dict[str, Any][source]¶
Return
settingswith text-written numbers as their declared type.A settings CSV round-trip makes every value a string, and so does a number typed into a GUI field.
expected_typesis the contract those values are meant to satisfy, so converting them to it is restoring what the settings already claim to be – not reinterpreting them.Doing it HERE, once, at the boundary, rather than at each point of use, is what stops the next consumer from being the one that crashes: mask generation died inside Cellpose on
diameter > 0withcell_diameter='60.0', and the same file had already been reported, three times, as an error the user was told to fix by hand – for a value that was perfectly well-formed.CONSERVATIVE BY CONSTRUCTION. Only
bool,intandfloatare converted, only from a string, only when the conversion is exact, and never whenstris itself an accepted type for the key. Anything that does not convert cleanly is left exactly as it was, forvalidate_settings()to report.- Parameters:
settings – the settings mapping.
app – the pipeline, for the same per-app type overrides
_check_types()honours.
- Returns:
a new dict; the input is not modified.
- spacr.validate.describe_plan(settings: Dict[str, Any], app_key: str = '') str[source]¶
Summarise what the run would actually do, without doing it.
Reports the app and the function behind it, the resolved source folder, how many files were found, which objects would be segmented or measured with which channels and diameters, where output would land, and roughly how many images would be processed.
- Parameters:
settings – the settings dict about to be handed to a pipeline.
app_key – which pipeline, as for
validate_settings().
- Returns:
the plan as a single string, no trailing newline.
- spacr.validate.describe_resources(settings: Dict[str, Any], app_key: str = '') str[source]¶
Project what the run would COST, without running it.
The second half of the dry-run card.
describe_plan()says what would happen; this says whether this machine can see it through, which is the question that actually stops a run at three in the morning.- Parameters:
settings – the settings dict about to be handed to a pipeline.
app_key – which pipeline, as for
validate_settings().
- Returns:
the card as a single string, no trailing newline.
WHAT IS MEASURED, NOT GUESSED: the bytes already on disk, the shape and dtype of one real array, the free memory, and the free space on the volume that would be written to. Every one is read off the machine.
WHAT IS DERIVED, AND SAID TO BE: memory is reported as a FLOOR –
n_jobsworkers each holding one field is the least the run can use, before a single working copy. A floor is defensible and still catches the case that matters, which is eight workers holding a 900 MB field each on a 16 GB laptop. The measured peak is whatspacr.benchmark.benchmark()reports, and this card says so rather than inventing a multiplier for it.WHAT IS REFUSED: the size of the PNG crop tree. It is objects times crop modes, and the object count is the thing the run exists to discover. The card names it as unbounded rather than producing a number that would be wrong by an order of magnitude either way.
- spacr.validate.format_report(problems: Sequence[Problem], settings: Dict[str, Any] | None = None, app_key: str = '') str[source]¶
Render
validate_settings()output for a terminal.Errors come first as a group, then warnings, each entry followed by its fix line. A clean run says so explicitly.
- Parameters:
problems – what
validate_settings()returned.settings – the settings dict, used only for the header line.
app_key – the app the check was run for, used only for the header.
- Returns:
the report as a single string, no trailing newline.
- spacr.validate.run_preflight(settings: Dict[str, Any], app_key: str, printer=print, trailer: str = DRY_RUN_TRAILER) List[Problem][source]¶
Validate, print the report and the plan, and hand back the problems.
This is what the
dry_runbranch of every pipeline entry point calls, so the wording is identical wherever it is triggered from.- Parameters:
settings – the settings dict that would have been run.
app_key – which pipeline, as for
validate_settings().printer – where the text goes; defaults to
print.trailer – the closing line. The default names the in-pipeline
dry_runflag, which is the wrong advice for a caller reached some other way –spacr-run --dry-runpasses its own wording. Pass an empty string to suppress it entirely.
- Returns:
the list of
Problemfound.
- spacr.validate.validate_settings(settings: Dict[str, Any], app_key: str) List[Problem][source]¶
Check a settings dict against the data it points at.
Image data are checked through headers and directory listings. An enabled PSF also loads its size-limited kernel to validate calibration and values; no image processing or GPU inference is run.
- Parameters:
settings – the settings dict about to be handed to a pipeline.
app_key – which pipeline, e.g.
'mask','measure','classify','umap','map_barcodes'. Thesettings_typestrings used by the GUI all work, as do a few aliases ('sequencing','measure_crop', …).
- Returns:
list of
Problem; empty means nothing was found.
Nested helpers¶
- _check_required_paths._require_file(key: str, purpose: str, fix: str) None¶
Refuse a missing or unset path, saying what it was needed FOR.
Both the purpose and the fix are carried into the message: “not set” on its own tells a user what happened and not what to do about it.
spacr/validate.py:1477
- _describe_regression_plan.paths(value)¶
Normalize an optional scalar or sequence source into a list without inventing entries.
spacr/validate.py:2041