spacr.errors¶
Fail-loud error accounting for spaCR pipelines.
spaCR processes batches — 384 wells, thousands of fields, dozens of
image files. A single unreadable image must not abort the plate, so
almost every batch loop in the codebase wraps its body in
try/except Exception and carries on. The historical problem was
not the surviving; it was that the survival left no trace: forty wells
would fail to segment, forty lines would scroll past in a log nobody
was watching, measurements.db would be written anyway, and every
downstream regression would silently run on 344 wells while reporting
as if it had 384.
This module supplies the missing half — survive, but account for it, and make the accounting impossible to miss:
RunLedgerrecords every per-item success and failure, logs each failure atERRORwith the item id and traceback, and prints one loud, grouped block at the end of the run.RunLedger.stamp()writes that verdict into the artifact — arun_statustable inside a SQLite database, or a sibling<name>.run_status.jsonnext to any other output — so a later reader can tell that the result is partial.read_run_status()/run_is_complete()/assert_run_complete()let downstream code check before it trusts a file.ConfigurationErrormarks the failures that must not be survived. A wrongsrcpath or a missing metadata column is not a per-item problem — continuing past it only produces garbage.RunLedger.item()deliberately re-raises it.
Typical adoption inside a batch loop:
from .errors import RunLedger
ledger = RunLedger('convert_to_yokogawa')
for file in files:
with ledger.item(file, stage='convert'):
convert(file)
ledger.finalize(artifact=csv_path)
and downstream, before trusting the output:
from spacr.errors import run_is_complete
if not run_is_complete(db_path):
... # the numbers in here are computed on a subset
The module is deliberately stdlib-only. It is imported by
io/core/measure/deep_spacr/plot at module scope, so
it must never drag in torch, cellpose, pandas or numpy.
Exceptions¶
The run was set up wrongly and cannot produce a valid result. |
|
The data produced cannot be trusted. |
|
Raised by |
|
The stamp could not be read, so the run's verdict is unknown. |
|
Base class for every spaCR-raised error. |
Classes¶
Functions¶
|
Raise |
|
Raise |
|
Read back every |
|
True when no stamp on |
|
True when recoverable setup errors should be raised instead of printed. |
Module Contents¶
- exception spacr.errors.ConfigurationError[source]¶
Bases:
SpacrErrorThe run was set up wrongly and cannot produce a valid result.
A missing
srcfolder, an unparseable regex, a metadata column that does not exist. These are not per-item failures: continuing past one produces garbage for every item, soRunLedger.item()re-raises this instead of recording it.Initialize self. See help(type(self)) for accurate signature.
- exception spacr.errors.DataIntegrityError[source]¶
Bases:
SpacrErrorThe data produced cannot be trusted.
Raised when an artifact is internally inconsistent, or when a run failed on so many items that its output is not meaningful.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.errors.PartialRunError[source]¶
Bases:
DataIntegrityErrorRaised by
RunLedger.raise_if_worse_than()past the threshold.A subclass of
DataIntegrityErrorso callers that only care about “the answer is wrong” can catch the parent.Initialize self. See help(type(self)) for accurate signature.
- exception spacr.errors.RunStatusUnreadable[source]¶
Bases:
DataIntegrityErrorThe stamp could not be read, so the run’s verdict is unknown.
Distinct from “this artifact was never stamped”, which is a perfectly ordinary state and reads as
[]/ complete. This one means the reader was stopped: the database is locked by a writer that still holds it, or the file is truncated or otherwise not a database.Both of those are what an interrupted run leaves behind, and both used to be swallowed by one
except sqlite3.Error: return []and reported as “no stamps, therefore complete” — so a run killed mid-write, whose process still held the lock, read as finished. Measured on a real measurements.db stampedpartial:run_is_completesaidFalsewhen nothing held the file, andTrue— withassert_run_completepassing — while a second connection heldBEGIN EXCLUSIVE. The database said the run failed a field either way; the lock is what stopped anyone hearing it.Initialize self. See help(type(self)) for accurate signature.
- exception spacr.errors.SpacrError[source]¶
Bases:
ExceptionBase class for every spaCR-raised error.
Catch this to catch anything spaCR raises deliberately, as opposed to an incidental
ValueErrorfrom numpy or pandas.Initialize self. See help(type(self)) for accurate signature.
- class spacr.errors.Failure[source]¶
One recorded per-item failure.
- Parameters:
item – identifier of the thing that failed — a filename, a well id, a fold number. This is what makes the ledger actionable, so it is always stringified and never empty.
stage – pipeline stage the failure happened in.
exc_type – exception class name, used as the grouping key.
message –
str(exc).traceback_str – formatted traceback, kept so the
ERRORlog record carries the whole story.timestamp – Unix time the failure was recorded.
- class spacr.errors.RunLedger(name: str = 'run', logger: logging.Logger | None = None)[source]¶
Accounting for one batch run: what was attempted, what failed, and why.
A ledger is cheap — create one per pipeline invocation (or per source folder), wrap each loop body in
item(), and callfinalize()before returning.- Parameters:
name – run name, shown in the summary block and stored in the artifact stamp. Use the pipeline stage, e.g.
'measure_crop'.logger – logger to emit failure records on. Defaults to the module logger, which funnels into
~/.spacr/logs/spacr.logoncespacr.logging_util.setup_logging()has run.
Example
ledger = RunLedger('measure_crop') for well in wells: with ledger.item(well, stage='measure'): measure(well) ledger.finalize(artifact='measurements.db')
Initialize an empty, uniquely identified run ledger.
- finalize(artifact: str | os.PathLike | None = None, threshold: float | None = None, quiet_when_clean: bool = True) RunLedger[source]¶
Emit the summary, stamp the artifact, then optionally abort.
Call this as the last thing a pipeline function does, so the verdict is the last thing on screen rather than 400 lines up.
- Parameters:
artifact – path to the file this run produced. Stamped via
stamp()so the artifact itself records that it is partial.threshold – when given,
raise_if_worse_than()is applied after the artifact has been stamped, so the evidence survives the abort.quiet_when_clean – when True (default) a run with no failures prints nothing and only logs at
INFO. A run with failures always prints the loud block.
- Returns:
self.
- grouped_failures() OrderedDict[str, List[Failure]][source]¶
Group failures by exception type, in first-seen order.
Forty identical
FileNotFoundErrors are one problem, not forty, and this is what makes the summary readable.
- item(name: Any, stage: str | None = None, echo: str | None = None) Iterator[RunLedger][source]¶
Run one loop body: swallow and record its failure, keep the batch alive.
On a clean exit the item is counted as a success. On an ordinary exception the item is recorded as a failure and the loop carries on.
Two things are deliberately re-raised rather than recorded:
ConfigurationError— a wrongsrcpath is not a per-item failure, and pretending it is would turn one mistake into N recorded “data” errors.KeyboardInterrupt/SystemExit— Ctrl-C must abort, not be filed as a corrupt image.
- Parameters:
name – identifier of this item, recorded verbatim.
stage – pipeline stage; defaults to the ledger name.
echo – when set, a failure additionally prints
f"{echo}: {exc}"to stdout. Adoption sites use this to keep the exact console message users already rely on.
Example
for path in paths: with ledger.item(path, stage='load'): arrays.append(np.load(path))
- raise_if_worse_than(threshold: float, message: str | None = None) RunLedger[source]¶
Abort when the failure rate is strictly above
threshold.Use where a partial result is not merely incomplete but meaningless — a 5-fold cross-validation in which 3 folds died does not have a spread worth reporting.
- Parameters:
threshold – fraction in
[0, 1].0.5aborts when more than half the items failed; a rate exactly equal to the threshold does not abort.message – override for the error text.
- Raises:
PartialRunError – when the rate exceeds
threshold.- Returns:
selfwhen the run is acceptable.
- record_failure(item: Any, stage: str | None = None, exc: Any = None) Failure[source]¶
Record that
itemfailed, logging it loudly atERROR.- Parameters:
item – identifier of the item that failed — the thing a human needs in order to go and look at it.
stage – pipeline stage; defaults to the ledger name.
exc – the caught exception. A plain string is accepted for failures that were detected rather than raised.
- Returns:
the stored
Failure.
- record_success(item: Any, stage: str | None = None) RunLedger[source]¶
Record that
itemcompleted cleanly.- Parameters:
item – identifier of the processed item.
stage – pipeline stage; defaults to the ledger name.
- Returns:
self, so calls can be chained.
- stamp(artifact: str | os.PathLike) pathlib.Path[source]¶
Record this run’s verdict into the artifact it produced.
For a SQLite path (
.db/.sqlite/.sqlite3) a row is appended to theRUN_STATUS_TABLEtable. For anything else a sibling<stem>.run_status.jsonis written next to the file. Either way a later reader — a person orread_run_status()— can tell the result is partial.Stamps accumulate, so a database written by several stages ends up with one row per stage.
- Parameters:
artifact – path of the file this run produced.
- Returns:
the path actually written (the db, or the sidecar).
- summary(max_groups: int = 10, max_examples: int = 3) str[source]¶
Render the loud end-of-run block.
- Parameters:
max_groups – at most this many exception types are shown.
max_examples – at most this many distinct messages are shown per exception type.
- Returns:
a multi-line string, ready to print.
- to_json(path: str | os.PathLike) pathlib.Path[source]¶
Write
to_dict()topathas JSON, creating parent dirs.- Parameters:
path – destination file.
- Returns:
the written
Path.
- spacr.errors.assert_run_complete(artifact: str | os.PathLike, timeout: float = RUN_STATUS_READ_TIMEOUT) None[source]¶
Raise
DataIntegrityErrorifartifactis stamped partial.This guard prevents downstream analysis of incomplete output. An unreadable status raises
RunStatusUnreadable, a subclass ofDataIntegrityError, so callers may handle failed runs and unverifiable run status with one exception type.- Parameters:
artifact – path of a spaCR output.
timeout – seconds to wait for a locked database.
- Raises:
DataIntegrityError – when any stamp recorded a failure.
RunStatusUnreadable – when the status cannot be read at all.
- spacr.errors.raise_if_strict(message: str, exc: BaseException | None = None, settings: Any = None, error_type: type = ConfigurationError) bool[source]¶
Raise
error_type(message)in strict mode; otherwise log it atERROR.Used at the category-B sites that historically printed and carried on. The default path keeps the legacy behaviour but stops the problem from being invisible to the log; setting
SPACR_STRICT_ERRORS=1turns the same site into a hard stop.- Parameters:
message – what went wrong, and why the result is untrustworthy.
exc – the caught exception, chained onto the raise.
settings – settings dict consulted for
strict_errors.error_type – exception class to raise.
- Returns:
False when not strict, so callers can branch on it.
- spacr.errors.read_run_status(artifact: str | os.PathLike, timeout: float = RUN_STATUS_READ_TIMEOUT) List[Dict[str, Any]][source]¶
Read back every
RunLedger.stamp()recorded forartifact.Works for both stamp flavours: a SQLite path is read from its
RUN_STATUS_TABLE, anything else from its<stem>.run_status.jsonsidecar.Three outcomes, deliberately kept apart:
stamps exist — they are returned, oldest first;
the artifact exists and holds no stamp —
[], meaning “no information”. Stamping is opt-in, so this covers every output written before a ledger reached that code path;the artifact could not be read —
RunStatusUnreadable. A database still locked by its writer, or truncated by akillmid-write, is exactly what an interrupted run leaves, and folding it into the second case is how an interrupted run came back “complete”.
- Parameters:
artifact – path of a spaCR output — e.g.
measurements.db.timeout – seconds to wait for a locked database before giving up. Default
RUN_STATUS_READ_TIMEOUT.
- Returns:
one dict per recorded run, oldest first. Each has
status/n_attempted/n_succeeded/n_failed/failure_rate/summaryand afailureslist. An artifact that was never stamped returns[].- Raises:
RunStatusUnreadable – when the artifact exists but cannot be read — locked, truncated, corrupt, or malformed JSON.
Example
from spacr.errors import read_run_status for run in read_run_status('/data/plate1/measurements/measurements.db'): if run['n_failed']: print(run['summary'])
- spacr.errors.run_is_complete(artifact: str | os.PathLike, timeout: float = RUN_STATUS_READ_TIMEOUT) bool[source]¶
True when no stamp on
artifactrecorded a failure.An artifact that was never stamped reads as complete — stamping is opt-in and predates neither the older outputs on disk nor the code paths that have not adopted a ledger yet. Use
read_run_status()when you need to distinguish “verified clean” from “no information”.An artifact whose status cannot be read reads as not complete. That is the one case where the two answers differ in consequence: an unstamped file is silent, whereas a locked or truncated one is positive evidence that something was interrupted, and answering “complete” there is how a killed run passed for a finished one.
- Parameters:
artifact – path of a spaCR output.
timeout – seconds to wait for a locked database.
- spacr.errors.strict_errors(settings: Any = None) bool[source]¶
True when recoverable setup errors should be raised instead of printed.
Resolution order: an explicit
strict_errorskey insettingswins; otherwise theSTRICT_ENV_VARenvironment variable is consulted. Off by default, so adopting this never changes an existing pipeline’s behaviour.- Parameters:
settings – a spaCR settings dict, or None.