spacr.resume

Resume / checkpointing for spaCR’s two long batch stages.

A plate is a thousand fields and several hours. When a run dies at field 900 — a full disk, an OOM-killed worker, a laptop lid — spaCR currently throws away all 900 and starts again. This module lets the next invocation pick up where the last one stopped.

The skipping is the easy half. The half that decides whether resuming is safe is everything else in here:

A file that exists is not a file that finished. spacr.io._load_and_concatenate_arrays writes merged/<field>.npy with a bare np.save straight onto the destination path, so a crash mid-write leaves a short, structurally valid-looking .npy on disk. A resume that skips on os.path.exists would accept that truncated field, measure whatever bytes happened to land, and put the result in the same table as the good fields. So completed_fields_in_merged() validates: it parses the .npy header and checks that the file is actually as long as its own shape claims. Anything short is reported as rejected and re-queued, never skipped. (io.py now also writes merged arrays atomically, so files produced from here on cannot be truncated — but the ones already on disk from previous runs can be.)

Measure appends; therefore a resume must delete first. _merge_and_save_to_database and filepaths_to_database both use to_sql(..., if_exists='append'). Re-measuring a field that already wrote some rows does not overwrite them, it adds to them, and every per-well aggregate downstream is then computed over inflated counts with nothing anywhere to indicate it. clear_field_rows() is the delete-before-insert that makes re-running a field idempotent, and it is the reason this module exists at all. It runs in one transaction across every table the field touched, so a failure part-way leaves the database exactly as it was.

measurements.db is not measure’s private file. convert writes conversion_map into it, align writes align_coordinates, foreign writes foreign_*, timelapse writes its track table — all keyed on the same four columns as the measurements, deliberately, so that they join. Discovering “tables to clear” structurally therefore found all of them, and a resume deleted every pending field’s row from each: the only record of which vendor file became plate1_A01_1, of where each tile was stitched, of somebody else’s imported measurements. None of it can be recomputed from the database. The tables a resume may delete from are now an explicit allow-list, MEASURE_OWNED_TABLES, checked both when the list is discovered and again per table immediately before the DELETE is prepared.

Owned by name is not owned in fact, so ownership is decided per row. A name allow-list answers “could measure have written this table?”, and there is one case where the honest answer to “did it?” is no: foreign.run_import copies the imported rows into the canonical cell / nucleus / pathogen table when nothing of anyone else’s is there, so that a purely-imported project is readable by every spaCR tool. Those rows sit under spaCR’s own metadata columns in a table on the allow-list, and a resume used to clear them along with the pending field.

The signal that tells them apart was already in the database and is what is read now: foreign_columns, written by every release of the importer, names each column that importer put in each table. Its table-scoped use — foreign._importer_owns, “are this table’s columns a subset of what the importer recorded?” — has a row-scoped twin, and that is the rule here: a row is the importer’s when every column it is non-NULL in is one the importer recorded for this table. A measure_crop row always carries at least its own area and intensity columns, which the importer never wrote; an imported row carries only metadata and foreign_-prefixed measurements. So a table can be half one and half the other and each row is still attributable — which matters, because a project that was imported and then measured is exactly that. See measure_rows_clause().

Row-scoping is what makes the earlier, backed-out attempt’s failures impossible rather than merely unlikely. That attempt read foreign_import.canonical_table_written and refused any delete from a table the record claimed: table-scoped, so a field the import never covered could not be cleared either and the project could never be resumed; unreachable on the flow it existed for; and open on the databases most likely to hit the bug, because the first importer wrote the copy without that marker. foreign_columns has none of those properties — it is per table, it is per column, and it is as old as the importer.

A copy nobody is going to keep is released, not left to be mixed into. measure_crop appends, so the moment spaCR measures a field into a canonical table an import filled, its rows and theirs sit in the same table with nothing marking the seam and every per-well count becomes the sum of two populations. supersede_imported_copies() runs before the first insert and hands that copy back — but only when every field it covers is either already measured into that table or queued to be measured now, and only after spacr.foreign.release_canonical_copy() has proved, row by row, that each row it removes still exists in foreign_<object>. A half-released table is worse than an unreleased one, so anything less than a complete, lossless release is refused and reported instead, together with the two lines that do it by hand.

What bounds all of it: every imported row in the canonical table has an identical twin in foreign_<object>, and foreign_<object> is not on the allow-list, so no resume can touch it. That is checked rather than assumed — a row without a twin stops the release.

One half of this is still open and is not in this module: _merge_and_save_to_database appends into a canonical table an import filled without noticing, so a user who measures such a project without a resume still gets both populations in one table. tests/test_resume_owned_tables.py pins it with xfail(strict=True) and names the change spacr/utils.py needs.

A database already damaged by an earlier resume cannot be repaired on read — the rows are gone, and nothing in the file records what they were. It can be rewritten from the sources outside the database, and all three writers use if_exists='replace', so re-running one is a full repair rather than a second generation of rows: convert.populate_db_from_map(db, '<converted>/conversion_map.csv') restores the map from the CSV that sits beside the converted images, align.save_coordinates restores the stitch coordinates, and foreign.run_import restores the imported measurements.

A field is identified by its well coordinates, never by a name prefix. Deleting with LIKE 'plate1_A01_f1%' also matches f10 … f19 — nineteen innocent fields destroyed to clean up one. Every statement in here matches on equality over plateID / rowID / columnID / fieldID (plus timeID for timelapse), and refuses to touch a table that does not carry all four.

A resume across different settings is not a resume. Concatenating fields measured with one channel/diameter/crop configuration onto fields measured with another produces a dataset that is half one thing and half another, and nothing downstream can tell. check_settings_compatible() compares the current settings against the ones recorded by the previous run and raises ResumeRefused naming exactly what differs. Environment drift (a numpy version bump) does not block.

Resume is opt-in: everything here is inert unless settings['resume'] is true, and a run without it behaves exactly as it did before.

The module is deliberately stdlib-only — os/sqlite3/ast and nothing else at import time. It is consulted at the very top of measure_crop, long before any model is loaded, and must never be the thing that drags torch or cellpose into a process.

Typical use, from measure_crop:

from .resume import plan_measure_resume
plan = plan_measure_resume(settings)      # None when resume is off
files = [f for f in os.listdir(src)
         if f.endswith('.npy') and not f.startswith('.')]
if plan is not None:
    files = plan.filter_files(files)

Exceptions

ResumeRefused

The requested resume would produce a dataset that is not one dataset.

Classes

ResumeState

What a resume decided to do, before it does any of it.

SettingsComparison

What differs between the recorded settings and the current ones.

Functions

check_settings_compatible(→ SettingsComparison)

Raise ResumeRefused when resuming would mix two datasets.

clear_field_rows(→ int)

Delete every row belonging to one field, from every table, atomically.

compare_settings(→ SettingsComparison)

Bucket the difference between two settings dicts by consequence.

completed_fields_in_db(→ Set[str])

Fields that are already measured in db_path.

completed_fields_in_merged(→ Set[str])

Field stems in src whose .npy is verified complete.

discover_field_tables(→ List[str])

The measure tables in db_path that store rows per field of view.

expected_min_planes(→ Optional[int])

Smallest last-axis size a merged stack must have for these settings.

field_identity(→ Dict[str, str])

Parse a merged-stack name into the well coordinates the database stores.

format_resume(→ str)

Render the plan as the block printed before any work starts.

identity_to_prcf(→ str)

Render an identity dict as the prcf string the tables also carry.

importer_owns_table(→ bool)

True when every column of table is one the importer recorded.

importer_recorded_columns(→ Optional[Set[str]])

Every column spacr.foreign recorded itself as writing into table.

importer_rows_clause(→ Optional[str])

The complement of measure_rows_clause() — the importer's rows.

importer_written_columns(→ Optional[Set[str]])

The widest set of columns this database shows an import having written.

measure_rows_clause(→ Optional[str])

A SQL condition selecting the rows of table the measure stage wrote.

measurements_db_path(→ str)

Where measure_crop writes, derived exactly as it derives it.

plan_measure_resume(→ Optional[ResumeState])

Plan (and make safe) a resumed measure_crop. The one call measure makes.

plan_resume(→ ResumeState)

Work out what to run, what to skip, and record why.

read_npy_header(→ Dict[str, Any])

Parse a .npy header without loading (or allocating) the array.

read_recorded_settings(→ Dict[str, Any])

Read the settings a previous run wrote, from wherever it wrote them.

resume_enabled(→ bool)

True when the caller asked for a resume.

run_already_complete(→ bool)

True when db_path carries a RunLedger stamp saying so.

supersede_imported_copies(→ Tuple[int, List[str]])

Release an import's convenience copy that this run is about to replace.

validate_merged_field(→ Tuple[bool, str])

Decide whether a merged/*.npy is a finished field or a crash scar.

Module Contents

exception spacr.resume.ResumeRefused[source]

Bases: spacr.errors.ConfigurationError

The requested resume would produce a dataset that is not one dataset.

Raised when the fields already on disk were produced under settings that differ materially from the current ones. This is a ConfigurationError on purpose: it is a setup mistake, not a per-field failure, and continuing past it would append apples to oranges in a single table with nothing to mark the seam.

Initialize self. See help(type(self)) for accurate signature.

class spacr.resume.ResumeState[source]

What a resume decided to do, before it does any of it.

Variables:
  • total – number of candidate fields considered.

  • done – fields verified complete — these are skipped.

  • pending – fields that will be run, in input order.

  • skipped – the fields actually skipped (done restricted to the candidates, so the two counts always add up to total).

  • reasons – {field: reason} for every pending field — 'truncated', 'partial-db-rows', 'not-measured', … This is the audit trail: a field that a naive exists() check would have skipped shows up here with the reason it was not.

  • enabled – False when resume was not requested, in which case pending is every field and nothing is skipped.

  • cleared_rows – stale rows deleted by the delete-before-insert.

  • released_rows – rows of a foreign import’s convenience copy that this run superseded and removed from a canonical table. They are unchanged in foreign_<object>; see supersede_imported_copies().

  • notes – things the user has to act on that are not per-field — today, a canonical table an import filled that could not be released automatically. Empty on every ordinary run.

  • src – the merged folder inspected, for the report.

  • db_path – the database inspected, for the report.

filter_files(files: Iterable[str]) → List[str][source]

Restrict a list of .npy filenames to the pending fields.

Order and extensions are preserved, so the caller’s list is the same list minus the fields already done. With enabled False the list is returned unchanged — that is what keeps resume opt-in.

Parameters:

files – filenames as os.listdir produced them.

Returns:

the subset still to process.

rejection_counts() → Dict[str, int][source]

{reason: count} over rejected, for the report.

property n_done: int[source]

How many fields were verified complete.

property n_pending: int[source]

How many fields will be run.

property n_rejected: int[source]

How many present-but-unusable fields were re-queued.

property n_skipped: int[source]

How many fields are being skipped.

property rejected: Tuple[str, ...][source]

Fields that were present but rejected — the important number.

A truncated .npy, or a field with rows in some tables and not others. These would have been silently skipped by an os.path.exists resume, and measuring them would have produced garbage that looked exactly like data.

class spacr.resume.SettingsComparison[source]

What differs between the recorded settings and the current ones.

Bucketed by consequence, not by key, following the same reasoning as spacr.run_journal.diff_runs(): a flat dict diff of two runs is mostly schema noise, and the one change that matters drowns in it.

Variables:
  • changed – material differences — these block the resume.

  • drift – cosmetic differences (worker counts, plotting) that cannot alter a measured number.

  • env – package/platform differences. Reported, never blocking.

  • only_in_recorded – keys the old run had and this one does not.

  • only_in_current – keys this run has and the old one did not, with an inert (null/empty) value.

  • same – count of keys that agree.

describe() → str[source]

One-line-per-change rendering naming exactly what differs.

property blocks_resume: bool[source]

True when at least one material setting differs.

spacr.resume.check_settings_compatible(recorded: Mapping[str, Any], current: Mapping[str, Any], source: str = '', cosmetic: Iterable[str] = COSMETIC_SETTINGS) → SettingsComparison[source]

Raise ResumeRefused when resuming would mix two datasets.

The fields already on disk were produced under recorded. If current differs in anything that affects the numbers, appending new fields to them yields a table that is half one experiment and half another, with no column marking the boundary — the single worst outcome this whole module exists to prevent. Better to refuse and make the user say what they meant.

Parameters:
  • recorded – settings read back from the previous run.

  • current – the settings this run would use.

  • source – where recorded came from, quoted in the error.

  • cosmetic – keys treated as inconsequential.

Returns:

the SettingsComparison when compatible.

Raises:

ResumeRefused – naming every material difference.

spacr.resume.clear_field_rows(db_path: str, tables: Sequence[str] | None, field: Any, timelapse: bool = False) → int[source]

Delete every row belonging to one field, from every table, atomically.

This is the delete-before-insert. Measure appends (to_sql(if_exists='append')), so re-running a field that already wrote rows adds a second copy of every object rather than replacing the first. Downstream, count_cell doubles, per-well means are computed over the doubled population, and nothing in the artifact says so. Calling this immediately before re-measuring a field is what makes a resume idempotent.

Three safety properties, all tested:

  • All or nothing. Every table is deleted from inside one BEGIN IMMEDIATE … COMMIT. If any statement fails — a missing table, a trigger, a lock — the transaction is rolled back and the database is left exactly as it was. A half-cleared field would be worse than an uncleared one.

  • Keyed on the field, not on its name. The WHERE clause is equality over plateID/rowID/columnID/fieldID, so clearing f1 cannot touch f10–f19 the way a LIKE 'plate1_A01_f1%' would. A table missing any of the four raises rather than running a broader delete.

  • Only measure’s own tables. Every table is checked against MEASURE_OWNED_TABLES before anything is deleted, and one that is not on it aborts the whole call. measurements.db is shared — conversion_map, align_coordinates and foreign_* live there and carry the same four key columns so that they join — and clearing a field out of those destroys the only record of how the project’s files were named and registered.

  • Only rows measure itself wrote. Owned by name is not owned in fact: in a project built by foreign.run_import the rows in cell are the import’s, sitting under spaCR’s own metadata columns in a table that is on the allow-list. Every DELETE is therefore narrowed by measure_rows_clause(), which is true only for rows holding a value in a column the importer never wrote. Clearing a field takes measure’s rows for that field and leaves the import’s — in the same table, in the same statement.

Parameters:
  • db_path – path to measurements.db.

  • tables – tables to clear. None uses discover_field_tables(), which is the safe default — naming tables by hand is how one gets missed.

  • field – field stem ('plate1_A01_3'), filename, or an identity mapping.

  • timelapse – include the timepoint in the key, so one frame can be cleared without touching the rest of the movie.

Returns:

total number of rows deleted.

Raises:
  • ValueError – when the field cannot be identified, a named table lacks the key columns, or a named table is not one measure writes.

  • sqlite3.Error – propagated after rollback.

Example

n = clear_field_rows(db, None, 'plate1_A01_3')
print(f'cleared {n} stale rows before re-measuring')
spacr.resume.compare_settings(recorded: Mapping[str, Any], current: Mapping[str, Any], cosmetic: Iterable[str] = COSMETIC_SETTINGS) → SettingsComparison[source]

Bucket the difference between two settings dicts by consequence.

A key is material unless it is explicitly known to be cosmetic. That direction is deliberate: an unrecognised new setting that changed will block the resume, which is the outcome one wants to be wrong in. The alternative — allow-listing the settings that matter — silently permits every knob nobody thought of.

A key present on only one side is schema drift and does not block, unless it appears only in the current settings with a real (non-null) value: enabling organelle_channel on the resume run genuinely changes what is measured even though the old run never recorded the key.

Parameters:
  • recorded – settings read back from the previous run.

  • current – the settings this run is about to use.

  • cosmetic – keys treated as inconsequential.

Returns:

a SettingsComparison.

spacr.resume.completed_fields_in_db(db_path: str, tables: Sequence[str] | None = None, fields: Iterable[str] | None = None, timelapse: bool = False, require_all: bool = True, partial: Dict[str, str] | None = None) → Set[str][source]

Fields that are already measured in db_path.

“Already measured” deliberately means present in every table this run writes, not “present somewhere”. _measure_crop_core inserts into cell, then nucleus, then pathogen, then cytoplasm, then png_list, one call after another with a commit each — so a process killed between two of those calls leaves a field that has cell rows and no nucleus rows. Counting that field as done would permanently lose its nuclei; counting it as pending without clearing it first would duplicate its cells. It is reported through partial so the caller can do the third, correct thing: clear it and re-run it.

“Present” means present in a row measure wrote. Rows a foreign import copied into a canonical table do not count: a project built purely by import has cell as its only field table and a row in it for every field, so counting them made this function report the whole plate measured, measure_crop run nothing at all, and the collaborator’s numbers stand as spaCR’s output. A table holding nothing but imported rows is left out of the “every table this run writes” set entirely, so it cannot make every field look partial either; it rejoins the moment measure puts a row in it. See measure_rows_clause().

Parameters:
  • db_path – path to measurements.db.

  • tables – tables to consult. Defaults to discover_field_tables().

  • fields – candidate field stems (e.g. from merged/). When given, the return value is the subset of these that are complete — which is what a caller filtering a file list wants. When None, prcf strings assembled from the rows themselves are returned.

  • timelapse – include the timepoint in the field key.

  • require_all – when False, a field counts as done if it appears in any table. Only useful for inspection.

  • partial – optional dict, populated with {field: 'partial-db-rows'} for fields present in some tables but not all. These must be cleared before being re-run.

Returns:

set of completed field identifiers.

spacr.resume.completed_fields_in_merged(src: str, min_planes: int | None = None, reasons: Dict[str, str] | None = None, fields: Iterable[str] | None = None) → Set[str][source]

Field stems in src whose .npy is verified complete.

This is the function that stops a resume from turning into corrupt output. It never answers on the strength of os.path.exists: each candidate’s header is parsed and its length checked against its own declared shape, because np.save writes merged arrays in place and a run killed mid-write leaves a short file behind.

Parameters:
  • src – the merged/ folder (or any folder of .npy fields).

  • min_planes – minimum last-axis size, see expected_min_planes(). When None and at least three files validate, the modal plane count across the folder is used instead — every field in a merged folder is written with the same plane count, so an odd one out is a bad field.

  • reasons – optional dict, populated in place with {field: reason} for every candidate that was rejected. This is how the caller can report “3 fields rejected as truncated” rather than silently doing more work.

  • fields – restrict the scan to these stems. Defaults to every .npy in src whose name does not start with a dot: a macOS ._ sidecar ends in .npy too and is not a field.

Returns:

set of field stems (basename without .npy) that are safe to skip.

Example

rejected = {}
done = completed_fields_in_merged('/data/p1/merged',
                                  min_planes=7, reasons=rejected)
print(f'{len(done)} done, {len(rejected)} rejected: {rejected}')
spacr.resume.discover_field_tables(db_path: str, include_non_field: bool = False, owned_only: bool = True) → List[str][source]

The measure tables in db_path that store rows per field of view.

Enumerated from the database rather than hard-coded, because a delete that misses one table leaves orphan rows that join incorrectly forever — and the set genuinely varies with the settings: no organelle table without an organelle channel, no *_organelle_summary tables without summarize_organelles_by, no png_list without save_png.

A table qualifies when it is in MEASURE_OWNED_TABLES and carries all of FIELD_KEY_COLUMNS. Both halves are needed. The column test alone is not a test of ownership: conversion_map, align_coordinates and foreign_* all carry the same four columns, precisely so that they join to the measurements, and they passed it — which is how a measure resume came to delete three other modules’ provenance tables. The name test alone would delete from a table whose schema is not what this module thinks it is.

Both halves are about the table, and a table on this list can still hold rows measure did not write — the copy foreign.run_import puts in cell. That is decided per row, later and separately, by measure_rows_clause(); this function deliberately still lists such a table, because it is one a resume has business with.

Parameters:
  • db_path – path to measurements.db.

  • include_non_field – skip the NON_FIELD_TABLES filter. For inspection only — never pass this to clear_field_rows().

  • owned_only – keep only MEASURE_OWNED_TABLES. Pass False to see every per-field table in the database whoever wrote it — for inspection only, and clear_field_rows() refuses the extras by name if that list is handed to it.

Returns:

sorted table names.

spacr.resume.expected_min_planes(settings: Any) → int | None[source]

Smallest last-axis size a merged stack must have for these settings.

Every plane index the measure stage will subscript — channels, png_dims and the four *_mask_dim entries — has to exist in the array, so the stack needs at least max(index) + 1 planes.

Parameters:

settings – a measure settings dict.

Returns:

the minimum plane count, or None when the settings name no plane at all.

spacr.resume.field_identity(field: Any, timelapse: bool = False) → Dict[str, str][source]

Parse a merged-stack name into the well coordinates the database stores.

spacr.schema.parse_field_stem() is the parse; this function adds the mapping passthrough and the “refuse rather than guess” errors a resume needs. It used to be a hand-rolled copy of spacr.utils._map_wells, because importing spacr.utils would pull torch and cellpose into every process that merely wants to know whether it can skip a field — spacr.schema is stdlib-only precisely so that reason no longer forces a copy, and tests/test_resume.py still asserts this module imports nothing heavy.

The copy had drifted, in the direction that matters most here: it gave 'plate1_A_3' the identity ('r1', 'c0') while _map_wells wrote 'error' into the database for the same name, so a resume computed a delete key that matched no rows and the field was measured twice. 'AA01' was the same story. Both now raise, which plan_measure_resume() reports rather than silently mis-deleting.

Parameters:
  • field – 'plate1_A01_3', 'plate1_A01_3.npy', or an already-parsed identity mapping (returned filtered, so callers can pass either).

  • timelapse – also parse the 4th component as timeID.

Returns:

dict with plateID / rowID / columnID / fieldID (and timeID when timelapse). Values match what _merge_and_save_to_database wrote, e.g. {'plateID': 'plate1', 'rowID': 'r1', 'columnID': 'c1', 'fieldID': 'f3'}.

Raises:

ValueError – when the name has too few underscore-separated parts to identify a field, or when the well cannot be read. Guessing here would produce a delete key that matches the wrong rows. (spacr.schema.SchemaError is a ValueError, so callers guarding on ValueError keep working.)

Example

>>> field_identity('plate1_A01_3')['fieldID']
'f3'
spacr.resume.format_resume(state: ResumeState, max_examples: int = 4) → str[source]

Render the plan as the block printed before any work starts.

Reports the four numbers that decide whether the resume is doing what the user thinks: how many fields exist, how many are being skipped, how many will run, and — the one that matters — how many were found on disk and rejected anyway, with the reason.

Parameters:
  • state – the plan from plan_resume().

  • max_examples – field names shown per rejection reason.

Returns:

a multi-line string ready to print.

spacr.resume.identity_to_prcf(identity: Mapping[str, str]) → str[source]

Render an identity dict as the prcf string the tables also carry.

Parameters:

identity – canonical field-identity components to join.

Used only for display and for the “no candidate list” mode of completed_fields_in_db(); never as a delete key.

spacr.resume.importer_owns_table(conn: sqlite3.Connection, table: str) → bool[source]

True when every column of table is one the importer recorded.

The table-scoped question, kept here beside its row-scoped twin so the two cannot drift: spacr.foreign._importer_owns calls this rather than repeating it. A table a measure_crop has appended to has grown columns the importer never wrote, so it stops being the importer’s the moment spaCR measures into it.

Parameters:
  • conn – open connection to measurements.db.

  • table – table name.

Returns:

False for a table with no recorded import, for one whose provenance cannot be read, and for one that has since been widened by another writer.

spacr.resume.importer_recorded_columns(conn: sqlite3.Connection, table: str) → Set[str] | None[source]

Every column spacr.foreign recorded itself as writing into table.

Read from FOREIGN_COLUMNS_TABLE and from nothing else. Every release of the importer has written it, including the first, so no marker needs inventing here — a marker would only protect databases written after it was added, and the ones most likely to hold an imported canonical table are the oldest.

Unreadable provenance answers None, deliberately: this is what importer_owns_table() — and through it foreign._importer_owns — consults before dropping and rewriting a canonical table, and there the protective answer is “not ours”. importer_written_columns() is the other direction of the same question and has the opposite fallback for the opposite reason.

Parameters:
  • conn – open connection to measurements.db.

  • table – table name, e.g. 'cell'.

Returns:

the recorded column names, or None when this database records no import into table at all — in which case every row in it is the measure stage’s, which is the ordinary case and the one that must stay exactly as cheap as it was.

Example

# a purely-imported project
importer_recorded_columns(conn, 'cell')
# {'object_label', 'plateID', …, 'foreign_areashape_area'}
importer_recorded_columns(conn, 'nucleus')      # None
spacr.resume.importer_rows_clause(conn: sqlite3.Connection, table: str) → str | None[source]

The complement of measure_rows_clause() — the importer’s rows.

Parameters:
  • conn – open connection to measurements.db.

  • table – table name.

Returns:

None when nothing in table is the importer’s, else a SQL condition selecting exactly the rows that are.

spacr.resume.importer_written_columns(conn: sqlite3.Connection, table: str) → Set[str] | None[source]

The widest set of columns this database shows an import having written.

importer_recorded_columns() first; failing that, the columns of foreign_<table>. The canonical copy is written from the same frame as foreign_<table>, column for column, so that table is a faithful second record of what the importer put in this one — and on a database whose provenance an older importer replaced wholesale when a second object type was imported into it, it is the only record left.

The fallback exists because this is the answer a delete is gated on, and there the protective answer is the opposite of the one importer_recorded_columns() must give: a database whose provenance cannot be read must have its imported rows preserved, not deleted as though nobody had claimed them. That was the original bug, and it was still live for exactly the files most exposed to it.

Parameters:
  • conn – open connection to measurements.db.

  • table – table name, e.g. 'cell'.

Returns:

the column names, or None when nothing in this database suggests an import ever wrote into table.

spacr.resume.measure_rows_clause(conn: sqlite3.Connection, table: str) → str | None[source]

A SQL condition selecting the rows of table the measure stage wrote.

The row-scoped ownership test, as a fragment to be ANDed into a WHERE clause. A row is the measure stage’s when it holds a value in at least one column the importer never wrote — importer_written_columns() being what “never wrote” is read from, so a database whose provenance was replaced away still gets the protective answer.

Parameters:
  • conn – open connection to measurements.db.

  • table – table name.

Returns:

None when nothing in this database suggests an import ever wrote into table — meaning no condition is needed and none should be added, so a project that has never seen an import runs exactly the statements it always did. '0' when the importer wrote every column the table has, i.e. nothing in it is measure’s. Otherwise a parenthesised "col" IS NOT NULL OR … over the columns the importer did not write.

The residual, stated because it is a real if unreachable one: a measure row whose every measurement column is NULL is indistinguishable from an imported one and is treated as imported — preserved rather than cleared. _merge_and_save_to_database always writes at least the object’s area and one intensity column, so producing such a row means producing a measurement that holds no measurements; and being wrong in that direction preserves data rather than deleting it.

spacr.resume.measurements_db_path(settings: Mapping[str, Any]) → str[source]

Where measure_crop writes, derived exactly as it derives it.

settings['src'] points at <experiment>/merged by the time the resume runs, and the database is <experiment>/measurements/ measurements.db — the same os.path.dirname hop _measure_crop_core and _save_settings_to_db make.

spacr.resume.plan_measure_resume(settings: Any, verbose: bool = True) → ResumeState | None[source]

Plan (and make safe) a resumed measure_crop. The one call measure makes.

Returns None immediately unless settings['resume'] is set, so a default run does exactly what it always did.

When resume is requested, in order:

  1. Read the previous run’s settings out of the settings table in measurements.db and refuse loudly if anything material differs. This must happen before spacr.io._save_settings_to_db overwrites that table with the current settings — hence the placement of the call in measure_crop.

  2. Enumerate the fields in merged/ and validate each one: a truncated .npy from a crash mid-np.save is pending, not done, and is counted as rejected.

  3. Ask the database which fields already have rows in every table the run writes. Fields with rows in only some tables are partial — a process killed between two inserts — and are re-queued.

  4. Delete every row belonging to a pending field before it is re-run. This is what stops the resume from doubling rows.

  5. Print the plan.

Parameters:
  • settings – the measure settings dict.

  • verbose – print the plan block.

Returns:

a ResumeState whose ResumeState.filter_files() turns the caller’s file list into the pending subset, or None when resume is off.

Raises:

ResumeRefused – when the recorded settings differ materially.

spacr.resume.plan_resume(all_fields: Iterable[str], done: Iterable[str], reasons: Mapping[str, str] | None = None, enabled: bool = True, src: str = '', db_path: str = '', cleared_rows: int = 0) → ResumeState[source]

Work out what to run, what to skip, and record why.

Nothing is done here — that is the point. The plan is computed and reported first, so a resume that is about to skip 900 fields says so before it skips them, and a resume that rejected three truncated files says that too.

Parameters:
  • all_fields – every candidate field stem (or filename), in the order they should be processed.

  • done – field stems verified complete and safe to skip.

  • reasons – known {field: reason} entries, e.g. from completed_fields_in_merged()’s rejection dict.

  • enabled – when False, every field is pending and none is skipped — the default-behaviour path.

  • src – merged folder, recorded for the report.

  • db_path – database, recorded for the report.

  • cleared_rows – rows already deleted by the caller’s delete-before-insert, recorded for the report.

Returns:

a ResumeState.

spacr.resume.read_npy_header(path: str) → Dict[str, Any][source]

Parse a .npy header without loading (or allocating) the array.

Reading the header alone is what makes validating a thousand 100-megabyte fields cheap enough to do on every resume — and it is strictly better at catching truncation than np.load, which has to allocate the full array before it discovers the file is short.

Parameters:

path – path to a .npy file.

Returns:

dict with shape (tuple), descr, fortran_order, header_end (byte offset of the data), itemsize (None when the dtype is not a simple numeric one), expected_bytes (None when itemsize is unknown) and actual_bytes.

Raises:

ValueError – when the file is empty, or is not a .npy at all, or its header is unparseable.

Example

>>> read_npy_header('merged/plate1_A01_3.npy')['shape']
(1080, 1080, 7)
spacr.resume.read_recorded_settings(source: str) → Dict[str, Any][source]

Read the settings a previous run wrote, from wherever it wrote them.

Three shapes are understood, which between them cover both stages:

  • a measurements.db — read from the settings table (setting_key / setting_value) that spacr.io._save_settings_to_db writes at the top of every measure_crop. Note that call happens before field enumeration and uses if_exists='replace', so a resume must read this table before the current run overwrites it.

  • a Key,Value CSV — <src>/settings/gen_mask_settings.csv or measure_crop_settings.csv, as written by spacr.utils.save_settings.

  • a run-journal folder containing settings.json.

Parameters:

source – path to any of the above.

Returns:

the recorded settings dict, or {} when nothing is recorded (an artifact that predates settings capture — which must read as “no information”, not as “everything matches”).

spacr.resume.resume_enabled(settings: Any) → bool[source]

True when the caller asked for a resume.

Resume is off unless explicitly requested, so a settings dict that has never heard of the key behaves exactly as it always did.

Parameters:

settings – a spaCR settings dict (or anything else, which reads as “off”).

spacr.resume.run_already_complete(db_path: str, name: str | None = None) → bool[source]

True when db_path carries a RunLedger stamp saying so.

A run that finished cleanly needs no resume at all, and this is the cheapest way to know: measure_crop ends with ledger.finalize(artifact=db_path), which appends a run_status row recording how many fields were attempted and how many failed.

Note this is stricter than spacr.errors.run_is_complete(), which reads an unstamped artifact as complete (stamping is newer than most data on disk). For a resume, “nobody ever said” must mean “go and check the files”, not “nothing to do”.

Parameters:
  • db_path – path to measurements.db.

  • name – only consider stamps from this ledger, e.g. 'measure_crop'. None considers the most recent stamp of any stage.

Returns:

True only when a matching stamp exists and its latest entry recorded zero failures over at least one attempted item.

spacr.resume.supersede_imported_copies(db_path: str, tables: Sequence[str], pending: Iterable[str], timelapse: bool = False) → Tuple[int, List[str]][source]

Release an import’s convenience copy that this run is about to replace.

foreign.run_import copies the imported rows into the canonical cell / nucleus / pathogen table when nothing of anyone else’s is there, so that a purely-imported project is readable by every spaCR tool. measure_crop appends, so the moment spaCR measures the same fields its rows land in that table beside theirs, in different columns, with nothing to tell a reader which population a row belongs to — and count_cell becomes the sum of the two.

The copy is released here, before any of that is written, and only when releasing it is complete:

  • every field the copy covers must either already hold rows this measure stage wrote in that table, or be queued to be measured now. A field that is neither would be left with nothing at all, so the copy stays whole and the caller is told what to do instead — a table that is half released is worse than one that is not;

  • the removal must be lossless, which spacr.foreign.release_canonical_copy() verifies row by row against foreign_<object> before deleting anything.

Nothing here can lose a measurement: what is removed is a duplicate of foreign_<object>, which no resume may touch, and the <object>_with_foreign view is created so it is still one query away.

Parameters:
  • db_path – path to measurements.db.

  • tables – the measure-owned tables discovered for this run.

  • pending – field stems this run is about to measure.

  • timelapse – when true nothing is released. The importer writes no timeID, so a timelapse project cannot have a copy this function would understand, and guessing at one would be a delete keyed on fewer columns than the writer used.

Returns:

(rows released, notes). The notes are for the resume report and are non-empty exactly when a user has to act.

Example

released, notes = supersede_imported_copies(
    db, ['cell'], ['plate1_A01_1', 'plate1_A01_2'])
# (4, [])  -> `cell` is spaCR's to fill; theirs are unchanged
#             in foreign_cell, joined by cell_with_foreign
spacr.resume.validate_merged_field(path: str, min_planes: int | None = None) → Tuple[bool, str][source]

Decide whether a merged/*.npy is a finished field or a crash scar.

Parameters:
  • path – path to the candidate .npy.

  • min_planes – when given, the array’s last axis must be at least this large. Derive it from the settings with expected_min_planes() — a stack with too few planes cannot be measured with the requested cell_mask_dim / channels regardless of whether its bytes are all there.

Returns:

(ok, reason). reason is REASON_DONE when ok, otherwise one of REASON_MISSING, REASON_EMPTY, REASON_TRUNCATED, REASON_UNREADABLE, REASON_TOO_FEW_PLANES.

Example

>>> validate_merged_field('merged/plate1_A01_3.npy', min_planes=7)
(True, 'done')