spacr.pipeline_v2¶
Streaming mask pipeline (v2).
Replaces the multi-copy disk chain that preprocess_generate_masks
has run since day one:
- originals
→ renamed + split into channel folders → orig/ backup → per-channel npy → batch npz on disk → cellpose → per-field mask npy → concatenated into merged/
…with a two-pass streaming pipeline that keeps only what the downstream measure module actually reads:
- Pass 1 — assemble
walk originals, parse metadata regex, build one npy stack per field with all image channels in the C axis. Emit
filename_map.csvrecording every original → stack mapping.- Pass 2 — segment
stream the plate in batches of N fields, hand each batch to Cellpose, append the mask channels to the SAME stack file. Each batch is written to a scratch NPZ on the way through and deleted again unless
keep_npz=True.
Output — merged/ folder holds one file per field, each shape
(H, W, C_image + C_mask) in uint16, plus:
channel_order.json {“channels”: […]} filename_map.csv original path, plate/well/field/…, stack idx
Public API:
from spacr.pipeline_v2 import (
FilenameMapper, stream_originals_to_stack,
stream_masks_from_stack, run_v2,
)
# High-level (one call):
run_v2(src_folder, channels=(0,1,2,3), model="cyto", diameter=60)
# Low-level (two passes, run each explicitly):
mapper = FilenameMapper.discover(src_folder,
metadata_type="cellvoyager")
stacks = stream_originals_to_stack(src_folder, mapper, channels=(0,1,2,3))
stream_masks_from_stack(stacks, model="cyto", diameter=60)
This module is opt-in for one release cycle. Once the follow-up commit
wires it as the default in spacr.core.preprocess_generate_masks()
the whole disk chain above collapses to merged/ alone.
Classes¶
Walks a folder of microscopy images, parses each filename's |
|
One entry in the filename map. |
|
One field's on-disk stack: |
Functions¶
|
Run the entire v2 pipeline against |
|
Batch the field stacks through Cellpose, then append the mask |
|
Write one |
Module Contents¶
- class spacr.pipeline_v2.FilenameMapper(records: List[FilenameRecord], metadata_type: str, regex: str)[source]¶
Walks a folder of microscopy images, parses each filename’s metadata via a regex, and records the mapping to a per-plate CSV.
The CSV is written next to the
merged/folder (at the plate root) so users can Excel-openfilename_map.csvand see the original path of every image in the run.- Parameters:
records – parsed filename records in file-system order.
metadata_type – name of the metadata convention used to parse them.
regex – regular-expression source used for parsing.
- Variables:
records – list of
FilenameRecordin file-system order.metadata_type – which regex was used (
"cellvoyager"/"yokogawa"/"custom").regex – compiled regex pattern that matched.
Store parsed filename records and the metadata rule that made them.
- by_field() Dict[str, List[FilenameRecord]][source]¶
Group records by
stack_field_id— one entry per field, with one record per channel inside.
- classmethod discover(src: pathlib.Path, metadata_type: str = 'auto', custom_regex: str | None = None, exts: Sequence[str] = ('.tif', '.tiff', '.png', '.jpg', '.jpeg')) FilenameMapper[source]¶
Scan
srcfor images + parse each name with the metadata regex. Falls back throughcellvoyager→yokogawaonmetadata_type="auto".- Parameters:
src – folder to scan (not recursive; we expect images at the top level as the current spacr layout does).
metadata_type –
"auto"/"cellvoyager"/"yokogawa"/"custom". When"custom",custom_regexmust be given.custom_regex – user-supplied regex; required for
metadata_type="custom".exts – image file extensions to include.
- Returns:
a populated
FilenameMapper.- Raises:
ValueError – when no images are found or no regex fits.
- classmethod load_csv(path: pathlib.Path) FilenameMapper[source]¶
Rehydrate a mapper from a previously-saved CSV.
- Parameters:
path – mapping CSV previously written by
save_csv().
- save_csv(path: pathlib.Path) pathlib.Path[source]¶
Write the mapping to
pathas a CSV that Excel opens cleanly. One row per (original image, resulting stack slot).- Parameters:
path – destination CSV path; its parent directory is created.
- class spacr.pipeline_v2.FilenameRecord[source]¶
One entry in the filename map.
- Variables:
original_path – absolute path to the source image on disk.
plate – plate id parsed from the filename.
well – well id parsed from the filename.
field – field index parsed from the filename.
channel – channel index parsed from the filename.
time – time index parsed from the filename (defaults to 1).
z – z-slice index parsed from the filename (defaults to 1).
stack_field_id – the
fieldid used inmerged/stack_<X>.npy.
- class spacr.pipeline_v2.StackFile[source]¶
One field’s on-disk stack:
merged/stack_<id>.npywith shape(H, W, C).Populated by
stream_originals_to_stack()before Cellpose runs (C = image channels only). Afterstream_masks_from_stack()the same file has additional mask channels appended.- Variables:
field_id – stable field identifier used in the stack filename.
path – path to the on-disk NumPy stack.
shape –
(height, width, channels)shape at write time.channels – human-readable channel names in array order.
- spacr.pipeline_v2.run_v2(src: pathlib.Path, channels: Sequence[int] = (0, 1, 2, 3), channel_names: Sequence[str] | None = None, model_name: str = 'cyto', channels_for_cellpose: Sequence[int] = (0, 0), diameter: float | None = None, batch_fields: int = 8, metadata_type: str = 'auto', custom_regex: str | None = None, keep_npz: bool = False, cellprob_threshold: float = 0.0, flow_threshold: float = 0.4, min_size: int = 15, resample: bool = True, postprocess_settings: Dict[str, Any] | None = None, object_type: str = 'cell', illumination_settings: Dict[str, Any] | None = None) Dict[str, Any][source]¶
Run the entire v2 pipeline against
src. Convenience wrapper.Equivalent to:
mapper = FilenameMapper.discover(src, metadata_type, custom_regex) stacks = stream_originals_to_stack(src, mapper, channels, channel_names) stream_masks_from_stack(stacks, model_name, channels_for_cellpose, diameter, batch_fields, keep_npz=keep_npz)
- Parameters:
src – plate folder holding the originals, scanned non-recursively.
filename_map.csvlands here and output goes to<src>/merged; neither path is overridable from this wrapper.channels – channel numbers as parsed from the filename (
C01gives 1), not C-axis positions. The default(0, 1, 2, 3)therefore misfits stock CellVoyager/Yokogawa names: channel 0 never exists, so plane 0 is all zeros, andC04is dropped. Pass(1, 2, 3, 4)for those layouts.channel_names – names recorded in
channel_order.jsonand on eachStackFile. A length mismatch withchannelstrips a bareassert, so it goes unchecked underpython -O.model_name – resolved by
spacr.utils._resolve_cellpose_pretrained(), not passed as Cellpose’smodel_type. On Cellpose 4 every legacy name ("cyto","nuclei", …) collapses tocpsam, so only a fine-tuned checkpoint path actually changes the weights, and a path with no file behind it raises instead of falling back.channels_for_cellpose – despite the name this never reaches Cellpose’s
channels=argument; it selects C-axis indices from the assembled stack, taken modulo the channel count (7 with C=4 becomes 3) and de-duplicated in order. The default(0, 0)collapses to one plane, and an empty sequence falls back to plane 0.diameter – forwarded to
model.eval;Noneleaves Cellpose to size objects itself. Unlikemodel_nameit is still honoured on Cellpose 4, which rescales the image by30 / diameter.batch_fields – fields loaded and segmented per batch, the memory-versus-speed dial.
0raisesValueError; a negative value segments nothing at all, yetchannel_order.jsonis still rewritten to claim a mask channel that no stack received.metadata_type –
"cellvoyager","yokogawa","custom"or"auto". Only"custom"is special-cased; every other unrecognised string quietly behaves as"auto"instead of raising.custom_regex – required when
metadata_type="custom", elseValueError. It must supply the named groups (plateID,wellID,fieldID,chanID,timeID,sliceID); any group it omits silently defaults, so a regex withoutchanIDmakes every image channel 1 and therefore its own single-plane field.keep_npz – the per-batch NPZ is compressed into
merged/_scratcheither way — this only decides whether it and the scratch folder survive, soFalsedoes not save the write.cellprob_threshold – forwarded to
model.evalthroughfloat().flow_threshold – forwarded to
model.evalthroughfloat().min_size – forwarded to
model.evalthroughint();NoneraisesTypeErrorrather than meaning “no minimum”.resample – coerced with
bool(), so any non-empty string —"false"included — is True, whileNoneis False.postprocess_settings –
Nonehands the raw selected planes to Cellpose and skips post-processing entirely. Any dict,{}included, switches on bothspacr.io._normalize_img_batch()andspacr.object.merge_split_filter_masks(). The four*_channelrole keys are rewritten on a copy (the caller’s dict is left alone): all cleared to None, thenobject_typeset to 0, plusnucleus_channel=1whenobject_typeis"cell"and two or more planes were selected.object_type – picks the weights during
model_nameresolution, names the role given channel 0 in normalisation, and is passed to the mask post-processor. An unrecognised value does not raise — it adds a dead<value>_channelkey and leaves every real role unset.illumination_settings – full Mask settings mapping. When it enables
illumination_correction, one model is prepared from the raw persisted stacks after pass 1 and its session corrects only the private Cellpose inputs in pass 2; omitted/off preserves the previous byte-level output contract.
- Returns:
dict with
mapper(FilenameMapper),stacks(list ofStackFile), anddst(Path tomerged/).- Raises:
ValueError – from
FilenameMapper.discover()whensrcholds no images, whenmetadata_type="custom"has nocustom_regex, or frombatch_fields=0.
- spacr.pipeline_v2.stream_masks_from_stack(stacks: List[StackFile], model_name: str = 'cyto', channels_for_cellpose: Sequence[int] = (0, 0), diameter: float | None = None, batch_fields: int = 8, mask_channel_name: str = 'mask', keep_npz: bool = False, npz_dir: pathlib.Path | None = None, cellprob_threshold: float = 0.0, flow_threshold: float = 0.4, min_size: int = 15, resample: bool = True, postprocess_settings: Dict[str, Any] | None = None, object_type: str = 'cell', illumination_session: Any | None = None, psf_session: Any | None = None) List[StackFile][source]¶
Batch the field stacks through Cellpose, then append the mask channel(s) to the SAME npy files.
- Parameters:
stacks – list produced by
stream_originals_to_stack().model_name – Cellpose model to use (
"cyto","nuclei", …).channels_for_cellpose – C-axis indices selected out of each assembled stack before it is handed to Cellpose — taken modulo the channel count and de-duplicated in order. It is NOT forwarded as Cellpose’s
channels=argument;[0, 0]therefore yields a single plane, not a grayscale pair. The first selected channel is the object’s own channel for absolute mean-intensity filtering.diameter – expected object diameter in px (None → Cellpose auto).
batch_fields – how many field stacks to load into memory at once. Larger = faster but more RAM.
mask_channel_name – human name to record for the appended mask channel (default
"mask").keep_npz – the intermediate batch is compressed to an NPZ under
npz_diron every batch regardless; this flag only decides whether that file and the scratch folder survive the run.npz_dir – where to write the (optional) intermediate NPZ files. Defaults to a scratch subfolder under the stack folder.
illumination_session – optional
spacr.illumination.SegmentationIlluminationSession. Its corrector sees private selected-channel copies immediately before normalisation/Cellpose; persisted intensity planes and scratch NPZs remain raw, and completion is recorded only after the combined stack has been atomically replaced.psf_session – optional captured PSF session. Unmixes each whole field first when unmixing is on, then processes selected intensities after illumination and before normalization, padding or Cellpose. Stored image channels stay raw; only the appended labels depend on PSF processing.
- Returns:
the same list, with each
StackFile.shape/.channelsupdated to reflect the appended mask channel.
- spacr.pipeline_v2.stream_originals_to_stack(src: pathlib.Path, mapper: FilenameMapper, channels: Sequence[int] = (0, 1, 2, 3), channel_names: Sequence[str] | None = None, dst: pathlib.Path | None = None) List[StackFile][source]¶
Write one
merged/stack_<field>.npyper field.Reads originals directly (no rename-into-channel-folders step), stacks the selected channels along the C axis, and writes one npy per field. Also emits a
channel_order.jsonsidecar describing which C-index holds which channel.- Parameters:
src – plate folder containing the original images.
mapper –
FilenameMapperproduced fromsrc.channels – which channel numbers (as parsed from filenames) to include, in the order they should occupy the C axis.
channel_names – human names for those channels (must match
channelslength). Default:["ch0", "ch1", …].dst – override the output folder; defaults to
<src>/merged.
- Returns:
list of
StackFile, one per field written.