spacr.pipeline_v2

Streaming mask pipeline (v2).

Replaces the multi-copy disk chain that preprocess_generate_masks has run since day one:

originals

→ renamed + split into channel folders → orig/ backup → per-channel npy → batch npz on disk → cellpose → per-field mask npy → concatenated into merged/

…with a two-pass streaming pipeline that keeps only what the downstream measure module actually reads:

Pass 1 — assemble

walk originals, parse metadata regex, build one npy stack per field with all image channels in the C axis. Emit filename_map.csv recording every original → stack mapping.

Pass 2 — segment

stream the plate in batches of N fields, hand each batch to Cellpose, append the mask channels to the SAME stack file. Each batch is written to a scratch NPZ on the way through and deleted again unless keep_npz=True.

Output — merged/ folder holds one file per field, each shape (H, W, C_image + C_mask) in uint16, plus:

channel_order.json {“channels”: […]} filename_map.csv original path, plate/well/field/…, stack idx

Public API:

from spacr.pipeline_v2 import (
    FilenameMapper, stream_originals_to_stack,
    stream_masks_from_stack, run_v2,
)

# High-level (one call):
run_v2(src_folder, channels=(0,1,2,3), model="cyto", diameter=60)

# Low-level (two passes, run each explicitly):
mapper = FilenameMapper.discover(src_folder,
                                   metadata_type="cellvoyager")
stacks = stream_originals_to_stack(src_folder, mapper, channels=(0,1,2,3))
stream_masks_from_stack(stacks, model="cyto", diameter=60)

This module is opt-in for one release cycle. Once the follow-up commit wires it as the default in spacr.core.preprocess_generate_masks() the whole disk chain above collapses to merged/ alone.

Classes

FilenameMapper

Walks a folder of microscopy images, parses each filename's

FilenameRecord

One entry in the filename map.

StackFile

One field's on-disk stack: merged/stack_<id>.npy with

Functions

run_v2(, channel_names, model_name, ...)

Run the entire v2 pipeline against src. Convenience wrapper.

stream_masks_from_stack(, diameter, batch_fields, ...)

Batch the field stacks through Cellpose, then append the mask

stream_originals_to_stack(, channel_names, dst)

Write one merged/stack_<field>.npy per field.

Module Contents

class spacr.pipeline_v2.FilenameMapper(records: List[FilenameRecord], metadata_type: str, regex: str)[source]

Walks a folder of microscopy images, parses each filename’s metadata via a regex, and records the mapping to a per-plate CSV.

The CSV is written next to the merged/ folder (at the plate root) so users can Excel-open filename_map.csv and see the original path of every image in the run.

Parameters:
  • records – parsed filename records in file-system order.

  • metadata_type – name of the metadata convention used to parse them.

  • regex – regular-expression source used for parsing.

Variables:
  • records – list of FilenameRecord in file-system order.

  • metadata_type – which regex was used ("cellvoyager" / "yokogawa" / "custom").

  • regex – compiled regex pattern that matched.

Store parsed filename records and the metadata rule that made them.

by_field() → Dict[str, List[FilenameRecord]][source]

Group records by stack_field_id — one entry per field, with one record per channel inside.

classmethod discover(src: pathlib.Path, metadata_type: str = 'auto', custom_regex: str | None = None, exts: Sequence[str] = ('.tif', '.tiff', '.png', '.jpg', '.jpeg')) → FilenameMapper[source]

Scan src for images + parse each name with the metadata regex. Falls back through cellvoyager → yokogawa on metadata_type="auto".

Parameters:
  • src – folder to scan (not recursive; we expect images at the top level as the current spacr layout does).

  • metadata_type – "auto" / "cellvoyager" / "yokogawa" / "custom". When "custom", custom_regex must be given.

  • custom_regex – user-supplied regex; required for metadata_type="custom".

  • exts – image file extensions to include.

Returns:

a populated FilenameMapper.

Raises:

ValueError – when no images are found or no regex fits.

field_ids() → List[str][source]

Return the sorted list of unique stack_field_id values.

classmethod load_csv(path: pathlib.Path) → FilenameMapper[source]

Rehydrate a mapper from a previously-saved CSV.

Parameters:

path – mapping CSV previously written by save_csv().

save_csv(path: pathlib.Path) → pathlib.Path[source]

Write the mapping to path as a CSV that Excel opens cleanly. One row per (original image, resulting stack slot).

Parameters:

path – destination CSV path; its parent directory is created.

class spacr.pipeline_v2.FilenameRecord[source]

One entry in the filename map.

Variables:
  • original_path – absolute path to the source image on disk.

  • plate – plate id parsed from the filename.

  • well – well id parsed from the filename.

  • field – field index parsed from the filename.

  • channel – channel index parsed from the filename.

  • time – time index parsed from the filename (defaults to 1).

  • z – z-slice index parsed from the filename (defaults to 1).

  • stack_field_id – the field id used in merged/stack_<X>.npy.

class spacr.pipeline_v2.StackFile[source]

One field’s on-disk stack: merged/stack_<id>.npy with shape (H, W, C).

Populated by stream_originals_to_stack() before Cellpose runs (C = image channels only). After stream_masks_from_stack() the same file has additional mask channels appended.

Variables:
  • field_id – stable field identifier used in the stack filename.

  • path – path to the on-disk NumPy stack.

  • shape – (height, width, channels) shape at write time.

  • channels – human-readable channel names in array order.

spacr.pipeline_v2.run_v2(src: pathlib.Path, channels: Sequence[int] = (0, 1, 2, 3), channel_names: Sequence[str] | None = None, model_name: str = 'cyto', channels_for_cellpose: Sequence[int] = (0, 0), diameter: float | None = None, batch_fields: int = 8, metadata_type: str = 'auto', custom_regex: str | None = None, keep_npz: bool = False, cellprob_threshold: float = 0.0, flow_threshold: float = 0.4, min_size: int = 15, resample: bool = True, postprocess_settings: Dict[str, Any] | None = None, object_type: str = 'cell', illumination_settings: Dict[str, Any] | None = None) → Dict[str, Any][source]

Run the entire v2 pipeline against src. Convenience wrapper.

Equivalent to:

mapper = FilenameMapper.discover(src, metadata_type, custom_regex)
stacks = stream_originals_to_stack(src, mapper, channels, channel_names)
stream_masks_from_stack(stacks, model_name, channels_for_cellpose,
                        diameter, batch_fields, keep_npz=keep_npz)
Parameters:
  • src – plate folder holding the originals, scanned non-recursively. filename_map.csv lands here and output goes to <src>/merged; neither path is overridable from this wrapper.

  • channels – channel numbers as parsed from the filename (C01 gives 1), not C-axis positions. The default (0, 1, 2, 3) therefore misfits stock CellVoyager/Yokogawa names: channel 0 never exists, so plane 0 is all zeros, and C04 is dropped. Pass (1, 2, 3, 4) for those layouts.

  • channel_names – names recorded in channel_order.json and on each StackFile. A length mismatch with channels trips a bare assert, so it goes unchecked under python -O.

  • model_name – resolved by spacr.utils._resolve_cellpose_pretrained(), not passed as Cellpose’s model_type. On Cellpose 4 every legacy name ("cyto", "nuclei", …) collapses to cpsam, so only a fine-tuned checkpoint path actually changes the weights, and a path with no file behind it raises instead of falling back.

  • channels_for_cellpose – despite the name this never reaches Cellpose’s channels= argument; it selects C-axis indices from the assembled stack, taken modulo the channel count (7 with C=4 becomes 3) and de-duplicated in order. The default (0, 0) collapses to one plane, and an empty sequence falls back to plane 0.

  • diameter – forwarded to model.eval; None leaves Cellpose to size objects itself. Unlike model_name it is still honoured on Cellpose 4, which rescales the image by 30 / diameter.

  • batch_fields – fields loaded and segmented per batch, the memory-versus-speed dial. 0 raises ValueError; a negative value segments nothing at all, yet channel_order.json is still rewritten to claim a mask channel that no stack received.

  • metadata_type – "cellvoyager", "yokogawa", "custom" or "auto". Only "custom" is special-cased; every other unrecognised string quietly behaves as "auto" instead of raising.

  • custom_regex – required when metadata_type="custom", else ValueError. It must supply the named groups (plateID, wellID, fieldID, chanID, timeID, sliceID); any group it omits silently defaults, so a regex without chanID makes every image channel 1 and therefore its own single-plane field.

  • keep_npz – the per-batch NPZ is compressed into merged/_scratch either way — this only decides whether it and the scratch folder survive, so False does not save the write.

  • cellprob_threshold – forwarded to model.eval through float().

  • flow_threshold – forwarded to model.eval through float().

  • min_size – forwarded to model.eval through int(); None raises TypeError rather than meaning “no minimum”.

  • resample – coerced with bool(), so any non-empty string — "false" included — is True, while None is False.

  • postprocess_settings – None hands the raw selected planes to Cellpose and skips post-processing entirely. Any dict, {} included, switches on both spacr.io._normalize_img_batch() and spacr.object.merge_split_filter_masks(). The four *_channel role keys are rewritten on a copy (the caller’s dict is left alone): all cleared to None, then object_type set to 0, plus nucleus_channel=1 when object_type is "cell" and two or more planes were selected.

  • object_type – picks the weights during model_name resolution, names the role given channel 0 in normalisation, and is passed to the mask post-processor. An unrecognised value does not raise — it adds a dead <value>_channel key and leaves every real role unset.

  • illumination_settings – full Mask settings mapping. When it enables illumination_correction, one model is prepared from the raw persisted stacks after pass 1 and its session corrects only the private Cellpose inputs in pass 2; omitted/off preserves the previous byte-level output contract.

Returns:

dict with mapper (FilenameMapper), stacks (list of StackFile), and dst (Path to merged/).

Raises:

ValueError – from FilenameMapper.discover() when src holds no images, when metadata_type="custom" has no custom_regex, or from batch_fields=0.

spacr.pipeline_v2.stream_masks_from_stack(stacks: List[StackFile], model_name: str = 'cyto', channels_for_cellpose: Sequence[int] = (0, 0), diameter: float | None = None, batch_fields: int = 8, mask_channel_name: str = 'mask', keep_npz: bool = False, npz_dir: pathlib.Path | None = None, cellprob_threshold: float = 0.0, flow_threshold: float = 0.4, min_size: int = 15, resample: bool = True, postprocess_settings: Dict[str, Any] | None = None, object_type: str = 'cell', illumination_session: Any | None = None, psf_session: Any | None = None) → List[StackFile][source]

Batch the field stacks through Cellpose, then append the mask channel(s) to the SAME npy files.

Parameters:
  • stacks – list produced by stream_originals_to_stack().

  • model_name – Cellpose model to use ("cyto", "nuclei", …).

  • channels_for_cellpose – C-axis indices selected out of each assembled stack before it is handed to Cellpose — taken modulo the channel count and de-duplicated in order. It is NOT forwarded as Cellpose’s channels= argument; [0, 0] therefore yields a single plane, not a grayscale pair. The first selected channel is the object’s own channel for absolute mean-intensity filtering.

  • diameter – expected object diameter in px (None → Cellpose auto).

  • batch_fields – how many field stacks to load into memory at once. Larger = faster but more RAM.

  • mask_channel_name – human name to record for the appended mask channel (default "mask").

  • keep_npz – the intermediate batch is compressed to an NPZ under npz_dir on every batch regardless; this flag only decides whether that file and the scratch folder survive the run.

  • npz_dir – where to write the (optional) intermediate NPZ files. Defaults to a scratch subfolder under the stack folder.

  • illumination_session – optional spacr.illumination.SegmentationIlluminationSession. Its corrector sees private selected-channel copies immediately before normalisation/Cellpose; persisted intensity planes and scratch NPZs remain raw, and completion is recorded only after the combined stack has been atomically replaced.

  • psf_session – optional captured PSF session. Unmixes each whole field first when unmixing is on, then processes selected intensities after illumination and before normalization, padding or Cellpose. Stored image channels stay raw; only the appended labels depend on PSF processing.

Returns:

the same list, with each StackFile.shape / .channels updated to reflect the appended mask channel.

spacr.pipeline_v2.stream_originals_to_stack(src: pathlib.Path, mapper: FilenameMapper, channels: Sequence[int] = (0, 1, 2, 3), channel_names: Sequence[str] | None = None, dst: pathlib.Path | None = None) → List[StackFile][source]

Write one merged/stack_<field>.npy per field.

Reads originals directly (no rename-into-channel-folders step), stacks the selected channels along the C axis, and writes one npy per field. Also emits a channel_order.json sidecar describing which C-index holds which channel.

Parameters:
  • src – plate folder containing the original images.

  • mapper – FilenameMapper produced from src.

  • channels – which channel numbers (as parsed from filenames) to include, in the order they should occupy the C axis.

  • channel_names – human names for those channels (must match channels length). Default: ["ch0", "ch1", …].

  • dst – override the output folder; defaults to <src>/merged.

Returns:

list of StackFile, one per field written.