spacr.stream_dataset

Create reproducible training datasets from merged image arrays.

The workflow first records a deterministic object selection table, then extracts the selected crops from merged arrays. Object types, channels, crop shape, and train/test allocation are therefore explicit settings rather than properties of a previously exported directory. Crop names use the same naming function as the standard object-crop exporter.

Functions

build_selection(→ Tuple[pandas.DataFrame, str])

Build and save the object-selection table.

coordinate_column(→ str)

Return the identifier column for an object-array type.

crop_name([pathogen_ids])

Return the standard exported-crop name for an object.

cut() → Optional[numpy.ndarray])

Extract one labeled object from an image stack.

selection_from_arrays(→ pandas.DataFrame)

Build a streaming selection from labels in merged arrays.

selection_from_objects(→ pandas.DataFrame)

Build a streaming selection from an object table.

settings_for_method(→ Tuple[str, ...])

Return settings used by a dataset-selection method.

stream(, mask_index, bounding_box, crop_mode[, write])

Write crops described by a saved selection table.

stream_dataset(→ Dict[str, Any])

Build, save, and execute a streamed-dataset selection.

Module Contents

spacr.stream_dataset.build_selection(dst: str, *, objects: pandas.DataFrame | None = None, merged_folder: str = '', object_array: str = 'cell', test_split: float = 0.2, seed: int = 0) → Tuple[pandas.DataFrame, str][source]

Build and save the object-selection table.

An available object table is preferred; otherwise labels are read from merged arrays.

Parameters:

dst (str) – Output folder; created if missing, and stream_selection.csv is written into it.

Returns:

  • pandas.DataFrame – Selected objects and deterministic split assignments.

  • str – Path to the written stream_selection.csv file.

spacr.stream_dataset.coordinate_column(object_array: str) → str[source]

Return the identifier column for an object-array type.

Parameters:

object_array (str) – Object type such as "cell" or "nucleus".

Returns:

str – Canonical identifier column.

Raises:

KeyError – If the object type is unsupported.

spacr.stream_dataset.crop_name(field_stem: str, object_id, *, crop_mode: str = 'cell', nucleus_ids=(), pathogen_ids=(), timelapse: bool = False) → str[source]

Return the standard exported-crop name for an object.

Naming is delegated to spacr.utils._generate_names(), ensuring that streamed and pre-exported crops from the same object are compatible.

Parameters:
  • field_stem (str) – File stem of the image field the object comes from; escaped, it begins the crop name.

  • object_id (int) – Label id of the object; it follows the stem in the name.

spacr.stream_dataset.cut(stack: numpy.ndarray, mask: numpy.ndarray, label: int, *, bounding_box: bool = True, channels: Sequence[int] = ()) → numpy.ndarray | None[source]

Extract one labeled object from an image stack.

Parameters:
  • stack (numpy.ndarray) – Two- or three-dimensional image data.

  • mask (numpy.ndarray) – Integer label mask aligned with the image dimensions.

  • label (int) – Object label to extract.

  • bounding_box (bool, default=True) – Return the complete bounding box when true. When false, pixels outside the selected object are set to zero within the box.

  • channels (sequence of int, optional) – Channel indices to retain. All channels are used when omitted.

Returns:

numpy.ndarray or None – Extracted crop, or None when the label is absent.

spacr.stream_dataset.selection_from_arrays(merged_folder: str, *, object_array: str = 'cell', test_split: float = 0.2, seed: int = 0, mask_index: int | None = None) → pandas.DataFrame[source]

Build a streaming selection from labels in merged arrays.

Parameters:
  • merged_folder (str) – Directory containing merged .npy arrays.

  • object_array (str, default="cell") – Object type represented by the selected mask plane.

  • test_split (float, default=0.2) – Fraction assigned to the test set.

  • seed (int, default=0) – Random seed controlling deterministic split assignment.

  • mask_index (int, optional) – Mask plane index. The final plane is used when omitted.

Returns:

pandas.DataFrame – Canonical identifiers, object type, split, and source array for every nonzero label.

Raises:

FileNotFoundError – If no merged arrays or nonzero object labels are available.

spacr.stream_dataset.selection_from_objects(frame: pandas.DataFrame, *, object_array: str = 'cell', test_split: float = 0.2, seed: int = 0, ask: Any | None = None) → pandas.DataFrame[source]

Build a streaming selection from an object table.

Parameters:
  • frame (pandas.DataFrame) – One row per measured object.

  • object_array (str, default="cell") – Object type to select.

  • test_split (float, default=0.2) – Fraction assigned to the test set.

  • seed (int, default=0) – Random seed controlling deterministic split assignment.

Returns:

pandas.DataFrame – Canonical identifiers, object type, split, and source for each object.

Raises:

ValueError – If the table is empty or contains no usable object identifier, and either nobody was asked or the answer did not help.

Notes

ask is called only when the table names no object, and is INJECTED rather than imported so this module never depends on Qt. A caller with nobody in front of it passes none and gets exactly the error it always got, which makes “never prompt in a batch run” structural rather than a rule each call site has to remember. It is called ask(tried=..., object_array=...) and returns ((database, table, column), why) – see spacr.qt.ask_for_the_path.ask_for_a_database_column().

spacr.stream_dataset.settings_for_method(method: str) → Tuple[str, ...][source]

Return settings used by a dataset-selection method.

Parameters:

method (str) – Selection method name ('column' or 'array'), matched case-insensitively after stripping whitespace.

Raises:

KeyError – If method is unsupported.

spacr.stream_dataset.stream(selection: pandas.DataFrame, merged_folder: str, dst: str, *, channel_arrays: Sequence[int] = (0, 1, 2), mask_index: int | None = None, bounding_box: bool = True, crop_mode: str = 'cell', write=None) → Dict[str, Any][source]

Write crops described by a saved selection table.

Each merged field is loaded once and all selected objects from that field are extracted before advancing. Missing or ambiguous field matches are reported as missing crops; a different field is never used as a fallback.

Parameters:
  • selection (pandas.DataFrame) – Selection produced by build_selection().

  • merged_folder (str) – Directory containing merged .npy arrays.

  • dst (str) – Destination for split-specific crop directories.

  • channel_arrays (sequence of int, default=(0, 1, 2)) – Image channels to retain.

  • mask_index (int, optional) – Mask plane index. The final plane is used when omitted.

  • bounding_box (bool, default=True) – Whether crops retain every pixel in the object bounding box.

  • crop_mode (str, default="cell") – Object type used for crop naming.

  • write (callable, optional) – Function accepting (path, array). The default writes NumPy arrays.

Returns:

dict – Written and missing crop counts, processed field count, destination folders, and per-field problems.

spacr.stream_dataset.stream_dataset(settings: Mapping[str, Any], dst: str, *, objects: pandas.DataFrame | None = None) → Dict[str, Any][source]

Build, save, and execute a streamed-dataset selection.

Parameters:
  • settings (mapping) – Source, object type, split, channel, and crop-shape settings.

  • dst (str) – Dataset destination.

  • objects (pandas.DataFrame, optional) – Object table used to build the selection. Merged-array labels are used when the table is unavailable.

Returns:

dict – Streaming report augmented with the saved selection-table path.

Nested helpers

stream.write(path, array)

Save one extracted crop using the default NumPy format.

Parameters:
  • path – proposed output path; its extension is replaced with .npy.

  • array – crop data to coerce to an array before saving.

Returns:

None.

spacr/stream_dataset.py:454