spacr.stream_dataset¶
Create reproducible training datasets from merged image arrays.
The workflow first records a deterministic object selection table, then extracts the selected crops from merged arrays. Object types, channels, crop shape, and train/test allocation are therefore explicit settings rather than properties of a previously exported directory. Crop names use the same naming function as the standard object-crop exporter.
Functions¶
|
Build and save the object-selection table. |
|
Return the identifier column for an object-array type. |
|
Return the standard exported-crop name for an object. |
|
Extract one labeled object from an image stack. |
|
Build a streaming selection from labels in merged arrays. |
|
Build a streaming selection from an object table. |
|
Return settings used by a dataset-selection method. |
|
Write crops described by a saved selection table. |
|
Build, save, and execute a streamed-dataset selection. |
Module Contents¶
- spacr.stream_dataset.build_selection(dst: str, *, objects: pandas.DataFrame | None = None, merged_folder: str = '', object_array: str = 'cell', test_split: float = 0.2, seed: int = 0) Tuple[pandas.DataFrame, str][source]¶
Build and save the object-selection table.
An available object table is preferred; otherwise labels are read from merged arrays.
- Parameters:
dst (str) – Output folder; created if missing, and
stream_selection.csvis written into it.- Returns:
pandas.DataFrame – Selected objects and deterministic split assignments.
str – Path to the written
stream_selection.csvfile.
- spacr.stream_dataset.coordinate_column(object_array: str) str[source]¶
Return the identifier column for an object-array type.
- spacr.stream_dataset.crop_name(field_stem: str, object_id, *, crop_mode: str = 'cell', nucleus_ids=(), pathogen_ids=(), timelapse: bool = False) str[source]¶
Return the standard exported-crop name for an object.
Naming is delegated to
spacr.utils._generate_names(), ensuring that streamed and pre-exported crops from the same object are compatible.
- spacr.stream_dataset.cut(stack: numpy.ndarray, mask: numpy.ndarray, label: int, *, bounding_box: bool = True, channels: Sequence[int] = ()) numpy.ndarray | None[source]¶
Extract one labeled object from an image stack.
- Parameters:
stack (numpy.ndarray) – Two- or three-dimensional image data.
mask (numpy.ndarray) – Integer label mask aligned with the image dimensions.
label (int) – Object label to extract.
bounding_box (bool, default=True) – Return the complete bounding box when true. When false, pixels outside the selected object are set to zero within the box.
channels (sequence of int, optional) – Channel indices to retain. All channels are used when omitted.
- Returns:
numpy.ndarray or None – Extracted crop, or
Nonewhen the label is absent.
- spacr.stream_dataset.selection_from_arrays(merged_folder: str, *, object_array: str = 'cell', test_split: float = 0.2, seed: int = 0, mask_index: int | None = None) pandas.DataFrame[source]¶
Build a streaming selection from labels in merged arrays.
- Parameters:
merged_folder (str) – Directory containing merged
.npyarrays.object_array (str, default="cell") – Object type represented by the selected mask plane.
test_split (float, default=0.2) – Fraction assigned to the test set.
seed (int, default=0) – Random seed controlling deterministic split assignment.
mask_index (int, optional) – Mask plane index. The final plane is used when omitted.
- Returns:
pandas.DataFrame – Canonical identifiers, object type, split, and source array for every nonzero label.
- Raises:
FileNotFoundError – If no merged arrays or nonzero object labels are available.
- spacr.stream_dataset.selection_from_objects(frame: pandas.DataFrame, *, object_array: str = 'cell', test_split: float = 0.2, seed: int = 0, ask: Any | None = None) pandas.DataFrame[source]¶
Build a streaming selection from an object table.
- Parameters:
frame (pandas.DataFrame) – One row per measured object.
object_array (str, default="cell") – Object type to select.
test_split (float, default=0.2) – Fraction assigned to the test set.
seed (int, default=0) – Random seed controlling deterministic split assignment.
- Returns:
pandas.DataFrame – Canonical identifiers, object type, split, and source for each object.
- Raises:
ValueError – If the table is empty or contains no usable object identifier, and either nobody was asked or the answer did not help.
Notes
askis called only when the table names no object, and is INJECTED rather than imported so this module never depends on Qt. A caller with nobody in front of it passes none and gets exactly the error it always got, which makes “never prompt in a batch run” structural rather than a rule each call site has to remember. It is calledask(tried=..., object_array=...)and returns((database, table, column), why)– seespacr.qt.ask_for_the_path.ask_for_a_database_column().
- spacr.stream_dataset.settings_for_method(method: str) Tuple[str, ...][source]¶
Return settings used by a dataset-selection method.
- spacr.stream_dataset.stream(selection: pandas.DataFrame, merged_folder: str, dst: str, *, channel_arrays: Sequence[int] = (0, 1, 2), mask_index: int | None = None, bounding_box: bool = True, crop_mode: str = 'cell', write=None) Dict[str, Any][source]¶
Write crops described by a saved selection table.
Each merged field is loaded once and all selected objects from that field are extracted before advancing. Missing or ambiguous field matches are reported as missing crops; a different field is never used as a fallback.
- Parameters:
selection (pandas.DataFrame) – Selection produced by
build_selection().merged_folder (str) – Directory containing merged
.npyarrays.dst (str) – Destination for split-specific crop directories.
channel_arrays (sequence of int, default=(0, 1, 2)) – Image channels to retain.
mask_index (int, optional) – Mask plane index. The final plane is used when omitted.
bounding_box (bool, default=True) – Whether crops retain every pixel in the object bounding box.
crop_mode (str, default="cell") – Object type used for crop naming.
write (callable, optional) – Function accepting
(path, array). The default writes NumPy arrays.
- Returns:
dict – Written and missing crop counts, processed field count, destination folders, and per-field problems.
- spacr.stream_dataset.stream_dataset(settings: Mapping[str, Any], dst: str, *, objects: pandas.DataFrame | None = None) Dict[str, Any][source]¶
Build, save, and execute a streamed-dataset selection.
- Parameters:
settings (mapping) – Source, object type, split, channel, and crop-shape settings.
dst (str) – Dataset destination.
objects (pandas.DataFrame, optional) – Object table used to build the selection. Merged-array labels are used when the table is unavailable.
- Returns:
dict – Streaming report augmented with the saved selection-table path.
Nested helpers¶
- stream.write(path, array)¶
Save one extracted crop using the default NumPy format.
- Parameters:
path – proposed output path; its extension is replaced with
.npy.array – crop data to coerce to an array before saving.
- Returns:
None.
spacr/stream_dataset.py:454