spacr.core

Workflow inputs and outputs

Mask

Select channels and models, inspect a preview, then run Mask. Measure consumes the merged arrays; the counts database is not yet a feature table.

Folder watching uses watch_normalization_pool='per_field' by default. For a declared static projected acquisition, 'fixed_map' waits for every image in conversion_map.csv before running the ordinary Mask or Mask/Measure batch with shared percentile normalization and padding. Set randomize=False and batch_size greater than one. A timeout never releases an incomplete pool; restart checkpoints bind the declared members, recipes and original input hashes.

Native time-volume Mask accepts a complete fixed Convert map of explicitly identified C/Z/T planes with both z_stack and t_stack enabled. Declare TZYX axis order, physical Z spacing and frame interval. Batch ingest and folder watching retain the original planes and build TZYXC archives for 4-D segmentation. This route supports Mask only; it does not enable Measure, tracking or Classify for native time volumes.

Open: Home → Mask.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.

  • Segmentation checkpoint — Saved Cellpose-compatible checkpoint or a compatible installed backend selected with its own configuration.

Outputs

  • Images and label masks — merged/*.npy in the project; channels and integer label planes share each field array.

  • Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.

  • Object counts — measurements/measurements.db; Mask counts alone are not per-object feature measurements.

Before this module

  • Cellpose Workbench: Select the saved compatible checkpoint in Mask.

  • Import Images: Use imported image planes and identities; image-only imports still need segmentation.

  • Format Converter: Use the converted layout and preserve source identity mappings.

  • Import: For image-only imports, use Import Images or Format Converter and point Mask at the formatted image project. External measurements alone are not segmentation input.

After this module

  • Measure: Use the same project and the correct image/mask channel indices.

  • Timelapse: Enable the nested time-series route before generating linked labels.

API reference.

Module tutorial.

Image UMAP

Project measured features or supplied encoder features and inspect representative crops. A cluster is a candidate grouping, not a validated phenotype. Use the lasso and annotation controls to write reviewed selections to an annotation column in the matching measurement database. A geometric selection alone does not establish a biological phenotype.

Open: Home → Image UMAP.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route: cell, nucleus, pathogen, cytoplasm. Relevant columns, depending on the route: plateID, rowID, columnID, fieldID.

  • Object crops — data/**/*_png when save_png is enabled; png_list indexes saved crops. Supported workflows can instead stream crops from merged arrays and masks. Relevant tables, depending on the route: png_list. Relevant columns, depending on the route: png_path, prcfo.

  • Image embeddings — Object-indexed encoder features, with channel policy and encoder provenance. The encoding API returns features; persistence is caller-dependent.

Outputs

  • Projection and clusters — Image UMAP/PCA coordinate tables, selected clusters and figures for the loaded measurement data.

  • Training annotations — A chosen annotation column in measurements/measurements.db, table png_list; labels belong to object identities. Relevant tables, depending on the route: png_list. Relevant columns, depending on the route: prcfo.

Before this module

  • Measure: Choose feature columns and inspect representative crops.

  • Embeddings: Supply the encoder feature table with matching object IDs; do not assume every GUI route automatically persists it.

After this module

  • Gate Editor: Supply the matching coordinate/feature columns when defining a selection.

  • Classify: Write reviewed lasso selections to an annotation column in the matching object database, then select that column in Classify. Inspect crops and validate labels; embedding clusters are not ground truth.

API reference.

Module tutorial.

Timelapse

Open Timelapse within Mask to segment and link objects across ordered frames; inspect links before downstream motility analysis.

Lineage exports retain generation times in frames and add generation_time_hours with calibrated summary statistics. Supply a positive frame_interval_s or complete, increasing time_s values in the tracks table. If both are present, they must agree; irregular timestamps are supported without a fixed interval. Missing calibration and incomplete cell cycles have no generation time in hours. Exported tables record the calibration source and units. The tree axis and Newick branch lengths remain in frames.

Open: Mask → Timelapse.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.

Outputs

  • Images and label masks — merged/*.npy in the project; channels and integer label planes share each field array.

  • Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.

  • Linked time-series objects — Tracked labels and frame/object associations from the time-series project, with frame interval and units.

Before this module

  • Mask: Enable the nested time-series route before generating linked labels.

After this module

  • Measure: Use the time-series project with stable frame/object identities.

  • Motility Assay: Supply frame interval and pixel calibration.

API reference.

Module tutorial.

Create microscopy masks and explore measured phenotypes in two dimensions.

WHAT IT IS FOR

This landing page currently serves two spaCR tiles. Mask runs preprocess_generate_masks() to turn raw multichannel acquisitions into preprocessed arrays and Cellpose masks for cells, nuclei, pathogens, and organelles. Image UMAP runs generate_image_umap() to reduce measured single-object features, cluster them, and optionally place representative image crops on the embedding. The same module also exposes the timelapse-mask entry point; the two main workflows remain independent, and UMAP does not segment images.

WHAT IT NEEDS

Mask generation needs one or more source folders, a filename metadata scheme (cellvoyager or an automatic/custom regex), zero-based channel indices for the objects to segment, and suitable object diameters and model choices. At least one segmentation channel must be enabled. Start with dry_run=True to validate paths, channels, models, and the planned writes without loading a model or changing the project.

Image UMAP needs an existing measurements/measurements.db for every source, the object tables and features to include, and reduction/clustering settings. It embeds numeric measurements rather than raw pixels. Thumbnail images come from the measured png_list table when crop_source='png' or are cut from merged/*.npy on demand when crop_source='merged'; the latter is useful when measurement crops were not saved.

WHAT IT PRODUCES

A normal Mask run writes preprocessed stacks, object masks, overlays and segmentation-QC artifacts, settings CSVs, counts in measurements.db, and a run manifest beneath the source tree; it normally returns None. A dry run instead returns its preflight problem list. Timelapse mask generation also writes movies and masks relabelled with track identities.

Image UMAP returns an annotated DataFrame containing the two-dimensional coordinates and cluster labels, or a Matplotlib figure when return_fig=True. Depending on the save and plotting settings, it also writes the embedding, cluster views, representative-crop grids, feature summaries, and the resolved settings alongside the project.

WHAT TO DO NEXT

After Mask finishes, inspect overlays and segmentation-QC flags before running spacr.measure.measure_crop(); inaccurate masks make every downstream feature inaccurate. After measurement, use Image UMAP to inspect phenotype structure, colour by plate or condition to expose batch effects, and validate clusters against their representative crops before treating them as biology. Use the Mask tile for preprocess_generate_masks() and the Image UMAP tile for generate_image_umap() until those tiles receive separate API pages.

Three details are deliberately explicit. Channel numbers are zero-based and diameters are pixels, so values copied from one magnification are not portable without conversion. The default v1 mask pipeline preserves the directory layout expected by downstream tools; pipeline_style='v2' is opt-in and writes a different streaming layout. Finally, removing UMAP cluster noise also removes those objects from the returned frame, keeping the table and the visible embedding aligned rather than silently returning different samples.

Functions

display(*args, **kwargs)

Discard display payloads when IPython's helper is unavailable.

generate_image_umap([settings, return_fig])

Reduce per-object features and plot the resulting 2-D embedding.

generate_screen_graphs(settings)

Build recruitment-metric summary graphs per source and for the combined data.

preprocess_generate_masks(settings)

Turn a folder of raw microscopy images into per-channel Cellpose masks ready for spacr.measure.measure_crop().

preprocess_generate_masks_timelapse(settings)

Entry point for the standalone Timelapse module.

reducer_hyperparameter_search([settings, ...])

Sweep UMAP/tSNE and DBSCAN/KMeans hyperparameters over the feature table.

Module Contents

spacr.core.display(*args, **kwargs)[source]

Discard display payloads when IPython’s helper is unavailable.

spacr.core.generate_image_umap(settings=None, return_fig=False)[source]

Reduce per-object features and plot the resulting 2-D embedding.

Reads measurements from the SQLite backend(s), applies preprocessing and dimensionality reduction, clusters the embedding, and renders scatter/grid plots of the resulting clusters.

The thumbnails overlaid on the embedding come from whichever crop source crop_source names: 'png' (and 'auto' wherever a crop folder exists) reads the pre-generated PNGs, 'merged' (and 'auto' with no folder) cuts each one out of merged/*.npy on demand through spacr.crops, so the embedding can be drawn on a project that never kept a crop folder — png_list is not required in that case. Either way the pixels arrive through spacr.crops, so a legacy (pre-341f446) crop folder is channel-corrected on load and the montage shows the same image the Annotate screen does.

Parameters:
  • settings – Configuration dict; canonicalized via spacr.settings.set_default_umap_image_settings(). Common keys: src, tables, row_limit, clustering, reduction_method (UMAP, t-SNE, PCA, Isomap or Spectral), embedding_by_controls, col_to_compare, pos, neg, plot_images, save_figure, exclude, crop_source.

  • return_fig – When True, return the Matplotlib figure instead of the annotated DataFrame.

Returns:

DataFrame of the input rows plus a cluster column, or a Matplotlib Figure when return_fig is True. With remove_cluster_noise the noise objects are dropped from the frame as well as from the embedding, so the two always describe the same objects.

Raises:
  • ValueError – when a source has no usable measurements.db — see _validate_umap_source_db() for exactly what is required.

  • NotImplementedError – when resnet_features is set; embedding raw crops with ResNet features is not implemented.

spacr.core.generate_screen_graphs(settings)[source]

Build recruitment-metric summary graphs per source and for the combined data.

Reads per-object measurements, annotates conditions, computes the recruitment metric, and generates one plot per source folder plus one combined plot.

Parameters:

settings – Config dict with keys src (path or list of paths), tables, cells, controls, controls_loc, graph_type, summary_func, y_axis_start, error_bar_type, theme, representation, nuclei_limit, pathogen_limit.

Returns:

None. Figures and CSVs are written under each source’s results/ folder.

spacr.core.preprocess_generate_masks(settings)[source]

Turn a folder of raw microscopy images into per-channel Cellpose masks ready for spacr.measure.measure_crop().

Given a source folder src of multi-channel images, this pipeline (1) optionally consolidates inputs from nested folders, (2) renames files to the Yokogawa layout used downstream, (3) preprocesses per-channel arrays, (4) generates masks for cell / nucleus / pathogen / organelle channels via Cellpose (SAM variant), (5) optionally reconciles the cell mask against nuclei + pathogen overlays, and (6) writes overlay plots and a gen_mask_settings.csv next to the outputs.

Parameters:

settings –

Settings dict; canonicalized via spacr.settings.set_default_settings_preprocess_generate_masks(). Must include src and at least one of cell_channel, nucleus_channel, pathogen_channel, or organelle_channel. Key entries the function reads:

  • src (str or list of str) — image folder(s) to process.

  • metadata_type — 'cellvoyager' or 'auto' (uses custom_regex when set).

  • cell_channel / nucleus_channel / pathogen_channel / organelle_channel — 0-based channel indices; None skips.

  • cell_diameter / nucleus_diameter / pathogen_diameter — Cellpose object diameters in pixels.

  • pathogen_model — path to a Cellpose-SAM checkpoint to segment pathogens with, instead of stock cpsam. The pre-SAM toxo_pv_lumen / toxo_cyto names are gone and resolve to cpsam; a PATH to a fine-tune is honoured, and a path that is not there stops the run rather than silently using stock weights.

  • consolidate — copy nested images into src/consolidated before processing.

  • preprocess / masks — toggle the two pipeline halves.

  • adjust_cells — reconcile cell masks against nuclei+pathogen.

  • timelapse — enable trackpy linking; forces randomize=False.

  • motility_analysis — when timelapse is enabled, analyze the completed merged frames once per plate, rebuilding measurements from the current masks rather than reusing an older assay table.

  • robustness_report — after the masks exist, re-segment a few sampled fields over a small grid of diameters, thresholds and contrast enhancement and write a stability report per object to qc/segmentation_robustness_<object>.csv, flagging the settings the results are fragile to.

  • dry_run — validate only: inspect the input folders, print the preflight report and plan and return, without writing anything or loading a model.

  • watch_folder — keep watching src and analyse each field as its files arrive and stop changing, with watch_pipeline, watch_measure_settings, watch_settle_seconds, watch_poll_seconds and watch_idle_minutes. Results gather in src/spacr_watch, and a record there lets a restarted watch skip the fields already analysed. watch_normalization_pool='fixed_map' instead waits for the entire declared static acquisition and retains the ordinary batch recipe’s shared normalization and padding.

  • microscope_feedback — during a watch_folder run with watch_pipeline='mask_measure', send the objects matching microscope_event_query back to the microscope named by microscope_driver to be imaged again, at stage positions worked out from microscope_positions and microscope_stage_transform.

  • src may also be a cloud address (s3://, gs://, az://, https://). An OME-Zarr plate there has only the wells, fields and pyramid level named by cloud_wells, cloud_fields and cloud_level fetched, as TIFFs, into a folder under cloud_cache, and the run analyses that folder; a cloud folder of images is mirrored there instead. Credentials come from the standard places, chosen with cloud_anonymous, cloud_profile and cloud_endpoint, and cloud_results copies the measurements folder back to cloud storage.

  • save, plot, verbose, test_mode, n_jobs.

Returns:

None on a normal run, having written masks, overlays, measurements.db counts and settings CSVs into subfolders of src. When dry_run is set, returns instead the list of problems from spacr.validate.run_preflight(). When watch_folder is set, returns a dict naming the fields analysed, failed and never completed when the watch ends.

Raises:

ValueError – if src is missing or of the wrong type, or no segmentation channel is defined.

Example

from spacr.core import preprocess_generate_masks
settings = {
    'src': '/data/plate01',
    'cell_channel': 0, 'nucleus_channel': 1, 'pathogen_channel': 2,
    'cell_diameter': 60, 'nucleus_diameter': 20, 'pathogen_diameter': 8,
    'magnification': 20, 'save': True, 'plot': True,
}
preprocess_generate_masks(settings)

See also

spacr.io.preprocess_img_data() — the preprocessing half only. spacr.measure.measure_crop() — downstream feature extraction.

Mask writes which object sits in which as its own table in measurements.db. That table is built from Measure’s object tables, so before Measure has run there are none, and that is not a failure.

spacr.core.preprocess_generate_masks_timelapse(settings)[source]

Entry point for the standalone Timelapse module.

Identical to preprocess_generate_masks() except that timelapse is forced on, so every well/field is grouped into a time stack, randomization is switched off, per-channel movies are written, and the objects listed in timelapse_objects are linked across frames and relabelled with their track IDs.

Timelapse is a first-class spaCR workflow, not a checkbox on mask generation — that is why it has its own module, its own settings group (spacr.settings.get_timelapse_settings()) and this entry point.

Parameters:

settings – Settings dict; canonicalized via spacr.settings.get_timelapse_settings(). Same keys as preprocess_generate_masks() plus the timelapse_* tracking group. timelapse is overwritten with True.

Returns:

None. Same outputs as preprocess_generate_masks(), plus <src>/movies and track-relabelled masks.

See also

spacr.timelapse.automated_motility_assay() — the Motility Assay module, which consumes the tracked merged/*.npy this produces.

Sweep UMAP/tSNE and DBSCAN/KMeans hyperparameters over the feature table.

Renders a grid of embeddings, one cell per (reduction, clustering) pair, so the caller can eyeball the impact of each parameter combination.

Parameters:
  • settings – Config dict; canonicalized via spacr.settings.set_default_umap_image_settings().

  • reduction_params – Dict or list of dicts of parameters for the reduction method. Presence of n_neighbors selects UMAP, perplexity selects tSNE.

  • dbscan_params – Dict or list of DBSCAN parameter dicts (each with eps and min_samples).

  • kmeans_params – Dict or list of KMeans parameter dicts.

  • save – When True, save the grid figure to <src>/results.

  • show – When True and not saving, call plt.show.

  • return_fig – When True, return the Matplotlib figure.

Returns:

The figure when return_fig is True, otherwise None.

Raises:

ValueError – when reduction_params is missing or empty, when it contains neither n_neighbors nor perplexity, when it mixes the two, or when settings['reduction_method'] is neither UMAP nor tSNE. All four are checked before any data is read.