spacr.utils

Shared image, model, database, statistics, and pipeline utilities.

Exceptions

ImportedCopyNotReleased

The table being appended to holds an import's copy of the same field.

MeasurementUnitsMismatch

A measurement frame's units differ from the ones already in the table.

OptionalDependencyCompatibilityError

An installed optional dependency is too old for spaCR's API contract.

Classes

Cache

LRU cache with a fixed maximum size.

CustomCellClassifier

Small classifier stacking EarlyFusion and a multi-scale attention block.

EarlyFusion

1x1 convolution that fuses input channels down to 64 feature maps.

FocalLossWithLogits

Focal loss for binary, multiclass, and multilabel targets.

GradCAM

Named-hook Grad-CAM implementation for arbitrary target layers.

GradCAMGenerator

Grad-CAM (and variants) map generator for binary classifiers.

IntegratedGradients

Compute integrated-gradients attributions for a classifier.

MultiScaleBlockWithAttention

Dilated conv block followed by a 1x1 attention convolution.

ResNet

ResNet backbone with a two-layer spaCR binary-classification head.

SaliencyMapGenerator

Generate saliency maps and predictions for a binary classifier.

ScaledDotProductAttention

Standard scaled dot-product attention layer.

SelectChannels

Callable transform that zeroes out image channels not present in channels.

SelfAttention

Linear-projected self-attention layer.

SpatialAttention

Spatial attention gate that reweights features by pooled channel statistics.

TorchModel

Thin wrapper around TorchVision classification backbones that:

TorchModel_v2

TorchVision backbone with a spaCR linear head (streamlined variant of TorchModel).

Functions

MLR(merged_df, refine_model)

Fit a multiple-linear regression on gene:grna interactions plus plate/row/column terms.

activation_correlations_to_database(df, img_paths, ...)

Merge per-image correlation stats with parsed well IDs and insert into the dataset DB.

activation_maps_to_database(img_paths, source_folder, ...)

Insert activation-map PNG paths and parsed well IDs into the dataset DB.

add_column_to_database(settings)

Adds a new column to the database table by matching on a common column from the DataFrame.

add_images_to_tar(paths_chunk, tar_path, total_images)

Add paths_chunk images to tar_path, updating the shared counter for progress.

adjust_cell_masks(parasite_folder, cell_folder, ...[, ...])

Run process_mask_file_adjust_cell() in parallel across matching mask files.

all_elements_match(list1, list2)

Return True if every element of list1 is contained in list2.

annotate_conditions(df[, cells, cell_loc, pathogens, ...])

Annotate df with host cell, pathogen, treatment, and combined condition columns.

annotate_predictions(csv_loc)

Read prediction CSV and add plate/well/field/object columns plus a cond label.

apply_mask(image[, output_value])

Zero out (or set to output_value) pixels outside a circular mask fit to image.

assign_colors(unique_labels, random_colors)

Return colors and their positional mapping for the unique labels.

augment_classes(dst, nc, pc[, generate, move, ...])

Augment negative and positive class images and split them into train/test folders.

augment_dataset(dataset[, is_grayscale])

Expand dataset by 8x through rotation and horizontal reflection of every image tensor.

augment_image(image)

Return a list of PIL images covering 4 rotations x 2 horizontal reflections of image.

augment_images(file_paths, dst)

Run augment_single_image() in parallel over file_paths.

augment_single_image(args)

Save six augmentations of one image (original, 90/180/270 rotations, H/V flips).

boundary_f1_score(mask_true, mask_pred[, dilation_radius])

Return the boundary F1 score between two masks with tolerance dilation_radius.

build_loss([loss_type, num_classes, class_counts, ...])

Return a closure loss_fn(logits, target) implementing the requested loss.

calculate_activation_correlations(inputs, ...[, ...])

Compute per-image Pearson and Manders correlations between input and activation channels.

calculate_iou(mask1, mask2)

Return the intersection-over-union of two binary masks after zero-padding to a common shape.

calculate_loss(output, target[, prefer_focal, gamma, ...])

Auto-select and return a loss for binary, multiclass, or multilabel problems.

calculate_shortest_distance(df, object1, object2)

Calculate the shortest edge-to-edge distance between two objects (e.g., pathogen and nucleus).

canonicalize_measurement_columns(df)

Rename legacy column spellings on an in-memory measurement frame.

check_index(df[, elements, split_char])

Validate that every index label in df splits into elements parts on split_char.

check_mask_folder(src, mask_fldr[, resume])

Return True if masks in src/masks/mask_fldr still need generating.

check_multicollinearity(x)

Checks multicollinearity of the predictors by computing the VIF.

check_normality(series)

Helper function to check if a feature is normally distributed.

check_overlap(current_position, other_positions, threshold)

Return True if current_position is within threshold of any point in other_positions.

choose_model(→ Optional[torch.nn.Module])

Instantiate a classification model by name for binary or multiclass problems.

class_visualization(target_y, model_path, dtype[, ...])

Synthesize an input image that maximizes the classifier score for target_y.

classification_metrics(all_labels, ...)

Return a one-row DataFrame of accuracy, PR-AUC, and optimal-threshold stats.

cleanup_pipeline_folders(src[, keep_intermediate, ...])

Delete the intermediate mask-pipeline folders once merged/ is built.

close_file_descriptors()

Close file descriptors from 3 up to the soft NOFILE limit.

close_multiprocessing_processes()

Terminate all detected multiprocessing child processes and close file descriptors.

cluster_feature_analysis(all_df[, cluster_col])

Perform Random Forest feature importance, ANOVA for normally distributed features,

combine_results(rf_df, anova_df, kruskal_df)

Combine the results into a single DataFrame.

compute_ap_over_iou_thresholds(true_masks, pred_masks, ...)

Return the area under the precision-recall curve swept over iou_thresholds.

compute_average_precision(matches, num_true_masks, ...)

Return (precision, recall) given match count, true count, and predicted count.

compute_irm_penalty(losses, dummy_w, device)

Return the IRM penalty as the sum of squared gradient dot-products across environments.

compute_segmentation_ap(true_masks, pred_masks[, ...])

Return the COCO-style segmentation AP by matching connected components across IoU thresholds.

console_can_encode(text[, stream])

Return True when text can be printed to stream as-is.

console_encoding([stream])

Return the codec text printed to stream has to survive.

console_safe(text[, stream])

Return text with anything the console cannot encode replaced by ?.

control_filelist(folder[, mode, values])

Return filenames in folder whose row or column ID matches one of values.

convert_and_relabel_masks(folder_path)

Converts all int64 npy masks in a folder to uint16 with relabeling to ensure all labels are retained.

copy_images_to_consolidated(image_path_map, root_folder)

Copies images from their original locations to a 'consolidated' folder,

correct_masks(src)

Convert cell masks under src/masks/cell_mask_stack to uint16 and re-stack arrays.

correct_metadata(df)

Normalize a metadata DataFrame to the canonical spaCR names and plate ids.

correct_metadata_column_names(df)

Renamed legacy metadata columns. Defined in spacr.schema.

correct_paths(df, base_path[, folder])

Rewrite PNG paths (in a DataFrame or list) so they live under base_path/folder.

count_reads_in_fastq(fastq_file)

Return the number of reads in a gzipped FASTQ file.

create_circular_mask(h, w[, center, radius])

Return a boolean circular mask of shape (h, w) centered on center.

debug([enabled, logger_name])

Decorator that temporarily sets the given logger to DEBUG for the wrapped call.

delete_folder(folder_path)

Recursively delete folder_path if it exists (files and subdirectories included).

delete_intermedeate_files(settings)

Remove intermediate per-channel and stack folders under settings['src'].

dense_mask_channel_positions(settings)

Map each RAW channel index to its position on the merged stack's axis.

dice_coefficient(mask1, mask2)

Return the Dice similarity of two masks, treating any nonzero value as foreground.

display(*args, **kwargs)

Do nothing: IPython is unavailable, so there is nowhere to display to.

download_models([repo_id, retries, delay])

Downloads all model files from Hugging Face and stores them in the resources/models directory

estimate_class_counts(→ torch.Tensor)

Return per-class sample counts as a LongTensor of length num_classes.

extract_boundaries(mask[, dilation_radius])

Return the boundary of a binary mask via morphological dilation minus erosion.

extract_features(image_paths[, resnet])

Extract features from images using a pre-trained ResNet model.

extract_tar_bz2_files(folder_path)

Extracts all .tar.bz2 files in the given folder into subfolders with the same name as the tar file.

feature_columns(columns, selection)

Which of columns the selection keeps. Order preserved.

feature_folder_name(→ str)

A folder name for one feature selection. Safe on every filesystem.

filepaths_to_database(img_paths, settings, ...)

Insert cropped PNG filepaths and parsed well/object IDs into the measurements DB.

fill_holes_in_mask(mask)

Fill the holes inside each object of a label mask, keeping every id.

filter_and_save_csv(input_csv, output_csv, ...)

Reads a CSV into a DataFrame, keeps the rows whose column value falls OUTSIDE

filter_columns(df, filter_by)

Return df restricted to columns matching filter_by (or morphology columns).

filter_dataframe_features(df, channel_of_interest[, ...])

Restrict a features DataFrame to a channel of interest and clean up correlated/low-variance columns.

find_non_overlapping_position(x, y, image_positions, ...)

Return a nearby (x, y) jittered position that does not collide with image_positions.

fishers_odds(df[, threshold, phenotyp_col])

Fisher's exact test per mutant column against a binarized phenotype label.

format_path_for_system(path)

Takes a file path and reformats it to be compatible with the current operating system.

generate_colors(num_clusters, black_background)

Return a deterministic Viridis RGBA palette for cluster points.

generate_cytoplasm_mask(nucleus_mask, cell_mask)

Generates a cytoplasm mask from nucleus and cell masks.

generate_fraction_map(df, gene_column[, min_frequency])

Return a wells-by-genes fraction matrix, dropping columns below min_frequency.

generate_image_path_map(root_folder[, valid_extensions])

Recursively scans a folder and its subfolders for images, then creates a mapping of:

generate_path_list_from_db(db_path, file_metadata)

Return all png_path values from db_path optionally filtered by file_metadata substrings.

get_cuda_version()

Return the installed CUDA toolkit version as a digit-only string, or None.

get_db_paths(src)

Return the standard measurements/measurements.db paths for one or more source roots.

get_files_from_dir(dir_path[, file_extension])

Return glob matches for dir_path/file_extension.

get_ml_results_paths(src[, model_type, ...])

Return the standard set of ML output paths for the given model and channel selection.

get_paths_from_db(df, png_df[, image_type])

Return rows of png_df whose path contains image_type and whose prcfo is in df.

get_sequencing_paths(src)

Return the standard sequencing/sequencing_data.csv paths for one or more source roots.

get_submodules(model[, prefix])

Return all dotted submodule names of model in traversal order.

group_feature_class(df[, feature_groups, name])

Add a column tagging each feature with its compartment (or other group) label.

initiate_counter(counter_, lock_)

Initialize shared multiprocessing counter and lock globals.

invert_image(image)

Return the intensity-inverted image, reflected through the dtype range.

is_list_of_lists(var)

Return True if var is a list whose every element is also a list.

is_multiprocessing_process(process)

Return True if process cmdline contains multiprocessing.

jaccard_index(mask1, mask2)

Return the Jaccard/IoU index of two binary masks.

lasso_reg(merged_df[, alpha_value, reg_type])

Fit Lasso or Ridge on one-hot-encoded gene/grna/plate/row/column predictors.

load_image(image_path)

Load and preprocess an image.

load_image_paths(c, visualize)

Load the png_list table into a DataFrame indexed by prcfo and optionally filter by object.

load_settings(csv_file_path[, show, setting_key, ...])

Reload a spacr settings CSV (written by save_settings()) back into a Python dict.

map_condition(col_value[, neg, pos, mix])

Map a column-ID value to one of 'neg', 'pos', 'mix', or 'screen'.

mask_object_count(mask)

Return the number of nonzero labeled objects in mask.

match_masks(true_masks, pred_masks, iou_threshold)

Greedy match each predicted mask to a still-unmatched true mask above iou_threshold.

measure_test_mode(settings)

Copy a random subset of source files into a test/merged folder when test_mode is on.

merge_dataframes(df, image_paths_df, verbose)

Merge df into image_paths_df on the shared prcfo index.

merge_regression_res_with_metadata(results_file, ...)

Merge regression outputs with gene metadata on the parsed gene column.

merge_split_objects(mask_src[, intensity_img_src, ...])

Merge by perimeter and filter labeled objects across a directory of masks.

merge_touching_objects(mask[, threshold])

Merge touching labeled objects whose shared boundary exceeds threshold of the smaller perimeter.

model_metrics(model)

Print RMSE/MAE/Durbin-Watson and show residual/QQ/scale-location diagnostic plots.

normalize_feature_filter(filter_by)

Normalize text representations of an unfiltered feature selection.

normalize_src_path(src)

Ensures that the 'src' value is properly formatted as either a list of strings or a single string.

normalize_to_dtype(array[, p1, p2, percentile_list, ...])

Percentile-normalize each channel of an image stack into the target dtype range.

object_label_from_png_id(values)

Migrate png_list's 'o<N>' text ids onto the integer object label.

pad_to_same_shape(mask1, mask2)

Zero-pad mask1 and mask2 to their element-wise maximum shape.

perform_statistical_tests(all_df[, cluster_col])

Perform ANOVA or Kruskal-Wallis tests depending on normality of features.

pick_best_model(src)

Return the strongest checkpoint anywhere below src.

plot_clusters(ax, embedding, labels, colors, ...[, ...])

Draw cluster outlines, points, and centroid labels onto ax for a 2-D embedding.

plot_clusters_grid(embedding, labels, image_nr, ...[, ...])

Plot a grid of example images per cluster label discovered in labels.

plot_embedding(embedding, image_paths, labels, ...[, ...])

Plot a 2-D embedding with cluster outlines, points, and optional image overlays.

plot_grid(cluster_images, colors, figuresize, ...[, ...])

Render one column per cluster of representative images with colored borders and labels.

plot_image(ax, x, y, img, img_zoom[, remove_image_canvas])

Place a zoomed thumbnail of img at (x, y) on ax.

plot_images_by_cluster(ax, image_paths, embedding, ...)

Overlay up to image_nr images per cluster on the embedding in ax.

plot_umap_images(ax, image_paths, embedding, labels, ...)

Overlay sample images from image_paths on the UMAP embedding in ax.

prepare_batch_for_segmentation(batch)

Cast a batch to float32 and per-image max-normalize any image whose max exceeds 1.

preprocess_data(df, filter_by, ...[, column_list, ...])

Prepare a feature matrix by filtering, decorrelating, log-transforming, and scaling df.

preprocess_image(image_path[, normalize, image_size, ...])

Load and preprocess image_path into a batched tensor ready for classification.

pretty_print_settings(settings[, title])

Print a settings dict to the console as a tidy, aligned table.

print_progress(files_processed, files_to_process, n_jobs)

Print a one-line progress report with an ETA derived from mean step time.

process_mask_file_adjust_cell(file_name, ...[, ...])

Load one triple of parasite/cell/nuclei masks, merge cells in place, and return the elapsed time.

process_masks(mask_folder, image_folder, channel[, ...])

Cluster object morphology/intensity across a mask folder and keep the largest cluster in place.

process_vision_results(df[, threshold])

Split image paths into well identifiers and binarize the pred column.

random_forest_feature_importance(all_df[, cluster_col])

Rank features by how well they predict the cluster label.

recommend_target_layers(model)

Return ([last_conv_layer], all_conv_layers) from model.

reduction_and_clustering(numeric_data, n_neighbors, ...)

Reduce numeric_data to 2-D and cluster the embedding.

remove_canvas(img)

Return img as RGBA with zero-valued pixels made transparent.

remove_highly_correlated_columns(df[, threshold, verbose])

Drop numeric columns whose absolute correlation with a prior column exceeds threshold.

remove_intensity_objects(image, mask, ...)

Drop labeled objects whose mean intensity is on the wrong side of intensity_threshold.

remove_low_variance_columns(df[, threshold, verbose])

Drop numeric columns whose variance is below threshold.

remove_noise(embedding, labels)

Drop rows of embedding (and labels) whose label is DBSCAN noise (-1).

remove_outliers_by_group(df, group_col, value_col[, ...])

Removes outliers from value_col within each group defined by group_col.

rename_columns_in_db(db_path)

Rename legacy column spellings across every table in a SQLite database.

reset_cellpose_model_reports()

Forget which Cellpose model notices have already been printed.

reset_mp()

Set the multiprocessing start method appropriate for the current OS.

resize_images_and_labels(images, labels, ...[, ...])

Resize aligned image/label lists to target_height x target_width.

resize_labels_back(labels, orig_dims)

Resize a list of label masks back to their original (width, height).

save_file_lists(dst, data_set, ls)

Write ls as a single-column CSV named <data_set>.csv under dst.

save_settings(settings[, name, show])

Persist a settings dict to <src>/settings/<name>.csv so a spacr run can be reproduced later.

search_reduction_and_clustering(numeric_data, ...[, ...])

Variant of reduction_and_clustering() accepting extra reducer kwargs via reduction_param.

setup_plot(figuresize, black_background[, theme_colors])

Create a square Matplotlib figure using scoped theme colors.

show_cam_on_image(img, mask)

Return img overlaid with a jet colormap of mask as an 8-bit RGB image.

smooth_hull_lines(cluster_data)

Return the x, y coordinates of a smoothed convex-hull outline of a 2-D point set.

split_my_dataset(dataset[, split_ratio])

Randomly split dataset into (train, val) subsets.

suggest_training_changes(dst[, train_csv, val_csv, ...])

Inspect saved training/validation progress CSVs and propose concrete training changes.

Module Contents

exception spacr.utils.ImportedCopyNotReleased[source]

Bases: ValueError

The table being appended to holds an import’s copy of the same field.

foreign.run_import copies the imported frame into the canonical cell / nucleus / pathogen table when the destination is empty, so a project built purely by import is readable by every spaCR tool. That copy is a convenience and it stops being one the moment spaCR measures the same field: _merge_and_save_to_database() appends, so its rows land beside theirs, in different columns, with nothing in the row marking the seam – and every count_cell downstream becomes the sum of two populations.

A resume supersedes the copy before measuring (see spacr.resume.supersede_imported_copies()) or refuses and says so, which covers every path through a resume. This covers the path around one: a direct measure_crop with resume off.

Raised only when the copy cannot be handed back provably losslessly – no foreign_<object> to check the rows against, a row in the copy with no twin in it, a timelapse whose frames the importer never keyed, or a delete that did not act on the rows the count cleared. In every one of those cases the field’s rows are not written, because a refused write can be re-run and a mixed table cannot be un-mixed.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.utils.MeasurementUnitsMismatch[source]

Bases: ValueError

A measurement frame’s units differ from the ones already in the table.

A 2-D field measures areas in px^2; a 3-D field measures volumes, in voxels or um^3, and writes them into the same <object>_area column, because that column is read by name by every downstream selector, model and threshold ever written against a spaCR database and renaming it would break all of them silently. Appending both into one table would therefore leave a numeric column that mixes two incompatible quantities with nothing in the row to tell them apart, which no amount of downstream care could recover from. So it is refused here instead.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.utils.OptionalDependencyCompatibilityError[source]

Bases: ImportError

An installed optional dependency is too old for spaCR’s API contract.

Initialize self. See help(type(self)) for accurate signature.

class spacr.utils.Cache(max_size)[source]

LRU cache with a fixed maximum size.

Parameters:

max_size – maximum number of entries retained; oldest is evicted on overflow.

Store the size limit and initialize an empty OrderedDict.

get(key)[source]

Return and refresh key, or None when it is not cached.

Parameters:

key – cache key to look up.

put(key, value)[source]

Insert value under key, evicting the oldest entry if full.

Parameters:
  • key – cache key under which to store the value.

  • value – object to cache.

class spacr.utils.CustomCellClassifier(num_classes, pathogen_channel, use_attention, use_checkpoint, dropout_rate)[source]

Bases: torch.nn.Module

Small classifier stacking EarlyFusion and a multi-scale attention block.

Parameters:
  • num_classes – output class count.

  • pathogen_channel – reserved for downstream use; kept for API compatibility.

  • use_attention – reserved for downstream use; kept for API compatibility.

  • use_checkpoint – run the forward pass through torch.utils.checkpoint.

  • dropout_rate – reserved for downstream use; kept for API compatibility.

Build the fusion, multi-scale, and linear classifier submodules.

custom_forward(x)[source]

Return the class logits for a batch x of shape (B, 3, H, W).

Parameters:

x – three-channel image batch to classify.

forward(x)[source]

Forward pass, optionally through activation checkpointing.

Parameters:

x – three-channel image batch to classify.

class spacr.utils.EarlyFusion(in_channels)[source]

Bases: torch.nn.Module

1x1 convolution that fuses input channels down to 64 feature maps.

Parameters:

in_channels – number of input channels.

Create the 1x1 fusion convolution.

forward(x)[source]

Return the 64-channel fused feature map.

Parameters:

x – image-feature tensor accepted by the 1x1 convolution.

class spacr.utils.FocalLossWithLogits(alpha=1.0, gamma=2.0, reduction='mean')[source]

Bases: torch.nn.Module

Focal loss for binary, multiclass, and multilabel targets.

Auto-selects the BCE or cross-entropy branch based on the shapes of logits and target:

  • binary: logits (N,) or (N,1); target float (N,) in {0,1}.

  • multiclass: logits (N,C); target long (N,) in [0..C-1].

  • multilabel: logits (N,C); target float (N,C) in {0,1}.

Parameters:
  • alpha – class-balancing factor (float or 1-D tensor of shape (C,)).

  • gamma – focusing parameter.

  • reduction – one of 'mean', 'sum', 'none'.

Store the focal-loss hyperparameters.

forward(logits, target)[source]

Return the focal loss value for the chosen reduction mode.

Parameters:
  • logits – unnormalized binary, multiclass, or multilabel scores.

  • target – labels shaped for the corresponding logits branch.

class spacr.utils.GradCAM(model, target_layers=None, use_cuda=True)[source]

Named-hook Grad-CAM implementation for arbitrary target layers.

Parameters:
  • model – trained model to inspect.

  • target_layers – list of dotted layer names to hook.

  • use_cuda – if true, move the model and inputs to CUDA unconditionally; the caller must ensure CUDA is available.

Store the model and move it to CUDA if requested.

__call__(x, index=None)[source]

Return the normalized CAM heatmap for input x, targeting class index.

forward(input)[source]

Return the model output for input.

Parameters:

input – tensor passed directly to the wrapped model.

class spacr.utils.GradCAMGenerator(model, target_layer, cam_type='gradcam')[source]

Grad-CAM (and variants) map generator for binary classifiers.

Parameters:
  • model – trained model to inspect.

  • target_layer – dotted attribute path to the convolutional layer to probe.

  • cam_type – variant identifier, e.g. 'gradcam'.

Store the model, resolve the target layer, and register activation/gradient hooks.

compute_gradcam_and_predictions(X)[source]

Return Grad-CAM maps and predictions for every sample.

Parameters:

X – differentiable input image batch to classify and probe.

compute_gradcam_maps(X, y)[source]

Return a normalized Grad-CAM map for one sample.

Parameters:
  • X – single-sample differentiable input batch to probe.

  • y – binary label selecting the signed output score.

get_layer(model, target_layer)[source]

Resolve a dotted attribute path into the referenced submodule.

Parameters:
  • model – root model from which attribute traversal starts.

  • target_layer – dot-separated submodule attribute path.

hook_layers()[source]

Register forward/backward hooks that capture activations and gradients.

percentile_normalize(img, lower_percentile=2, upper_percentile=98)[source]

Per-channel percentile-normalize img into [0, 1].

Parameters:

img – channels-last image array to normalize.

plot_activation_grid(X, gradcam, predictions, overlay=True, normalize=False)[source]

Render a grid overlaying Grad-CAM maps on inputs with predicted-class labels.

The grid is always eight columns wide with ceil(N / 8) rows, and unused panels in an incomplete last row are hidden. The figure is returned, never shown.

Parameters:
  • X – batch tensor shaped (N, C, H, W); N fixes the grid size, and an empty batch raises ValueError from subplots. The pixels are read only under overlay, where the sample is permuted to (H, W, C) – C of 1, 3 or 4 renders, C of 2 raises TypeError from imshow.

  • gradcam – torch tensor with at least N entries. Each map may be (H, W), (1, H, W) or (3, H, W), with the same shape normalization used by the saliency twin. It is indexed and moved to the CPU on every iteration, so it must be a tensor either way.

  • predictions – sequence supporting predictions[i].item(), whose scalar is stamped in each panel’s corner. A plain Python list of ints raises AttributeError.

  • overlay – true draws the input beneath a translucent map; false draws the map alone. Default True.

  • normalize – percentile-stretch the input image only; the Grad-CAM map is always drawn raw. Has no effect unless overlay is true, and a channel that is flat between its 2nd and 98th percentiles divides by zero and comes out NaN. Default False.

Returns:

the Matplotlib Figure.

class spacr.utils.IntegratedGradients(model)[source]

Compute integrated-gradients attributions for a classifier.

Parameters:

model – trained PyTorch model.

Store the model and switch it to eval mode.

generate_integrated_gradients(input_tensor, target_label_idx, baseline=None, num_steps=50)[source]

Return integrated gradients from baseline to input_tensor for target_label_idx.

Parameters:
  • input_tensor – input sample tensor.

  • target_label_idx – target class index whose logit is attributed.

  • baseline – reference tensor (defaults to zeros of the same shape).

  • num_steps – number of Riemann-sum interpolation steps.

Returns:

attribution ndarray with the shape of input_tensor.

class spacr.utils.MultiScaleBlockWithAttention(in_channels, out_channels)[source]

Bases: torch.nn.Module

Dilated conv block followed by a 1x1 attention convolution.

Parameters:
  • in_channels – input channel count.

  • out_channels – output channel count.

Build the dilated convolution and 1x1 spatial-attention convolution.

custom_forward(x)[source]

Apply dilated conv + ReLU followed by the 1x1 spatial attention.

Parameters:

x – input feature map for the convolutional block.

forward(x)[source]

Forward pass; delegates to custom_forward().

Parameters:

x – input feature map for the convolutional block.

class spacr.utils.ResNet(resnet_type='resnet50', dropout_rate=None, use_checkpoint=False, init_weights='imagenet')[source]

Bases: torch.nn.Module

ResNet backbone with a two-layer spaCR binary-classification head.

Parameters:
  • resnet_type – one of 'resnet18'/'resnet34'/'resnet50'/'resnet101'/'resnet152'.

  • dropout_rate – dropout probability before the final linear layer; None disables.

  • use_checkpoint – enable gradient checkpointing through the ResNet backbone.

  • init_weights – 'imagenet' for pretrained weights or 'none' for random init.

Raises:

ValueError – if resnet_type is unsupported, or if init_weights is neither 'imagenet' nor 'none'.

Select the backbone and delegate head construction to initialize_base().

forward(x)[source]

Return the flattened single-logit prediction for input batch x.

Parameters:

x – image batch for the configured ResNet backbone.

initialize_base(base_model_dict, dropout_rate, use_checkpoint, init_weights)[source]

Build the backbone (with or without pretrained weights) and the two-layer head.

Parameters:
  • base_model_dict – dict with keys func (model constructor) and weights.

  • dropout_rate – dropout probability applied between the two linear layers.

  • use_checkpoint – enable gradient checkpointing through the backbone.

  • init_weights – 'imagenet' or 'none'.

Raises:

ValueError – if init_weights is neither 'imagenet' nor 'none'.

class spacr.utils.SaliencyMapGenerator(model)[source]

Generate saliency maps and predictions for a binary classifier.

Parameters:

model – trained PyTorch model with a single-logit binary output.

Store the model to be probed.

compute_saliency_and_predictions(X)[source]

Return saliency maps and the model’s own predicted classes.

Parameters:

X – differentiable input image batch to classify and probe.

compute_saliency_maps(X, y)[source]

Return absolute-gradient saliency maps for inputs X.

Parameters:
  • X – differentiable input image batch to probe.

  • y – binary labels selecting the signed output scores.

percentile_normalize(img, lower_percentile=2, upper_percentile=98)[source]

Per-channel percentile-normalize img into [0, 1].

Parameters:

img – channels-last image array to normalize.

plot_activation_grid(X, saliency, predictions, overlay=True, normalize=False)[source]

Render a grid overlaying saliency maps on inputs with predicted-class labels.

The grid is always eight columns wide with ceil(N / 8) rows, and unused panels in an incomplete last row are hidden. The figure is returned, never shown.

Parameters:
  • X – batch tensor shaped (N, C, H, W); N fixes the grid size, and an empty batch raises ValueError from subplots. The pixels are read only under overlay, where the sample is permuted to (H, W, C) – C of 1, 3 or 4 renders, C of 2 reaches imshow as an invalid shape and raises TypeError.

  • saliency – torch tensor of at least N entries. It is indexed and moved to the CPU on every iteration even when overlay is false, so a numpy array raises AttributeError either way. Each entry may be (H, W), (1, H, W) or (3, H, W); a leading singleton is removed and a leading RGB dimension is transposed to channels-last. Other three-dimensional shapes reach imshow and raise TypeError.

  • predictions – sequence supporting predictions[i].item(), whose scalar is stamped in each panel’s corner. A plain Python list of ints raises AttributeError.

  • overlay – true draws the input beneath a translucent map; false draws the map alone. Default True.

  • normalize – percentile-stretch the input image only; the saliency map is always drawn raw. Has no effect unless overlay is true, and a channel that is flat between its 2nd and 98th percentiles divides by zero and comes out NaN. Default False.

Returns:

the Matplotlib Figure.

class spacr.utils.ScaledDotProductAttention(d_k)[source]

Bases: torch.nn.Module

Standard scaled dot-product attention layer.

Parameters:

d_k – dimensionality of key/query vectors used in the scaling factor.

Store d_k used to scale attention logits.

forward(Q, K, V)[source]

Return softmax(QK^T / sqrt(d_k)) V.

Parameters:
  • Q – query tensor.

  • K – key tensor.

  • V – value tensor.

Returns:

attention-weighted value tensor.

class spacr.utils.SelectChannels(channels)[source]

Callable transform that zeroes out image channels not present in channels.

Parameters:

channels – iterable of 1-based channel indices to keep (1=red, 2=green, 3=blue).

Store the list of channels to preserve.

__call__(img)[source]

Return img with unselected RGB channels zeroed.

class spacr.utils.SelfAttention(in_channels, d_k)[source]

Bases: torch.nn.Module

Linear-projected self-attention layer.

Parameters:
  • in_channels – input feature dimension.

  • d_k – projected key/query/value dimension.

Build the Q/K/V projections and the underlying attention layer.

forward(x)[source]

Return self-attention over x of shape (B, in_channels).

Parameters:

x – batch of input feature vectors.

class spacr.utils.SpatialAttention(kernel_size=7)[source]

Bases: torch.nn.Module

Spatial attention gate that reweights features by pooled channel statistics.

Parameters:

kernel_size – convolution kernel width used to fuse average+max pooled maps.

Build the fusion convolution and sigmoid gate.

forward(x)[source]

Return the spatial attention map for x in [0, 1].

Parameters:

x – feature map whose channel statistics define the attention.

class spacr.utils.TorchModel(model_name: str = 'resnet50', pretrained: bool = True, dropout_rate: float | None = None, use_checkpoint: bool = False, num_classes: int = 2, multilabel: bool = False, image_size: int = 224)[source]

Bases: torch.nn.Module

Thin wrapper around TorchVision classification backbones that:
  1. Loads a requested backbone with (optional) pretrained weights

  2. Strips its classification head to expose features

  3. Adds a simple Linear ‘spacr’ classifier with num_classes outputs

  4. Optionally applies dropout before the final classifier

  5. Supports gradient checkpointing

Works with most TorchVision classification models. Non-classification (detection/segmentation) models are rejected with a clear error.

Build the backbone, strip its head, and attach the spaCR linear classifier.

Parameters:
  • model_name – TorchVision classification model to load.

  • pretrained – use ImageNet-pretrained weights when available.

  • dropout_rate – dropout probability applied to backbone and spaCR head; None disables.

  • use_checkpoint – enable gradient checkpointing through the backbone.

  • num_classes – output class count; 1 yields a BCE-style binary head.

  • multilabel – informational flag consumed by external loss/metrics code.

  • image_size – square input resolution used for the dummy forward pass that infers the backbone’s feature width.

Raises:

ValueError – if model_name is not a TorchVision model.

forward(x: torch.Tensor) → torch.Tensor[source]

Return classification logits of shape (N, num_classes).

Parameters:

x – input image batch for the configured TorchVision backbone.

class spacr.utils.TorchModel_v2(model_name: str = 'resnet50', pretrained: bool = True, dropout_rate: float = None, use_checkpoint: bool = False, num_classes: int = 2, multilabel: bool = False)[source]

Bases: torch.nn.Module

TorchVision backbone with a spaCR linear head (streamlined variant of TorchModel).

Parameters:
  • model_name – TorchVision classification model to load.

  • pretrained – use ImageNet-pretrained weights when available.

  • dropout_rate – dropout probability applied to backbone and spaCR head; None disables.

  • use_checkpoint – enable gradient checkpointing through the backbone.

  • num_classes – output class count.

  • multilabel – informational flag consumed by external loss/metrics code.

Raises:

ValueError – if model_name is not a TorchVision model.

Build the backbone, strip its head, and attach the spaCR classifier.

forward(x: torch.Tensor) → torch.Tensor[source]

Return classification logits of shape (N, num_classes).

Parameters:

x – input image batch for the configured TorchVision backbone.

spacr.utils.MLR(merged_df, refine_model)[source]

Fit a multiple-linear regression on gene:grna interactions plus plate/row/column terms.

Parameters:
  • merged_df – DataFrame with gene, grna, plate, row, column, pred columns.

  • refine_model – refit after removing outliers by residuals and Cook’s distance.

Returns:

tuple (max_effects, max_effects_pvalues, model, df).

spacr.utils.activation_correlations_to_database(df, img_paths, source_folder, settings)[source]

Merge per-image correlation stats with parsed well IDs and insert into the dataset DB.

Parameters:
  • df – DataFrame of correlation stats indexed by file_name.

  • img_paths – iterable of PNG paths matching rows of df.

  • source_folder – experiment root; DB written to measurements/<dataset>.db.

  • settings – settings dict; must contain dataset and cam_type.

Returns:

None.

spacr.utils.activation_maps_to_database(img_paths, source_folder, settings)[source]

Insert activation-map PNG paths and parsed well IDs into the dataset DB.

Parameters:
  • img_paths – iterable of PNG paths for activation-map images.

  • source_folder – experiment root; DB written to measurements/<dataset>.db.

  • settings – settings dict; must contain dataset and cam_type.

Returns:

None.

spacr.utils.add_column_to_database(settings)[source]

Adds a new column to the database table by matching on a common column from the DataFrame. If the column already exists in the database, it adds the column with a suffix. NaN values will remain as NULL in the database.

Parameters:

settings (dict) – A dictionary containing the following keys: csv_path (str): Path to the CSV file with the data to be added. db_path (str): Path to the SQLite database (or connection string for other databases). table_name (str): The name of the table in the database. update_column (str): The name of the new column in the DataFrame to add to the database. match_column (str): The common column used to match rows.

Returns:

None

spacr.utils.add_images_to_tar(paths_chunk, tar_path, total_images)[source]

Add paths_chunk images to tar_path, updating the shared counter for progress.

Parameters:
  • paths_chunk – list of image paths to add.

  • tar_path – destination tar archive path.

  • total_images – overall image count used to render progress.

Returns:

None.

spacr.utils.adjust_cell_masks(parasite_folder, cell_folder, nuclei_folder, organelle_folder=None, overlap_threshold=5, perimeter_threshold=30, n_jobs=None, *, output_folder=None)[source]

Run process_mask_file_adjust_cell() in parallel across matching mask files.

Parameters:
  • parasite_folder – folder of parasite masks.

  • cell_folder – folder of cell masks (overwritten in place).

  • nuclei_folder – folder of nuclei masks.

  • organelle_folder – optional folder of organelle masks.

  • overlap_threshold – fractional overlap threshold used by the merger.

  • perimeter_threshold – shared-perimeter threshold used by the merger.

  • n_jobs – worker count; None defaults to cpu_count() - 2 and values below two run inline without starting a child process.

  • output_folder – optional separate folder for adjusted masks. None preserves the historical in-place behavior. A separate folder keeps all source masks byte-identical, and every selected field is rebuilt from its source on each invocation, including after interrupted work. In place, each adjusted mask’s SHA-256 is recorded in ADJUSTED_CELLS_LEDGER as it lands, and a mask whose bytes still match its record is left alone on the next run: adjusting an adjusted mask merges it again, so a re-run used to change the result every time. A cell mask segmented again no longer matches and is adjusted afresh.

Returns:

None.

Raises:

ValueError – if the three folders contain different numbers of files or mismatched filenames, or a mask is truncated, nonnumeric, empty, not two-dimensional or has different dimensions from its partners. Available organelle masks are checked too. Header-only validation of every field finishes before any mask is changed or workers are started. An explicit output folder must differ from every source folder and must not contain masks outside the selected field set.

spacr.utils.all_elements_match(list1, list2)[source]

Return True if every element of list1 is contained in list2.

Parameters:
  • list1 – iterable of items to test.

  • list2 – iterable acting as the reference set.

Returns:

True when list1 is a subset of list2, else False.

spacr.utils.annotate_conditions(df, cells=None, cell_loc=None, pathogens=None, pathogen_loc=None, treatments=None, treatment_loc=None)[source]

Annotate df with host cell, pathogen, treatment, and combined condition columns.

Parameters:
  • df – DataFrame to annotate; must contain rowID/columnID.

  • cells – host cell types (str or list).

  • cell_loc – per-cell-type list-of-lists of row/column identifiers.

  • pathogens – pathogens (str or list).

  • pathogen_loc – per-pathogen list-of-lists of row/column identifiers.

  • treatments – treatments (str or list).

  • treatment_loc – per-treatment list-of-lists of row/column identifiers.

Returns:

annotated DataFrame with host_cells, pathogen, treatment, condition columns.

spacr.utils.annotate_predictions(csv_loc)[source]

Read prediction CSV and add plate/well/field/object columns plus a cond label.

Parameters:

csv_loc – path to a predictions CSV with a path column of PNG paths.

Returns:

DataFrame enriched with parsed metadata and a cond column ('screen'/'pc'/'nc' from the plate/well convention).

spacr.utils.apply_mask(image, output_value=0)[source]

Zero out (or set to output_value) pixels outside a circular mask fit to image.

The circle is not a parameter: create_circular_mask() is called with no center or radius, so it is always the largest circle inscribed in the frame. On a non-square image that means the short side sets the radius.

Parameters:
  • image – 2-D (H, W) or 3-D (H, W, C) array. The mask is built from the first two axes and broadcast across every channel, so all channels are cropped identically.

  • output_value – fill written outside the circle. It goes through np.where, so a value the input dtype cannot hold promotes the whole result – passing np.nan to a uint16 image returns a float array, not a masked integer one. Default 0.

spacr.utils.assign_colors(unique_labels, random_colors)[source]

Return colors and their positional mapping for the unique labels.

Parameters:
  • unique_labels – cluster labels in the order assigned palette indices.

  • random_colors – iterable of color values converted to tuples.

spacr.utils.augment_classes(dst, nc, pc, generate=True, move=True, group_by='well', test_size=0.1)[source]

Augment negative and positive class images and split them into train/test folders.

Parameters:
  • dst – destination root; augmented images land under aug_nc/aug_pc and move into aug/{train,test}/{nc,pc}.

  • nc – negative-class source image paths.

  • pc – positive-class source image paths.

  • generate – run augmentation before moving files.

  • move – split augmented images into train/test folders.

  • group_by – source identity held intact across train/test. Default 'well'; 'cell' permits sibling objects on both sides.

  • test_size – requested test fraction; whole groups make it approximate.

Returns:

None.

spacr.utils.augment_dataset(dataset, is_grayscale=False)[source]

Expand dataset by 8x through rotation and horizontal reflection of every image tensor.

Parameters:
  • dataset – iterable of (tensor, label, filename).

  • is_grayscale – informational flag (retained for API compatibility).

Returns:

list of augmented (tensor, label, filename) tuples.

Raises:

TypeError – if an image is not a torch.Tensor.

spacr.utils.augment_image(image)[source]

Return a list of PIL images covering 4 rotations x 2 horizontal reflections of image.

The 8 outputs are the dihedral group of the square and include the unmodified original as element 0, so the list is an 8x expansion, not 8 extra images. Ordering is rotation-major – [0deg, 0deg flipped, 90deg, 90deg flipped, ...] – which matters if you are keeping a parallel list of labels.

Parameters:

image – a PIL image or a numpy array. Arrays are used as-is; PIL images are converted first. A 2-D grayscale input is expanded to 3 channels via cv2.cvtColor, so every result is RGB even when the input was not – channel count is not preserved. The rotations go through cv2, which expects uint8 or another OpenCV-supported dtype; the final Image.fromarray likewise rejects the float or 16-bit arrays typical of raw microscopy, so convert to 8-bit before calling. Because 90-degree rotations swap height and width, a non-square input yields images of two different shapes in the same list.

spacr.utils.augment_images(file_paths, dst)[source]

Run augment_single_image() in parallel over file_paths.

Spawn workers so they inherit no locks held by other threads in this process. Close and join the pool after mapping; terminating forked workers can hang while their exit handlers wait on inherited locks.

Parameters:
  • file_paths – iterable of source image paths.

  • dst – destination folder (created if missing).

Returns:

None.

spacr.utils.augment_single_image(args)[source]

Save six augmentations of one image (original, 90/180/270 rotations, H/V flips).

Parameters:

args – (img_path, dst) tuple.

Returns:

None.

spacr.utils.boundary_f1_score(mask_true, mask_pred, dilation_radius=1)[source]

Return the boundary F1 score between two masks with tolerance dilation_radius.

Both masks are binarized before the boundary is taken, so this scores the outline of the foreground as a whole: boundaries where two labeled objects abut are interior to that foreground and do not appear. Split/merge errors between touching cells are therefore invisible to this metric.

Parameters:
  • mask_true – reference label or binary mask, reduced to its boundary by extract_boundaries().

  • mask_pred – predicted mask of the same shape; the two boundary images are intersected element-wise, so the masks must be pixel-registered.

  • dilation_radius – half-width of the square structuring element, giving a band 2 * dilation_radius + 1 pixels wide. This is the matching tolerance – raising it forgives localization error but also thickens both boundaries, so scores rise for every model and stop being comparable across different radii. Default 1.

spacr.utils.build_loss(loss_type: str = 'ce', num_classes: int = 2, class_counts: torch.Tensor | None = None, label_smoothing: float = 0.0, focal_gamma: float = 2.0, focal_alpha: float | None = None, logit_adjust_tau: float = 0.0, asl_gamma_pos: float = 0.0, asl_gamma_neg: float = 4.0, asl_clip: float = 0.05)[source]

Return a closure loss_fn(logits, target) implementing the requested loss.

Supported loss_type values: 'ce', 'ce_smooth', 'ce_weighted', 'focal_ce', 'bce', 'focal_bce', 'logit_adjust_ce', 'asl', 'auto'. num_classes==1 selects binary (BCE variants); >=2 selects multiclass (CE variants).

Parameters:
  • loss_type – loss identifier (see above).

  • num_classes – output class count.

  • class_counts – per-class sample counts used to derive weights or logit adjustment.

  • label_smoothing – label-smoothing epsilon for ce_smooth.

  • focal_gamma – focal-loss focusing parameter.

  • focal_alpha – focal-loss class-balancing factor (float or per-class tensor).

  • logit_adjust_tau – strength of the Menon-et-al. logit adjustment; 0 disables.

  • asl_gamma_pos – asymmetric-loss gamma for positives.

  • asl_gamma_neg – asymmetric-loss gamma for negatives.

  • asl_clip – asymmetric-loss negative-probability clip.

Returns:

loss_fn(logits, target) callable returning a scalar tensor.

Raises:

ValueError – if loss_type is unknown or incompatible with num_classes.

spacr.utils.calculate_activation_correlations(inputs, activation_maps, file_names, manders_thresholds=None)[source]

Compute per-image Pearson and Manders correlations between input and activation channels.

Parameters:
  • inputs – input image batch, tensor of shape (B, C, H, W).

  • activation_maps – activation-map batch, tensor of shape (B, C, H, W) or (B, H, W).

  • file_names – file names corresponding to each image in the batch.

  • manders_thresholds – intensity percentiles used for Manders coefficients. Default [15, 50, 75].

Returns:

DataFrame with one row per image and one column per channel-pair statistic.

spacr.utils.calculate_iou(mask1, mask2)[source]

Return the intersection-over-union of two binary masks after zero-padding to a common shape.

Unlike jaccard_index(), this returns 0 rather than nan when both masks are empty, which is what makes it safe to call inside the matching loop.

Parameters:
  • mask1 – 2-D array. Any nonzero value counts as foreground, so a multi-label crop is treated as one merged object – pass a single object’s mask if you want a per-object IoU.

  • mask2 – 2-D array compared against mask1. Shapes may differ; pad_to_same_shape() zero-pads both at the bottom and right, which assumes the two masks share a top-left origin. Two crops taken from different offsets in the same image will score meaninglessly low.

spacr.utils.calculate_loss(output, target, prefer_focal=False, gamma=2.0, alpha=1.0, reduction='mean')[source]

Auto-select and return a loss for binary, multiclass, or multilabel problems.

Dispatches based on the shapes/dtypes of output and target:
  • binary: logits (N,1), float targets in {0,1} -> BCE / focal-BCE.

  • multiclass: logits (N,C), long targets (N,) -> CE / focal-CE.

  • multilabel: logits (N,C), float targets (N,C) -> BCE / focal-BCE.

Parameters:
  • output – model logits.

  • target – ground-truth labels.

  • prefer_focal – use the focal-loss variant instead of plain CE/BCE.

  • gamma – focal-loss focusing parameter.

  • alpha – focal-loss class-balancing factor.

  • reduction – one of 'mean', 'sum', 'none'.

Returns:

scalar loss tensor (or per-sample tensor when reduction='none').

spacr.utils.calculate_shortest_distance(df, object1, object2)[source]

Calculate the shortest edge-to-edge distance between two objects (e.g., pathogen and nucleus).

Parameters: - df: Pandas DataFrame containing measurements - object1: String, name of the first object (e.g., “pathogen”) - object2: String, name of the second object (e.g., “nucleus”)

Returns: - df: Pandas DataFrame with a new column for shortest edge-to-edge distance.

Parameters:
  • df – measurement frame containing centroid and Feret-diameter columns for both objects.

  • object1 – prefix of the first object’s measurement columns.

  • object2 – prefix of the second object’s measurement columns.

spacr.utils.canonicalize_measurement_columns(df)[source]

Rename legacy column spellings on an in-memory measurement frame.

The DataFrame counterpart of rename_columns_in_db(), for frames that did not come from a spaCR database and so never passed through it — a CSV exported by an older release, or a frame a user assembled themselves.

Follows the same never-destructive rule: a rename whose target is already present is skipped, so a frame carrying both spellings keeps both rather than losing one to a silently dropped duplicate. The rule itself lives in spacr.schema.canonical_rename_plan(), which this and schema.canonicalise_columns both call so the two frame canonicalisers cannot drift apart again — and which folds case, because these frames are written with to_sql and SQLite compares identifiers case-insensitively.

Parameters:

df – A measurement DataFrame.

Returns:

df with legacy column names replaced (a copy is not made; the frame is renamed in place and returned).

spacr.utils.check_index(df, elements=5, split_char='_')[source]

Validate that every index label in df splits into elements parts on split_char.

Parameters:
  • df – DataFrame whose index labels are compound identifiers.

  • elements – Expected number of parts after splitting. Default 5.

  • split_char – Delimiter used to split each index label. Default '_'.

Returns:

None.

Raises:

ValueError – if any index label does not split into elements parts.

spacr.utils.check_mask_folder(src, mask_fldr, resume=False)[source]

Return True if masks in src/masks/mask_fldr still need generating.

Parameters:
  • src – experiment root containing masks/ and stack/ subfolders.

  • mask_fldr – subfolder name under masks/.

  • resume – accepted for the callers that pass it. Only structurally complete mask arrays are counted whether or not it is set, so an empty or truncated array left by an interrupted run is re-queued.

Returns:

True when the mask folder is missing or any expected stack filename lacks a complete mask. Unrelated masks cannot substitute for a missing field and do not force complete fields to run again.

spacr.utils.check_multicollinearity(x)[source]

Checks multicollinearity of the predictors by computing the VIF.

Parameters:

x – DataFrame of the design matrix – one row per observation, one column per predictor, and no response column. Every column is fed to variance_inflation_factor via x.values, so all columns must be numeric; categorical predictors have to be one-hot encoded first. Add an explicit constant column if you want the intercept accounted for, since without one the VIFs are inflated by the shared mean. Perfectly collinear columns yield inf.

spacr.utils.check_normality(series)[source]

Helper function to check if a feature is normally distributed.

This is a failure to reject at alpha 0.05, not evidence of normality: the answer is True whenever the D’Agostino-Pearson test does not find significant skew or kurtosis. Small samples therefore look normal for want of power, and very large ones fail on deviations too small to matter – which is what decides whether perform_statistical_tests() sends a feature to ANOVA or to Kruskal-Wallis.

Parameters:

series – one numeric column of observations, pooled across all groups. NaN propagates and makes the p-value NaN, which compares False against alpha and so is reported as normal – drop missing values first. Under 8 observations scipy cannot run the skew test, returns NaN, and the feature is likewise reported as normal; a constant column behaves the same way. Values are treated as one sample, so a strongly bimodal feature whose groups are each normal is judged on the mixture.

spacr.utils.check_overlap(current_position, other_positions, threshold)[source]

Return True if current_position is within threshold of any point in other_positions.

Parameters:
  • current_position – candidate point as a sequence of coordinates. Any dimensionality works as long as it matches the entries of other_positions; a genuine length mismatch raises ValueError from the subtraction, while a length-1 entry broadcasts silently and yields a meaningless distance.

  • other_positions – already-placed points to test against. Scanned linearly with an early return, so cost grows with the number of placed items – this is the inner loop of the image-scatter layout. An empty sequence returns False, so the first placement always succeeds.

  • threshold – minimum center-to-center Euclidean separation, in the same units as the coordinates (data units for an embedding, not pixels or points). The comparison is strict <, so a distance exactly equal to threshold counts as not overlapping. Because it measures centers, set it to roughly the thumbnail width; half of that still lets images overlap visually.

spacr.utils.choose_model(model_type: str, device: torch.device, init_weights: bool = True, dropout_rate: float = 0.0, use_checkpoint: bool = False, channels: int = 3, height: int = 224, width: int = 224, chan_dict: dict[str, Any] | None = None, num_classes: int = 2, verbose: bool = False) → torch.nn.Module | None[source]

Instantiate a classification model by name for binary or multiclass problems.

Parameters:
  • model_type – TorchVision model name (e.g. 'resnet50', 'vit_b_16'). 'custom' passes the name check but then raises NotImplementedError, as no custom builder is wired up.

  • device – unused; the model is built on the CPU and the caller moves it.

  • init_weights – load pretrained weights when available.

  • dropout_rate – dropout probability before the classifier head (None/0 disables).

  • use_checkpoint – enable gradient checkpointing for the backbone.

  • channels – unused; the forward sanity check always feeds 3 channels.

  • height – square input resolution, forwarded as TorchModel(image_size=...); it therefore fixes the dummy-forward size used to infer the backbone feature dimension (and so the size of the classifier head) as well as both dimensions of the square forward sanity check (falls back to 224 when falsy).

  • width – unused; height sets both dimensions.

  • chan_dict – unused; reserved for the unimplemented custom branch.

  • num_classes – output class count; 1 yields a single-logit BCE head.

  • verbose – print the model structure when True.

Returns:

The instantiated nn.Module.

Raises:

ValueError – model_type names no backbone, or the built model does not produce logits of the requested shape.

Unsupported names raise immediately and include close TorchVision matches when available, so configuration errors are reported before training.

spacr.utils.class_visualization(target_y, model_path, dtype, img_size=224, channels=None, l2_reg=0.001, learning_rate=25, num_iterations=100, blur_every=10, max_jitter=16, show_every=25, class_names=None)[source]

Synthesize an input image that maximizes the classifier score for target_y.

Parameters:
  • target_y – target class index.

  • model_path – path to the trained model checkpoint.

  • dtype – tensor dtype; overridden internally based on CUDA availability.

  • img_size – square image size (pixels).

  • channels – input channels (defaults to [0, 1, 2]).

  • l2_reg – L2 regularization weight on the pixel norm.

  • learning_rate – gradient-ascent step size.

  • num_iterations – optimization iteration count.

  • blur_every – interval (iterations) between periodic Gaussian blurs.

  • max_jitter – maximum pixel jitter applied per iteration.

  • show_every – interval (iterations) between preview plots.

  • class_names – display names for the classes (defaults to ['nc', 'pc']).

Returns:

deprocessed image as a numpy array.

spacr.utils.classification_metrics(all_labels, prediction_pos_probs, loss, epoch)[source]

Return a one-row DataFrame of accuracy, PR-AUC, and optimal-threshold stats.

Parameters:
  • all_labels – ground-truth binary labels.

  • prediction_pos_probs – predicted positive-class probabilities.

  • loss – loss tensor for the epoch (.item() is called).

  • epoch – epoch number used as the row index.

Returns:

DataFrame indexed by epoch with accuracy, per-class accuracy, loss, PR-AUC, and optimal threshold columns.

Raises:

ValueError – if all_labels and prediction_pos_probs have different lengths.

spacr.utils.cleanup_pipeline_folders(src, keep_intermediate=False, keep_original=False, verbose=True)[source]

Delete the intermediate mask-pipeline folders once merged/ is built.

By default spaCR keeps only merged/ (the concatenated image+mask arrays that Measure reads). This removes stack/ + masks/ (their data is embedded in merged/ and object labels are recorded in the database) and the raw orig/ backup, unless the caller opts to keep them.

Heavily guarded so it never destroys un-merged data: stack/ + masks/ are only removed when merged/ is non-empty AND every stack/*.npy has a matching merged/*.npy (i.e. every field of view was merged).

Parameters:
  • src – run root folder (holds merged/, stack/, masks/, orig/).

  • keep_intermediate – keep stack/ + masks/ when True.

  • keep_original – keep the raw orig/ backup when True.

Returns:

list of folder paths that were deleted.

spacr.utils.close_file_descriptors()[source]

Close file descriptors from 3 up to the soft NOFILE limit.

spacr.utils.close_multiprocessing_processes()[source]

Terminate all detected multiprocessing child processes and close file descriptors.

spacr.utils.cluster_feature_analysis(all_df, cluster_col='cluster')[source]

Perform Random Forest feature importance, ANOVA for normally distributed features, and Kruskal-Wallis for non-normally distributed features. Combine results into a single DataFrame.

Parameters:
  • all_df – DataFrame holding the numeric feature columns and cluster_col, one row per object. The same frame is passed to the Random Forest and to the statistical tests, so the two rankings describe the same rows. The features are selected by schema.model_feature_columns(..., allow_unknown=True), which means stray numeric bookkeeping columns (row/column indices, object IDs) are picked up as features unless you drop them first. Because each feature is routed to either ANOVA or Kruskal-Wallis, the merged output has exactly one of the two p-value pairs filled per row and NaN in the other.

  • cluster_col – column holding the group label, excluded from the features and used as the Random Forest target and the grouping variable for the tests. It needs at least two distinct labels, and every group needs enough rows for the test to run. DBSCAN’s -1 noise label is not special-cased here, so it is analysed as if it were a real cluster – strip it with remove_noise() first if you do not want that. Default 'cluster'.

spacr.utils.combine_results(rf_df, anova_df, kruskal_df)[source]

Combine the results into a single DataFrame.

All three frames are keyed on Feature and carry exactly one row per feature: rf_df is built from the feature list, and perform_statistical_tests() sends each feature to either ANOVA or Kruskal-Wallis, never both. Hence one_to_one. A repeated Feature – the signature of a frame with duplicated column names, or of two runs’ results concatenated by mistake – would multiply the importance rows and report the same feature several times as if independently ranked.

Parameters:
  • rf_df – random-forest results keyed uniquely by Feature.

  • anova_df – ANOVA results keyed uniquely by Feature.

  • kruskal_df – Kruskal-Wallis results keyed uniquely by Feature.

spacr.utils.compute_ap_over_iou_thresholds(true_masks, pred_masks, iou_thresholds)[source]

Return the area under the precision-recall curve swept over iou_thresholds.

Parameters:
  • true_masks – sequence of per-object ground-truth masks, one entry per object – not a single label image. Its length is the ground-truth count used for recall, so filtering objects out changes the denominator.

  • pred_masks – sequence of per-object predicted masks, matched greedily against true_masks at each threshold. Matching walks predictions in the order given and claims the first free true mask that clears the threshold, so the ordering can change which pairs form.

  • iou_thresholds – iterable of IoU cutoffs to sweep. The curve is the trapezoid over the resulting points sorted by recall, so a single threshold gives an area of 0 – pass at least two (COCO convention is np.linspace(0.5, 0.95, 10)). Duplicate thresholds contribute zero-width segments and do not count.

Raises:

ValueError – if a computed precision or recall falls outside [0, 1], which indicates the mask counts disagree with the matches.

spacr.utils.compute_average_precision(matches, num_true_masks, num_pred_masks)[source]

Return (precision, recall) given match count, true count, and predicted count.

Despite the name this computes a single precision/recall point, not an averaged precision; compute_ap_over_iou_thresholds() is what integrates those points into an AP.

Parameters:
  • matches – the pair list from match_masks(). Only its length is used, so any sized container works, but it must be the matched pairs rather than all candidate pairs – matching is greedy and one-to-one, so the length is the true-positive count.

  • num_true_masks – total ground-truth objects; drives false negatives as num_true_masks - len(matches). Passing a count smaller than the match count silently yields a recall above 1, which compute_ap_over_iou_thresholds() rejects with ValueError.

  • num_pred_masks – total predicted objects, used the same way for false positives. Both counts are the totals for the whole field, not per class. Zero denominators return 0 instead of raising.

spacr.utils.compute_irm_penalty(losses, dummy_w, device)[source]

Return the IRM penalty as the sum of squared gradient dot-products across environments.

Parameters:
  • losses – per-environment loss tensors.

  • dummy_w – scalar dummy weight used for gradient computation.

  • device – torch device on which to compute the penalty.

Returns:

scalar IRM penalty value.

spacr.utils.compute_segmentation_ap(true_masks, pred_masks, iou_thresholds=np.linspace(0.5, 0.95, 10))[source]

Return the COCO-style segmentation AP by matching connected components across IoU thresholds.

This is the whole-image entry point: unlike compute_ap_over_iou_thresholds() it takes label images and splits them into objects itself.

Parameters:
  • true_masks – ground-truth label or binary image for one field. It is re-run through label(), so existing IDs are discarded and touching objects that share a border merge into one component – the AP is computed on connected components, not on the IDs you supply.

  • pred_masks – predicted mask for the same field, treated identically. Each object is reduced to its bounding-box crop by regionprops, so objects are compared shape-to-shape with their positions dropped; two identically shaped cells in different corners score as a perfect match.

  • iou_thresholds – IoU cutoffs to sweep. Default np.linspace(0.5, 0.95, 10) is the COCO sweep. This default array is evaluated once at import and shared by every call, so do not mutate it in place.

spacr.utils.console_can_encode(text, stream=None)[source]

Return True when text can be printed to stream as-is.

Parameters:
  • text – the string about to be printed.

  • stream – text stream to test against; defaults to sys.stdout.

Returns:

bool.

spacr.utils.console_encoding(stream=None)[source]

Return the codec text printed to stream has to survive.

Parameters:

stream – a text stream; defaults to sys.stdout.

Returns:

a codec name, 'utf-8' when the stream does not declare one (a queue-backed GUI console, a StringIO, a captured pipe).

spacr.utils.console_safe(text, stream=None)[source]

Return text with anything the console cannot encode replaced by ?.

Console decoration must never be able to end a run. No Windows codepage encodes spaCR’s own output set – ▸ (U+25B8) is absent from cp1252, cp437, cp850, cp932 and cp936, and the box-drawing frame is absent from cp1252 – and neither does any of them encode the domain vocabulary that ends up in settings values, such as the parental strain Δku80 or a µm voxel size. Printing either to a non-UTF-8 stream raises UnicodeEncodeError, and on Windows that is the normal case the moment stdout is redirected: a batch-queue job, spacr-run, a legacy console.

Parameters:
  • text – the string about to be printed.

  • stream – text stream to encode against; defaults to sys.stdout.

Returns:

text unchanged when it is printable, otherwise a lossy but printable version of it.

spacr.utils.control_filelist(folder, mode='columnID', values=None)[source]

Return filenames in folder whose row or column ID matches one of values.

The filename is split on _ and the second token is inspected: characters after the first (mode='columnID') or the leading character (mode='rowID') are matched against values.

Parameters:
  • folder – Directory to scan.

  • mode – 'columnID' matches trailing digits, 'rowID' matches leading letter. Default 'columnID'.

  • values – Iterable of allowed ID strings. Defaults to ['01', '02'].

Returns:

List of matching filenames.

spacr.utils.convert_and_relabel_masks(folder_path)[source]

Converts all int64 npy masks in a folder to uint16 with relabeling to ensure all labels are retained.

Parameters: - folder_path (str): The path to the folder containing int64 npy mask files.

Returns: - None

Parameters:

folder_path – directory containing .npy masks to inspect and convert in place.

spacr.utils.copy_images_to_consolidated(image_path_map, root_folder)[source]

Copies images from their original locations to a ‘consolidated’ folder, renaming them according to the generated dictionary.

Parameters:
  • image_path_map (dict) – Dictionary mapping {original_path: new_path}.

  • root_folder (str) – The root directory where the ‘consolidated’ folder will be created.

spacr.utils.correct_masks(src)[source]

Convert cell masks under src/masks/cell_mask_stack to uint16 and re-stack arrays.

Relabels masks so they fit in uint16 and then re-concatenates the four array folders under src in the layout expected downstream.

Parameters:

src – Root folder of a spacr run containing a masks/ subfolder.

Returns:

None.

spacr.utils.correct_metadata(df)[source]

Normalize a metadata DataFrame to the canonical spaCR names and plate ids.

One call into spacr.schema.canonicalise_frame(), which is the whole vocabulary in one place:

  • every legacy spelling renamed, folding case and punctuation, so Plate/PLATE/plate/plateid/plate_name all arrive as plateID and the same for the row, the column, the field and the well;

  • one column per key. A file carrying well and wellID has two opinions about which well a row came from; they are compared row by row as stripped strings (so 1, 1.0 and ' 1 ' agree and a dtype difference is not a disagreement), one is kept, and the rest are dropped. Agreement prints; disagreement warns and says how many rows differ.

  • the pp plate repair, over every column that embeds the plate id.

THE pp REPAIR RUNS AFTER THE RENAMES, and over every plate-bearing column. It used to run first and only over plateID/prcfo, which meant it did nothing at all for the files that actually carry the artifact: a legacy score CSV has a plate column and NO plateID column, so the guard was false, the repair was skipped, and the very next line then copied the unrepaired plate value into plateID.

The cost was silent and total. Score files stamped pplate1 met count files stamped plate1, every prc differed by one character, the join produced ZERO rows, and the run died several steps later inside a plot with KeyError: 0 – nowhere near the mismatch, and with nothing on screen naming a plate.

Parameters:

df – Metadata DataFrame that may still use legacy naming.

Returns:

The DataFrame with canonical columns.

spacr.utils.correct_metadata_column_names(df)[source]

Renamed legacy metadata columns. Defined in spacr.schema.

Re-exported here because every existing caller imports it from utils, and moved there because importing this module costs torch, torchvision and cv2 – 6.7 seconds – for a function that needs none of them.

Parameters:

df – tabular frame whose legacy metadata names are canonicalized.

spacr.utils.correct_paths(df, base_path, folder='data')[source]

Rewrite PNG paths (in a DataFrame or list) so they live under base_path/folder.

A non-string entry is passed through untouched. png_list is LEFT-joined onto the object tables, so any object whose crop was never written arrives here with png_path = NaN – a state spacr.io._read_and_join_tables() documents as healthy (len(merged) == len(cell) > len(png_list): save_png off for a field, a crop that failed to write, an interrupted run, or a cell_id that could not be migrated). Testing base_path not in path on that NaN raised TypeError: argument of type 'float' is not iterable and took the whole embedding down over one missing thumbnail. There is no path to re-anchor for such a row, and it has to keep its position so the rewritten column still aligns with df.

Delegate rewriting to spacr.crops.reanchor_path(), which handles same-platform moves, Windows paths read on Linux, and old absolute paths. It finds the rightmost anchor component after normalizing separators and checks existing roots component by component. Paths with no matching folder component pass through unchanged; their count and one example are printed so unresolved paths remain visible.

Parameters:
  • df – DataFrame with a png_path column, or a list of paths.

  • base_path – destination root to prepend.

  • folder – intermediate folder name that anchors the rewrite.

Returns:

DataFrame + list, or list, mirroring the input type.

spacr.utils.count_reads_in_fastq(fastq_file)[source]

Return the number of reads in a gzipped FASTQ file.

Counts total lines and divides by four (the FASTQ record length).

Parameters:

fastq_file – Path to a .fastq.gz file.

Returns:

Integer read count.

spacr.utils.create_circular_mask(h, w, center=None, radius=None)[source]

Return a boolean circular mask of shape (h, w) centered on center.

Parameters:
  • h – image height.

  • w – image width.

  • center – (x, y) center; defaults to the image middle.

  • radius – circle radius; defaults to the largest circle fitting inside.

Returns:

boolean ndarray where True marks pixels within radius.

spacr.utils.debug(enabled=True, logger_name=None)[source]

Decorator that temporarily sets the given logger to DEBUG for the wrapped call.

Parameters:
  • enabled – no-op when False.

  • logger_name – logger name to tweak; defaults to the function’s module logger.

Returns:

decorator function.

spacr.utils.delete_folder(folder_path)[source]

Recursively delete folder_path if it exists (files and subdirectories included).

Parameters:

folder_path – directory to remove, contents and all. A missing path or a plain file is reported on stdout and ignored – the function never raises for those, so it cannot be used to confirm that a delete happened; check with os.path.isdir afterwards if that matters. Deletion is unconditional and unprompted, with no trash or dry-run, so a wrong path is not recoverable. Symlinked subdirectories are not descended into but are still handed to os.rmdir, which raises on a symlink, so a tree containing one aborts part-way through.

spacr.utils.delete_intermedeate_files(settings)[source]

Remove intermediate per-channel and stack folders under settings['src'].

Safeguarded to only run when a merged/ folder is present and the orig/ backup folder exists, so raw inputs are preserved.

Parameters:

settings – Dict with an 'src' key naming the run’s root folder.

Returns:

None.

spacr.utils.dense_mask_channel_positions(settings)[source]

Map each RAW channel index to its position on the merged stack’s axis.

Built the same way io.preprocess_img_data builds the stack, because that is the only thing that makes the answer true: walk the roles in MASK_CHANNEL_ROLE_ORDER and give each newly seen raw channel the next dense position.

THE TRAP THIS EXISTS TO CLOSE. Several callers computed the position as sorted({nucleus, cell, pathogen, organelle}) instead, which agrees with role order only when the roles happen to be in ascending channel order. With nucleus_channel=2, cell_channel=0, organelle_channel=1 the stack is [2, 0, 1] – raw channel 1 sits at position 2 – while the sorted reading says position 1, which holds the CELL image. Cellpose then segments organelles on the cell plane, silently, and every count and intensity downstream is measured from the wrong masks.

Parameters:

settings – the run settings, holding the raw *_channel keys.

Returns:

{raw_channel: dense_position}. Channels that are None or uncoercible are absent, matching the writer’s own behaviour.

spacr.utils.dice_coefficient(mask1, mask2)[source]

Return the Dice similarity of two masks, treating any nonzero value as foreground.

Parameters:
  • mask1 – array binarized with > 0, so negative values are counted as background – a signed difference image will not behave as expected.

  • mask2 – array of the same shape as mask1; like jaccard_index() there is no padding step, so shapes must already agree. Two empty masks return 1.0 here (defined as perfect agreement) rather than nan.

spacr.utils.display(*args, **kwargs)[source]

Do nothing: IPython is unavailable, so there is nowhere to display to.

THE FALLBACK IS THE POINT. IPython.display.display is imported at module scope, and IPython can be mid-init – partially imported by another thread – which makes that import raise. Letting it propagate would make importing this module fail for a reason that has nothing to do with what the module does. spaCR only calls display from notebook contexts; the Qt GUI ignores it.

Parameters:
  • args – whatever the caller would have displayed.

  • kwargs – likewise.

spacr.utils.download_models(repo_id='einarolafsson/models', retries=5, delay=5)[source]

Downloads all model files from Hugging Face and stores them in the resources/models directory within the installed spacr package.

Parameters:
  • repo_id (str) – The repository ID on Hugging Face (default is ‘einarolafsson/models’).

  • retries (int) – Number of retry attempts in case of failure.

  • delay (int) – Delay in seconds between retries.

Returns:

str – The local path to the downloaded models.

spacr.utils.estimate_class_counts(loader, num_classes: int, src=None, classes=None) → torch.Tensor[source]

Return per-class sample counts as a LongTensor of length num_classes.

When src and classes are provided the counts are taken from the file listings under src/<class>, avoiding a slow DataLoader iteration on NAS.

Parameters:
  • loader – fallback DataLoader iterated only when folder info is missing.

  • num_classes – number of output classes.

  • src – parent folder containing per-class subfolders.

  • classes – ordered class-folder names matching src.

Returns:

LongTensor of per-class counts.

spacr.utils.extract_boundaries(mask, dilation_radius=1)[source]

Return the boundary of a binary mask via morphological dilation minus erosion.

Parameters:
  • mask – label or binary mask.

  • dilation_radius – half-width of the structuring element.

Returns:

boolean boundary mask.

spacr.utils.extract_features(image_paths, resnet=resnet50)[source]

Extract features from images using a pre-trained ResNet model.

The classification head is stripped, so each image yields its pooled backbone features.

Parameters:
  • image_paths – Iterable of image paths, each loaded via load_image().

  • resnet – TorchVision model constructor (not an instance), called as resnet(pretrained=True). Default resnet50.

Returns:

(N, 2048) ndarray of pooled features, one row per image (2048 for ResNet-50; the width follows the chosen backbone).

spacr.utils.extract_tar_bz2_files(folder_path)[source]

Extracts all .tar.bz2 files in the given folder into subfolders with the same name as the tar file.

Parameters:

folder_path (str) – Path to the folder containing .tar.bz2 files.

spacr.utils.feature_columns(columns, selection)[source]

Which of columns the selection keeps. Order preserved.

The union over the selection’s members, so [1, 'morphology'] is channel 1’s intensities AND the shapes rather than the empty intersection of the two.

COLOCALISATION BELONGS TO BOTH CHANNELS IT MEASURES. A cell_channel_1_channel_2_pearsons column names two channels and survives a request for either – which is what makes “localization” reachable without a separate setting: ask for the channel and its relationships come with it.

Parameters:
  • columns – ordered column names available for selection.

  • selection – channel, morphology group, text filter, mixture, or None as accepted by feature_selection().

spacr.utils.feature_folder_name(channel_of_interest) → str[source]

A folder name for one feature selection. Safe on every filesystem.

None is all_features; a channel is channel_1; several are channels_1_2; morphology is itself; a mixture joins them in the order given; and a free-text filter is slugified, because a user may reasonably filter on mean_intensity and a column fragment can carry anything.

Parameters:

channel_of_interest – feature selection accepted by feature_selection().

spacr.utils.filepaths_to_database(img_paths, settings, source_folder, crop_mode)[source]

Insert cropped PNG filepaths and parsed well/object IDs into the measurements DB.

Parameters:
  • img_paths – iterable of PNG paths for cropped objects.

  • settings – settings dict; timelapse toggles time_id parsing.

  • source_folder – experiment root; DB is written to measurements/measurements.db.

  • crop_mode – a registered object role, including any organelle slot.

Returns:

None.

spacr.utils.fill_holes_in_mask(mask)[source]

Fill the holes inside each object of a label mask, keeping every id.

Delegates to spacr.qt.mask_engine.fill_label_holes(), the one hole filler for label images. This used to run ndimage.label over the mask first, which made every pair of touching objects one object: with Cellpose fill_in on (its default in Apply), a field of 74 adjacent cells was saved as 8. Now no object is merged or renumbered, and a hole takes the id of the object that encloses it.

Parameters:

mask (np.ndarray) – A labeled mask where each object has a unique integer value. A boolean mask is labelled by connectivity first.

Returns:

np.ndarray – The mask with holes filled and the original labels preserved, in the input’s dtype.

spacr.utils.filter_and_save_csv(input_csv, output_csv, column_name, upper_threshold, lower_threshold)[source]

Reads a CSV into a DataFrame, keeps the rows whose column value falls OUTSIDE the two thresholds, and saves the filtered DataFrame to a new CSV file.

The two tests are combined with OR, not AND, so this is a two-tailed selection that keeps the extremes and discards the middle. Both comparisons are strict, so a value exactly equal to either threshold is dropped, and so is NaN. Passing an upper_threshold below lower_threshold makes the two conditions cover the whole line and nothing is filtered out at all.

Parameters:
  • input_csv (str) – Path to the input CSV file, read with pd.read_csv.

  • output_csv (str) – Path to save the filtered CSV file, written without the index. Its parent directory must already exist – pandas raises OSError rather than creating it.

  • column_name (str) – Column the two comparisons are applied to. A name that is not in the frame raises KeyError, and a text column raises TypeError when compared against a numeric threshold.

  • upper_threshold (float) – Rows strictly greater than this are retained.

  • lower_threshold (float) – Rows strictly less than this are retained too; everything between the two bounds is discarded.

Returns:

None. The filtered frame is written to output_csv, shown with display for notebook users, and the destination is printed.

spacr.utils.filter_columns(df, filter_by)[source]

Return df restricted to columns matching filter_by (or morphology columns).

Parameters:
  • df – source DataFrame.

  • filter_by – substring required in column names, or 'morphology' to drop channel columns.

Returns:

column-filtered DataFrame.

spacr.utils.filter_dataframe_features(df, channel_of_interest, exclude=None, remove_low_variance_features=True, remove_highly_correlated_features=True, verbose=False)[source]

Restrict a features DataFrame to a channel of interest and clean up correlated/low-variance columns.

Parameters:
  • df – input DataFrame.

  • channel_of_interest – int, str, list, or 'morphology' to select feature groups.

  • exclude – feature(s) to drop from the final list. A single name or any number of them; the Qt ‘Exclude’ field collects a list.

  • remove_low_variance_features – apply remove_low_variance_columns().

  • remove_highly_correlated_features – apply remove_highly_correlated_columns().

  • verbose – print filter details.

Returns:

(filtered_df, features).

spacr.utils.find_non_overlapping_position(x, y, image_positions, threshold, max_attempts=100)[source]

Return a nearby (x, y) jittered position that does not collide with image_positions.

Parameters:
  • x – original x.

  • y – original y.

  • image_positions – previously placed points.

  • threshold – minimum allowed spacing.

  • max_attempts – retry budget before giving up.

Returns:

(x, y) tuple; original position if no non-overlapping spot is found.

spacr.utils.fishers_odds(df, threshold=0.5, phenotyp_col='mean_pred')[source]

Fisher’s exact test per mutant column against a binarized phenotype label.

Parameters:
  • df – DataFrame with per-mutant presence columns plus phenotyp_col.

  • threshold – cutoff below which phenotyp_col is called “high phenotype”.

  • phenotyp_col – name of the phenotype column.

Returns:

DataFrame with columns Mutant, OddsRatio, PValue, AdjustedPValue.

spacr.utils.format_path_for_system(path)[source]

Takes a file path and reformats it to be compatible with the current operating system.

Parameters:

path (str) – The file path to be formatted.

Returns:

str – The formatted path for the current operating system.

spacr.utils.generate_colors(num_clusters, black_background)[source]

Return a deterministic Viridis RGBA palette for cluster points.

Parameters:
  • num_clusters – how many colors to sample, evenly spaced across the Viridis range 0.08-0.92 (the extremes are trimmed so the darkest cluster stays visible on black and the lightest on white). Coerced with int() and floored at 1, so 0 or a negative still yields a one-entry palette rather than an empty one. Count the clusters you will actually plot: DBSCAN’s -1 noise label is not drawn, so including it shifts every real cluster’s color.

  • black_background – accepted for call compatibility with the plotting helpers but not used – the palette is Viridis either way. Set the background through setup_plot/theme_colors instead; changing this flag will not change the colors you get back.

spacr.utils.generate_cytoplasm_mask(nucleus_mask, cell_mask)[source]

Generates a cytoplasm mask from nucleus and cell masks.

Parameters: - nucleus_mask (np.array): Binary or segmented mask of the nucleus (non-zero values represent nucleus). - cell_mask (np.array): Binary or segmented mask of the whole cell (non-zero values represent cell).

Returns: - cytoplasm_mask (np.array): Copy of cell_mask with nucleus pixels set to 0, keeping the cell labels elsewhere (pathogens are not considered).

Parameters:
  • nucleus_mask – nucleus mask whose nonzero pixels are excluded.

  • cell_mask – labeled cell mask copied into the cytoplasm result.

spacr.utils.generate_fraction_map(df, gene_column, min_frequency=0.0)[source]

Return a wells-by-genes fraction matrix, dropping columns below min_frequency.

Parameters:
  • df – long-format DataFrame with prc, count, well_read_sum columns.

  • gene_column – column identifying the gene/guide.

  • min_frequency – drop columns whose maximum fraction is below this cutoff.

Returns:

DataFrame indexed by prc with per-gene fractions.

spacr.utils.generate_image_path_map(root_folder, valid_extensions=('tif', 'tiff', 'png', 'jpg', 'jpeg', 'bmp', 'czi', 'nd2', 'lif'))[source]

Recursively scans a folder and its subfolders for images, then creates a mapping of: {original_image_path: new_image_path}, where the new path includes all subfolder names.

Parameters:
  • root_folder (str) – The root directory to scan for images.

  • valid_extensions (tuple) – Tuple of valid image file extensions.

Returns:

dict – A dictionary mapping original image paths to their new paths.

spacr.utils.generate_path_list_from_db(db_path, file_metadata)[source]

Return all png_path values from db_path optionally filtered by file_metadata substrings.

Parameters:
  • db_path – path to the measurements SQLite DB.

  • file_metadata – substring or list of substrings to LIKE-match against png_path.

Returns:

list of PNG paths.

spacr.utils.get_cuda_version()[source]

Return the installed CUDA toolkit version as a digit-only string, or None.

Parses the nvcc --version output; the dots are stripped so 11.8 becomes "118".

Returns:

Version string without dots, or None if nvcc is missing or fails.

spacr.utils.get_db_paths(src)[source]

Return the standard measurements/measurements.db paths for one or more source roots.

Parameters:

src – plate folder, or list of plate folders for a multi-plate run. A bare string is wrapped, so the return type is always a list and a single-plate caller still has to index [0]. These are the run roots that Measure wrote into, not the measurements folder itself – the measurements/measurements.db suffix is appended here. Nothing is checked for existence, so a typo produces a path that only fails later at connect time.

spacr.utils.get_files_from_dir(dir_path, file_extension='*')[source]

Return glob matches for dir_path/file_extension.

Parameters:
  • dir_path – directory to list. It is joined to the pattern rather than walked, so subdirectories are never searched and a nonexistent path yields an empty list instead of an error.

  • file_extension – despite the name this is a full glob pattern, not a suffix – pass '*.tif', not '.tif' or 'tif', or nothing matches. Matching follows the filesystem’s case sensitivity, and dotfiles are excluded by glob semantics. Default '*' returns every non-hidden entry, directories included.

spacr.utils.get_ml_results_paths(src, model_type='xgboost', channel_of_interest=1)[source]

Return the standard set of ML output paths for the given model and channel selection.

Parameters:
  • src – experiment root.

  • model_type – model identifier (used in the results folder name).

  • channel_of_interest – int, list, 'morphology', or None (aliased to all_features).

Returns:

10-tuple of paths (data, permutation, feature_importance, model_metrics, permutation_fig, feature_importance_fig, shap_fig, plate_heatmap, settings, ml_features).

Raises:

ValueError – if channel_of_interest has an unsupported type.

spacr.utils.get_paths_from_db(df, png_df, image_type='cell_png')[source]

Return rows of png_df whose path contains image_type and whose prcfo is in df.

Parameters:
  • df – DataFrame indexed by prcfo identifiers.

  • png_df – DataFrame of PNG metadata with png_path and prcfo columns.

  • image_type – substring that must appear in png_path.

Returns:

filtered subset of png_df.

spacr.utils.get_sequencing_paths(src)[source]

Return the standard sequencing/sequencing_data.csv paths for one or more source roots.

Parameters:

src – plate folder, or list of plate folders, given in the same order as the corresponding get_db_paths() call – the two lists are zipped positionally when measurements are joined to barcode counts, so a reordered list silently pairs a plate’s images with another plate’s reads. A bare string is wrapped, so the result is always a list, and no path is checked for existence.

spacr.utils.get_submodules(model, prefix='')[source]

Return all dotted submodule names of model in traversal order.

Parameters:
  • model – PyTorch module to walk.

  • prefix – optional prefix prepended to returned names.

Returns:

list of dotted submodule names.

spacr.utils.group_feature_class(df, feature_groups=None, name='compartment')[source]

Add a column tagging each feature with its compartment (or other group) label.

Matches feature names against the tokens in feature_groups and stores the result in a new column name. When name == 'channel', unmatched features are relabeled 'morphology'.

Parameters:
  • df – DataFrame with a feature column.

  • feature_groups – Iterable of substrings/regex tokens to look for in each feature name. Defaults to ['cell', 'cytoplasm', 'nucleus', 'pathogen'].

  • name – Name of the column added to df. Default 'compartment'.

Returns:

df with the new group column populated.

spacr.utils.initiate_counter(counter_, lock_)[source]

Initialize shared multiprocessing counter and lock globals.

Parameters:
  • counter – shared multiprocessing.Value counter.

  • lock – shared multiprocessing.Lock guarding the counter.

Returns:

None.

spacr.utils.invert_image(image)[source]

Return the intensity-inverted image, reflected through the dtype range.

The pivot is iinfo.min + iinfo.max, which is the dtype maximum for every unsigned dtype – so uint8 and uint16 invert exactly as they always did – and -1 for a signed one, the same convention skimage.util.invert() uses. Reflecting through the range instead of subtracting from the ceiling is what keeps a signed image in range: under int8, -100 inverts to 99 rather than to 227, which used to wrap silently to -29.

Parameters:

image – array with an integer dtype. A float or boolean image raises ValueError rather than inverting – convert or rescale to an integer dtype first. The pivot is the dtype range, not the image range, so a dim uint16 image inverts against 65535 and comes back near-white; normalize to the dtype range first if you want a contrast-preserving inversion.

Returns:

the inverted image, in the input dtype. Every value stays in range, so nothing wraps.

Raises:

ValueError – image does not have an integer dtype.

spacr.utils.is_list_of_lists(var)[source]

Return True if var is a list whose every element is also a list.

Parameters:

var – value to test, including an empty or nested list.

spacr.utils.is_multiprocessing_process(process)[source]

Return True if process cmdline contains multiprocessing.

Parameters:

process – process object exposing the psutil cmdline API.

spacr.utils.jaccard_index(mask1, mask2)[source]

Return the Jaccard/IoU index of two binary masks.

Parameters:
  • mask1 – array of any shape; nonzero is foreground, so a multi-label mask collapses to one merged object.

  • mask2 – array that must already have the same shape as mask1 – there is no padding step here, so mismatched shapes either raise a broadcast error or, worse, broadcast silently against a length-1 axis. Use calculate_iou() when the shapes can differ; it also returns 0 for two empty masks, whereas this divides by zero and returns nan with a runtime warning.

spacr.utils.lasso_reg(merged_df, alpha_value=0.01, reg_type='lasso')[source]

Fit Lasso or Ridge on one-hot-encoded gene/grna/plate/row/column predictors.

Parameters:
  • merged_df – DataFrame with gene, grna, plateID, rowID, columnID, pred.

  • alpha_value – regularization strength.

  • reg_type – 'lasso' or 'ridge'.

Returns:

DataFrame with Feature and Coefficient columns.

spacr.utils.load_image(image_path)[source]

Load and preprocess an image.

The preprocessing is fixed to the ImageNet recipe used by extract_features(): resize to 224x224, then normalize with the ImageNet channel means and standard deviations. None of it is configurable.

Parameters:

image_path – path to any file PIL can open. It is forced through convert('RGB'), so a 16-bit or float microscopy TIFF is downcast to 8-bit and a single-channel image is replicated across three channels rather than rejected – the dynamic range of a raw scientific image is lost here, so rescale to 8-bit yourself if that matters. Aspect ratio is not preserved: Resize((224, 224)) takes both dimensions, so non-square crops are stretched, not letterboxed. Returns a (1, 3, 224, 224) tensor with the batch axis already added, so it can be fed to a model directly but must be concatenated, not stacked, to batch several images.

spacr.utils.load_image_paths(c, visualize)[source]

Load the png_list table into a DataFrame indexed by prcfo and optionally filter by object.

Parameters:
  • c – open sqlite3 cursor.

  • visualize – object-type prefix ('cell'/'nucleus'/…) or falsy to keep all rows.

Returns:

DataFrame of PNG metadata indexed by prcfo.

spacr.utils.load_settings(csv_file_path, show=False, setting_key='setting_key', setting_value='setting_value')[source]

Reload a spacr settings CSV (written by save_settings()) back into a Python dict.

Every spacr pipeline persists its resolved settings alongside its outputs so that a run can be reproduced. This helper re-parses that CSV, coercing each value into its original Python type (bool, int, float, None, list, tuple, dict, str).

Parameters:
  • csv_file_path – path to the CSV file.

  • show – display the raw DataFrame for debugging. Default False.

  • setting_key – name of the key column. Default 'setting_key'; 'Key' is accepted too, see below.

  • setting_value – name of the value column. Default 'setting_value'; 'Value' is accepted too.

Returns:

dict of parsed settings, ready to pass back into the original pipeline entry point.

Raises:

ValueError – if the required key / value columns are missing.

THE TWO SPELLINGS. save_settings() writes Key / Value, while this function’s defaults ask for setting_key / setting_value – so the documented inverse pair did not round-trip, and the example above raised. Callers had each worked around it separately (spacr/qt/dnd.py tries one spelling and catches the failure to try the other), which is how it survived: nothing that used the defaults was reading a file spacr had written.

Either spelling is now read. An explicitly named column still wins, so a caller that knows its file’s header is unaffected.

Example

from spacr.utils import load_settings
from spacr.core import preprocess_generate_masks
settings = load_settings('/data/plate01/settings/gen_mask_settings.csv')
preprocess_generate_masks(settings)

See also

save_settings() — inverse operation.

spacr.utils.map_condition(col_value, neg='c1', pos='c2', mix='c3')[source]

Map a column-ID value to one of 'neg', 'pos', 'mix', or 'screen'.

Parameters:
  • col_value – Column identifier from the plate metadata.

  • neg – Column ID that corresponds to negative controls. Default 'c1'.

  • pos – Column ID that corresponds to positive controls. Default 'c2'.

  • mix – Column ID that corresponds to mixed controls. Default 'c3'.

Returns:

Condition label; any unlisted column returns 'screen'.

spacr.utils.mask_object_count(mask)[source]

Return the number of nonzero labeled objects in mask.

Parameters:

mask – integer label image where 0 is background. The count is the number of distinct nonzero values, so label IDs need not be contiguous and gaps left by filtering are not counted. A purely binary mask therefore reports 1 however many blobs it contains.

spacr.utils.match_masks(true_masks, pred_masks, iou_threshold)[source]

Greedy match each predicted mask to a still-unmatched true mask above iou_threshold.

Parameters:
  • true_masks – iterable of ground-truth masks.

  • pred_masks – iterable of predicted masks.

  • iou_threshold – minimum IoU to count as a match.

Returns:

list of (true_mask, pred_mask) matched pairs.

spacr.utils.measure_test_mode(settings)[source]

Copy a random subset of source files into a test/merged folder when test_mode is on.

Fewer files than test_nr is not an error. test_mode is the setting a user reaches for on a SMALL plate, and random.sample raised ValueError: Sample larger than population or is negative on exactly that case – so the one folder you most want to smoke-test first was the one folder test_mode refused to run on.

Only visible .npy arrays are sampled, so a macOS ._ sidecar is never measured in place of a field. The folder’s .spacr_plane_layout.json is copied across as well when it exists: it is what says which plane is which, and a test/merged without it is read as a legacy folder, against the default plane order.

Parameters:

settings – settings dict; must contain src, test_mode, test_nr.

Returns:

settings dict with src optionally redirected to the test folder.

Raises:

ValueError – if there is nothing to sample – an empty src, or a test_nr below 1. Sampling zero files would point src at an empty test/merged and the run would report “no fields found”, blaming the wrong thing.

spacr.utils.merge_dataframes(df, image_paths_df, verbose)[source]

Merge df into image_paths_df on the shared prcfo index.

Parameters:
  • df – feature DataFrame with a prcfo column.

  • image_paths_df – DataFrame indexed by prcfo.

  • verbose – display the merged DataFrame.

Returns:

merged DataFrame.

spacr.utils.merge_regression_res_with_metadata(results_file, metadata_file, name='_metadata')[source]

Merge regression outputs with gene metadata on the parsed gene column.

Parameters:
  • results_file – path to a regression results CSV with a feature column.

  • metadata_file – path to a gene metadata CSV with a Gene ID column.

  • name – suffix appended to the output filename.

Returns:

merged DataFrame (also written to <results_file><name>.csv).

spacr.utils.merge_split_objects(mask_src, intensity_img_src=None, intensity_channel=None, perimeter_fraction=0.5, min_area=0, max_area=0, remove_border_objects=False, n_jobs=1, progress_callback=None, op_name='', *, min_intensity=0, max_intensity=0, filters=None)[source]

Merge by perimeter and filter labeled objects across a directory of masks.

Runs the shared in-memory merge/filter pipeline on each mask file in mask_src in parallel, overwriting each mask in place.

Parameters:
  • mask_src – directory containing mask .tif/.tiff/.npy files.

  • intensity_img_src – directory of matched original intensity images, required when either intensity bound is enabled.

  • intensity_channel – explicit channel-last index for multi-channel intensity images; unnecessary for single-channel planes.

  • perimeter_fraction – minimum shared-boundary fraction for perimeter-based merging.

  • min_area – remove objects smaller than this (px); 0 disables.

  • max_area – remove objects larger than this (px); 0 disables.

  • remove_border_objects – drop objects touching the image border.

  • n_jobs – parallel worker count.

  • progress_callback – optional callback(fov_index, total, duration, op_name).

  • op_name – label passed to the progress callback.

  • min_intensity – remove objects whose own-channel mean is below this raw-image value; equality is kept and 0 disables the lower bound.

  • max_intensity – remove objects whose own-channel mean is above this raw-image value; equality is kept and 0 disables the upper bound.

  • filters – object filter entries (any scalar regionprop with a minimum and a maximum), judged with the bounds above in one pass.

Returns:

None.

spacr.utils.merge_touching_objects(mask, threshold=0.25)[source]

Merge touching labeled objects whose shared boundary exceeds threshold of the smaller perimeter.

Parameters:
  • mask – labeled mask.

  • threshold – fraction of the smaller perimeter required to merge.

Returns:

merged label mask.

spacr.utils.model_metrics(model)[source]

Print RMSE/MAE/Durbin-Watson and show residual/QQ/scale-location diagnostic plots.

Parameters:

model – fitted statsmodels regression result.

Returns:

None.

spacr.utils.normalize_feature_filter(filter_by)[source]

Normalize text representations of an unfiltered feature selection.

Settings imported from CSV files and older Qt sessions can contain the literal string "None". Treating that as a feature-name substring removes every measurement column, although the UI means “all channels”.

Parameters:

filter_by – the raw setting value. A string is stripped and, if it case-insensitively matches one of the “no filter” spellings ("", "none", "null", "all", "all_channels", "all channels", "*"), collapsed to None – otherwise the stripped string is returned as a feature-name substring. Anything that is not a string (a real None, or a list of channels) passes through untouched, so this is safe to apply unconditionally to whatever the settings dict holds. Note "*" means no filter, not a glob: real patterns are not supported, and the surviving string is matched as a plain substring of the column name.

spacr.utils.normalize_src_path(src)[source]

Ensures that the ‘src’ value is properly formatted as either a list of strings or a single string.

Parameters:

src (str or list) – The input source path(s).

Returns:

list or str –

A correctly formatted list if the input was a list (or string representation of a list),

otherwise a single string.

spacr.utils.normalize_to_dtype(array, p1=2, p2=98, percentile_list=None, new_dtype=None)[source]

Percentile-normalize each channel of an image stack into the target dtype range.

Parameters:
  • array – input stack of shape (H, W, C).

  • p1 – lower percentile. Default 2.

  • p2 – upper percentile. Default 98.

  • percentile_list – per-channel (low, high) pairs; overrides p1/p2.

  • new_dtype – target dtype (np.uint8/np.uint16 or their string forms).

Returns:

normalized stack with the same shape as array.

spacr.utils.object_label_from_png_id(values)[source]

Migrate png_list’s 'o<N>' text ids onto the integer object label.

png_list stores an object id as text ('o5') because it is the last component of prcfo; every object table stores the same object as an integer object_label, and the child tables store their parent as an integer (in practice a float, since measure writes NaN for “no overlapping cell”) cell_id. Two types for one identity, which is why a plain SQL png_list.cell_id = nucleus.cell_id matches zero rows rather than failing: SQLite compares a TEXT value with an INTEGER one by type class, and text always sorts after numbers. Measured on a database built by the real writers: 6 crops, 6 nuclei, 0 rows joined.

The integer is canonical — it is what the measurement tables key on — so this is the one migration, applied on read. It replaces series.str[1:].astype(int), which crashed on four values the real writers genuinely produce:

  • 'omulti' and 'onone' — _generate_names() names a crop that overlaps several cells ..._multi.png and one that overlaps none ..._none.png. Both are ordinary outcomes of a real segmentation. ValueError: invalid literal for int() with base 10: 'multi';

  • 'error' — what _map_wells_png() writes for a name it cannot parse. .str[1:] turned it into 'rror', so the exception did not even name the problem;

  • NULL — every row of a different crop mode, in a database measured with more than one. TypeError: int() argument must be ... not 'NoneType';

  • an already-integer column, from a database whose ids were migrated elsewhere: .str raises AttributeError on a numeric Series.

All four now come back as NaN, which a caller can count and drop — losing the crop’s path for those objects, never the whole read.

Parameters:

values – a png_list object-id column (cell_id, nucleus_id, …), of any dtype.

Returns:

a float Series of object labels, NaN where the id holds no integer. Float rather than int because NaN has no int64.

spacr.utils.pad_to_same_shape(mask1, mask2)[source]

Zero-pad mask1 and mask2 to their element-wise maximum shape.

Padding is appended at the bottom and right only, so the two masks are aligned on their top-left corner. This is an alignment assumption, not a registration: crops taken from different offsets are not brought into correspondence by padding them.

Parameters:
  • mask1 – 2-D array. Only axes 0 and 1 are considered, so a 3-D stack is padded on its first two axes and left ragged on the third.

  • mask2 – 2-D array padded to the same element-wise maximum shape. Each mask is padded independently, so the larger one along a given axis is returned untouched on that axis.

spacr.utils.perform_statistical_tests(all_df, cluster_col='cluster')[source]

Perform ANOVA or Kruskal-Wallis tests depending on normality of features.

Each numeric feature is tested for normality and then sent to either ANOVA or Kruskal-Wallis across the groups defined by cluster_col, never both.

Parameters:
  • all_df – DataFrame with the numeric feature columns and the cluster column.

  • cluster_col – Column holding the cluster label, excluded from the features. Default 'cluster'.

Returns:

(anova_df, kruskal_df), with columns Feature/ANOVA_Statistic/ANOVA_pValue and Feature/Kruskal_Statistic/Kruskal_pValue respectively; together they cover the features exactly once.

spacr.utils.pick_best_model(src)[source]

Return the strongest checkpoint anywhere below src.

Current artifacts are ranked by their stored validation metric and role; legacy files fall back to their _acc_/_epoch_ filename fields.

Parameters:

src – model directory or a checkpoint path.

Returns:

absolute path to the top-ranked checkpoint.

spacr.utils.plot_clusters(ax, embedding, labels, colors, cluster_centers, plot_outlines, plot_points, smooth_lines, figuresize=10, dot_size=50, verbose=False, point_color='cluster', point_alpha=0.65, outline_width=1.0)[source]

Draw cluster outlines, points, and centroid labels onto ax for a 2-D embedding.

Parameters:
  • ax – Matplotlib axes to draw into.

  • embedding – (N, 2) array of 2-D points (e.g. UMAP output).

  • labels – length-N cluster labels; -1 denotes noise.

  • colors – iterable of per-cluster colors, one per unique label.

  • cluster_centers – iterable of (x, y) centroids, one per unique label.

  • plot_outlines – draw a hull/smoothed outline around each cluster.

  • plot_points – render the scatter points (otherwise plotted invisibly).

  • smooth_lines – use a smoothed hull polyline instead of the convex hull edges.

  • figuresize – base size in inches used to scale axis label and tick fonts. Default 10.

  • dot_size – scatter marker size in points. Default 50.

  • verbose – unused placeholder kept for API compatibility. Default False.

  • point_color – 'cluster'/'viridis' (or empty) colors points per cluster; any other Matplotlib color is applied to every point. Default 'cluster'.

  • point_alpha – scatter opacity, clamped to [0, 1]; ignored when plot_points is False. Default 0.65.

  • outline_width – hull line width in points, floored at 0.1. Default 1.0.

Returns:

None.

spacr.utils.plot_clusters_grid(embedding, labels, image_nr, image_paths, colors, figuresize, black_background, verbose, theme_colors=None)[source]

Plot a grid of example images per cluster label discovered in labels.

Parameters:
  • embedding – accepted and never read – the panels are built from labels and image_paths alone, so None works.

  • labels – cluster labels in the row order of image_paths. -1 is dropped as noise, and if nothing else remains the function prints No clusters found. and returns None instead of a figure.

  • image_nr – per-cluster cap on how many images are opened. A cluster no larger than this contributes all of its members.

  • image_paths – paths addressed positionally by the label array; a list shorter than labels raises IndexError.

  • colors – palette indexed downstream by the cluster LABEL itself rather than by its rank, so the palette has to be long enough to reach the largest label – labels 0 and 5 against a two-color palette raise IndexError. Entries need at least three components.

  • figuresize – per-cluster panel size in inches, shrunk downstream so the whole row never exceeds 200 inches.

  • black_background – picks the white-on-black fallback theme instead of black-on-white; theme_colors overrides it per role.

  • verbose – only ever prints for STRING cluster labels; silent for the integer labels DBSCAN and KMeans produce.

  • theme_colors – dict with background/foreground/border colors. Entries Matplotlib cannot parse are dropped silently and fall back to the black_background choice. Default None.

Returns:

the Matplotlib Figure, or None when every label is -1.

spacr.utils.plot_embedding(embedding, image_paths, labels, image_nr, img_zoom, colors, plot_by_cluster, plot_outlines, plot_points, plot_images, smooth_lines, black_background, figuresize, dot_size, remove_image_canvas, verbose, interactive_payload=None, theme_colors=None, point_color='cluster', point_alpha=0.65, outline_width=1.0)[source]

Plot a 2-D embedding with cluster outlines, points, and optional image overlays.

Parameters:
  • embedding – (N, 2) array of 2-D points (e.g. UMAP output).

  • image_paths – length-N image paths used for the overlays; None skips them.

  • labels – length-N cluster labels; -1 denotes noise.

  • image_nr – number of images to overlay (per cluster when plot_by_cluster).

  • img_zoom – zoom factor applied to each overlaid thumbnail.

  • colors – palette of per-cluster colors, one entry per unique label.

  • plot_by_cluster – sample the overlaid images per cluster instead of at random.

  • plot_outlines – draw a hull/smoothed outline around each cluster.

  • plot_points – render the scatter points (otherwise plotted invisibly).

  • plot_images – overlay the images from image_paths.

  • smooth_lines – use a smoothed hull polyline instead of the convex hull edges.

  • black_background – use the white-on-black default theme instead of black-on-white; entries in theme_colors override it per role.

  • figuresize – figure side length in inches; also scales label and tick fonts.

  • dot_size – scatter marker size in points.

  • remove_image_canvas – mask out zero-valued pixels of each overlaid image.

  • verbose – forwarded to the cluster and image helpers, which ignore it.

  • interactive_payload – optional object stashed on the figure as _spacr_umap_payload so the Qt bridge can keep point/image identities. Default None.

  • theme_colors – dict with background/foreground/border colors. Default None.

  • point_color – 'cluster'/'viridis' colors points per cluster; any other Matplotlib color is applied to every point. Default 'cluster'.

  • point_alpha – scatter opacity, clamped to [0, 1]. Default 0.65.

  • outline_width – hull line width in points, floored at 0.1. Default 1.0.

Returns:

matplotlib Figure.

spacr.utils.plot_grid(cluster_images, colors, figuresize, black_background, verbose, theme_colors=None)[source]

Render one column per cluster of representative images with colored borders and labels.

Parameters:
  • cluster_images – ordered mapping of cluster label to that cluster’s list of image arrays; one column per key, and an empty mapping raises ValueError from subplots.

  • colors – palette used consistently for both panel borders and legend swatches. Integer labels index it by label (wrapping when necessary), while string labels use their position in cluster_images. An empty palette falls back to neutral grey, and entries need at least three components.

  • figuresize – figure height in inches and the label font size; the width is this times the cluster count. It is shrunk to 200 / n_clusters when that product would exceed 200 inches, which silently caps the font size too.

  • black_background – picks the white-on-black fallback theme instead of black-on-white.

  • verbose – prints the label and its index for STRING cluster labels only; integer labels never print anything.

  • theme_colors – dict with background/foreground/border colors overriding the black_background fallback; values Matplotlib cannot parse are ignored. Default None.

Returns:

the Matplotlib Figure, which is also passed to plt.show.

spacr.utils.plot_image(ax, x, y, img, img_zoom, remove_image_canvas=True)[source]

Place a zoomed thumbnail of img at (x, y) on ax.

Parameters:
  • ax – axes the thumbnail is added to, as a frameless annotation box.

  • x – data-space x coordinate the thumbnail is anchored at.

  • y – data-space y coordinate the thumbnail is anchored at.

  • img – PIL image when remove_image_canvas is true, since img.mode is read; any array-like otherwise.

  • img_zoom – scale factor handed to OffsetImage. It sizes the thumbnail from the source’s pixel dimensions in display space, so the drawn size is unchanged by the axis limits.

  • remove_image_canvas – true swaps the image for an RGBA array whose alpha channel hides zero-valued pixels, which accepts only PIL modes L, I and RGB – RGBA and P raise ValueError, and a numpy array raises AttributeError because it has no mode. An all-zero L image divides by its own zero maximum and comes out NaN rather than raising. False just calls np.array. Default True.

Returns:

None.

spacr.utils.plot_images_by_cluster(ax, image_paths, embedding, labels, image_nr, img_zoom, colors, cluster_indices, remove_image_canvas, verbose)[source]

Overlay up to image_nr images per cluster on the embedding in ax.

Parameters:
  • ax – axes the thumbnails are added to, as frameless annotation boxes.

  • image_paths – paths addressed by the indices held in cluster_indices, so they must be in the embedding’s row order.

  • embedding – (N, 2) array supplying each thumbnail’s position.

  • labels – only np.unique(labels) is used, to decide which clusters to visit; -1 is skipped as noise.

  • image_nr – per-cluster cap. A cluster no larger than this contributes all of its members – no sampling happens.

  • img_zoom – scale factor handed to OffsetImage, applied to the file’s own pixel dimensions rather than to data units.

  • colors – accepted for caller compatibility but not read. Thumbnail overlays do not use a cluster color, and palette length no longer limits how many labels are visited.

  • cluster_indices – mapping of label to the row indices to draw from. Looked up with .get(label, []), so a label present in labels but absent here plots nothing instead of raising.

  • remove_image_canvas – forwarded to plot_image().

  • verbose – accepted and ignored; any value at all is tolerated.

Returns:

None.

spacr.utils.plot_umap_images(ax, image_paths, embedding, labels, image_nr, img_zoom, colors, plot_by_cluster, remove_image_canvas, verbose)[source]

Overlay sample images from image_paths on the UMAP embedding in ax.

Parameters:
  • ax – axes the thumbnails are added to, as frameless annotation boxes.

  • image_paths – paths addressed by the same positional index as embedding, so the two must share a row order; a short list raises IndexError.

  • embedding – (N, 2) array whose selected rows give each thumbnail its data-space position.

  • labels – cluster labels aligned with embedding. Read only when plot_by_cluster is true; None is accepted otherwise.

  • image_nr – with plot_by_cluster false, the exact number of rows sampled at random from the whole embedding, so a value above N raises ValueError from random.sample. With it true, a per-cluster cap – a cluster no larger than this contributes every member, unsampled.

  • img_zoom – scale factor handed to OffsetImage. It sizes the thumbnail from the file’s own pixel dimensions in display space, so rescaling the axes does not change how big the image is drawn.

  • colors – only zipped against np.unique(labels) to drive the iteration; the color itself is never drawn. Its LENGTH is therefore a silent limit, and because np.unique includes the -1 noise label a palette sized to the real clusters leaves the last cluster with no images. Unused (None is fine) when plot_by_cluster is false.

  • plot_by_cluster – true samples per cluster and skips label -1; false ignores labels and colors entirely and samples globally.

  • remove_image_canvas – forwarded to plot_image(); true masks zero-valued pixels out and restricts the inputs to PIL modes L, I and RGB.

  • verbose – accepted and ignored, here and in the helper it is passed to; any value at all is tolerated.

Returns:

None.

spacr.utils.prepare_batch_for_segmentation(batch)[source]

Cast a batch to float32 and per-image max-normalize any image whose max exceeds 1.

Parameters:

batch – (N, ...) numpy array of images.

Returns:

The same array cast to float32 with each image scaled to [0, 1].

spacr.utils.preprocess_data(df, filter_by, remove_highly_correlated, log_data, exclude, column_list=False, *, batch_correction='none', batch_column='plateID', batch_control_column=None, batch_control_values=None, batch_covariate_column=None, batch_combat_mean_only=False, batch_min_samples=3, batch_missing_control='error')[source]

Prepare a feature matrix by filtering, decorrelating, log-transforming, and scaling df.

Parameters:
  • df – input DataFrame.

  • filter_by – channel of interest passed to filter_dataframe_features(); None and its text forms disable filtering.

  • remove_highly_correlated – correlation cutoff (float) or True to use 0.95; False disables.

  • log_data – apply log(x + 1e-6) to numeric columns.

  • exclude – features to exclude from filtering.

  • column_list – optional explicit column subset applied before selecting numeric columns.

  • batch_correction – none, center, zscore, robust_zscore, control_center or combat.

  • batch_column – metadata column identifying acquisition batches.

  • batch_control_column – metadata column selecting reference controls.

  • batch_control_values – reference value(s) for control_center.

  • batch_covariate_column – metadata column(s) naming the biology combat must preserve. Required by combat, ignored by every other method — and left blank, combat refuses to run rather than removing the contrast along with the plate effect.

  • batch_combat_mean_only – correct only combat’s additive shift and leave each batch’s scale alone.

  • batch_min_samples – minimum rows/reference controls per batch.

  • batch_missing_control – error or skip when a batch lacks enough controls.

Returns:

standard-scaled ndarray of numeric features.

Raises:

ValueError – if no numeric columns remain after filtering.

spacr.utils.preprocess_image(image_path, normalize=True, image_size=224, channels=None)[source]

Load and preprocess image_path into a batched tensor ready for classification.

Parameters:
  • image_path – path to the source image.

  • normalize – apply ImageNet mean/std normalization.

  • image_size – square resize dimension.

  • channels – reserved for downstream use; kept for API compatibility.

Returns:

(pil_image, input_tensor) where the tensor has shape (1, 3, H, W).

spacr.utils.pretty_print_settings(settings, title='Settings')[source]

Print a settings dict to the console as a tidy, aligned table.

Nicer than dumping a truncated pandas DataFrame: values are grouped by the spacr settings categories, keys are aligned in a column, long values are clipped, and the whole thing sits under a boxed title. Purely cosmetic – used wherever “Saving settings” is shown.

Purely cosmetic, and it stays that way: the frame degrades to ASCII and every line goes out through console_safe(), so a console that cannot encode the decoration prints a plainer table instead of raising UnicodeEncodeError. spacr.measure.measure_crop() calls this (through save_settings()) before it does any work at all, so a decoration character was enough to end a whole run before the first field was read.

Parameters:
  • settings – the settings dict to render.

  • title – heading shown in the box.

Returns:

None.

spacr.utils.print_progress(files_processed, files_to_process, n_jobs, time_ls=None, batch_size=None, operation_type='')[source]

Print a one-line progress report with an ETA derived from mean step time.

Parameters:
  • files_processed – number of items done (int or list).

  • files_to_process – total items to do (int or list).

  • n_jobs – parallelism used to compute ETA.

  • time_ls – list of per-step durations (seconds) for ETA; None skips ETA.

  • batch_size – batch size when time_ls is per batch rather than per image.

  • operation_type – label printed alongside the progress line.

Returns:

None.

spacr.utils.process_mask_file_adjust_cell(file_name, parasite_folder, cell_folder, nuclei_folder, organelle_folder=None, overlap_threshold=5, perimeter_threshold=30, *, output_folder=None)[source]

Load one triple of parasite/cell/nuclei masks, merge cells in place, and return the elapsed time.

Parameters:
  • file_name – mask file name (must exist in all folders).

  • parasite_folder – folder of parasite masks.

  • cell_folder – folder of cell masks (overwritten in place). The adjusted mask replaces the old one atomically, so a run killed during the write leaves the previous whole mask, never a truncated one.

  • nuclei_folder – folder of nuclei masks.

  • organelle_folder – optional folder of organelle masks.

  • overlap_threshold – fractional overlap threshold used by the merger.

  • perimeter_threshold – shared-perimeter threshold used by the merger.

  • output_folder – optional separate destination for adjusted masks. None retains in-place adjustment. An explicit destination must differ from every source mask folder, including through directory symlinks.

Returns:

elapsed seconds.

Raises:

ValueError – if the matching cell or nuclei mask file is missing, or a mask file holds pickled objects: masks are plain arrays, and nothing is unpickled.

spacr.utils.process_masks(mask_folder, image_folder, channel, batch_size=50, n_clusters=2, plot=False)[source]

Cluster object morphology/intensity across a mask folder and keep the largest cluster in place.

Parameters:
  • mask_folder – folder of .npy masks.

  • image_folder – matching folder of .npy intensity images.

  • channel – channel index used for intensity measurements.

  • batch_size – number of files to load per batch.

  • n_clusters – number of KMeans clusters.

  • plot – show a PCA scatter of the clustered objects.

Returns:

None.

spacr.utils.process_vision_results(df, threshold=0.5)[source]

Split image paths into well identifiers and binarize the pred column.

Parameters:
  • df – DataFrame with path and pred columns.

  • threshold – cutoff used to derive cv_predictions.

Returns:

enriched DataFrame with plateID, rowID, columnID, fieldID, prc, cv_predictions.

spacr.utils.random_forest_feature_importance(all_df, cluster_col='cluster')[source]

Rank features by how well they predict the cluster label.

Z-scales the numeric feature columns and fits a 100-tree RandomForestClassifier against cluster_col.

Parameters:
  • all_df – DataFrame with the numeric feature columns and the cluster column.

  • cluster_col – Column holding the cluster label, excluded from the features. Default 'cluster'.

Returns:

DataFrame with Feature and Importance columns, sorted by descending importance.

spacr.utils.recommend_target_layers(model)[source]

Return ([last_conv_layer], all_conv_layers) from model.

Parameters:

model – PyTorch module to scan for Conv2d layers.

Returns:

tuple (recommended, all) of layer-name lists.

Raises:

ValueError – if the model contains no convolutional layers.

spacr.utils.reduction_and_clustering(numeric_data, n_neighbors, min_dist, metric, eps, min_samples, clustering, reduction_method='umap', verbose=False, embedding=None, n_jobs=-1, mode='fit', model=False, reducer_options=None, prefer_gpu=False, random_seed=42)[source]

Reduce numeric_data to 2-D and cluster the embedding.

Supported reducers are UMAP, t-SNE, PCA, Isomap and Spectral Embedding. reducer_options carries only method-specific settings; irrelevant options are never forwarded. RAPIDS is opt-in and applies to UMAP, t-SNE and PCA, with the actual backend retained on the fitted reducer.

Parameters:
  • numeric_data – rows of numeric features to embed and cluster.

  • n_neighbors – reducer neighborhood size, or a row fraction as a float; also supplies the default t-SNE perplexity.

  • min_dist – minimum embedding distance used by UMAP.

  • metric – distance metric used by the reducer and DBSCAN.

  • eps – DBSCAN neighborhood radius.

  • min_samples – DBSCAN minimum neighborhood size, or KMeans cluster count when clustering='kmeans'.

  • clustering – clustering algorithm, 'dbscan' or 'kmeans'.

spacr.utils.remove_canvas(img)[source]

Return img as RGBA with zero-valued pixels made transparent.

Parameters:

img – PIL image in L, I, or RGB mode.

spacr.utils.remove_highly_correlated_columns(df, threshold=0.95, verbose=False)[source]

Drop numeric columns whose absolute correlation with a prior column exceeds threshold.

Parameters:
  • df – input DataFrame.

  • threshold – correlation cutoff.

  • verbose – print the dropped column names.

Returns:

decorrelated DataFrame.

spacr.utils.remove_intensity_objects(image, mask, intensity_threshold, mode)[source]

Drop labeled objects whose mean intensity is on the wrong side of intensity_threshold.

Parameters:
  • image – intensity image.

  • mask – labeled mask aligned to image.

  • intensity_threshold – cutoff value.

  • mode – 'low' removes below-threshold objects, 'high' removes above.

Returns:

filtered label mask.

spacr.utils.remove_low_variance_columns(df, threshold=0.01, verbose=False)[source]

Drop numeric columns whose variance is below threshold.

Parameters:
  • df – input DataFrame.

  • threshold – variance cutoff.

  • verbose – print the dropped column names.

Returns:

filtered DataFrame.

spacr.utils.remove_noise(embedding, labels)[source]

Drop rows of embedding (and labels) whose label is DBSCAN noise (-1).

Rows are removed, not renumbered, so positional indices into the original data (an image_paths list, a DataFrame row order) no longer line up with the returned arrays – filter those alongside, using the same mask, or keep the identities before calling.

Parameters:
  • embedding – (N, D) ndarray of points. It is filtered by boolean mask, so a Python list or a DataFrame will not index correctly; pass a numpy array.

  • labels – length-N ndarray of cluster labels, aligned row-for-row with embedding. Only -1 is treated as noise, which is the DBSCAN convention – KMeans labels contain no -1 and pass through unchanged, making this a no-op rather than an error on KMeans output.

spacr.utils.remove_outliers_by_group(df, group_col, value_col, method='iqr', threshold=1.5)[source]

Removes outliers from value_col within each group defined by group_col.

Rows are selected, never modified: the original index is preserved and a new frame is returned. A row whose value is NaN fails the comparison and is always dropped, whichever method is used.

Parameters:
  • df (pd.DataFrame) – The input DataFrame.

  • group_col (str) – Column name to group by, or a list of column names. Grouping passes observed=False, so unused categories of a Categorical are kept. Rows whose group key is missing are discarded, because pandas drops NaN group keys and the per-row bound then comes back NaN.

  • value_col (str) – Column containing values to check for outliers. A name that is not in the frame raises KeyError.

  • method (str) – ‘iqr’ or ‘zscore’. Anything else raises ValueError. The two now agree on tiny groups: a one-row group has an undefined standard deviation, and since one row cannot be an outlier within its own group it is KEPT under both. It used to be dropped by ‘zscore’ and kept by ‘iqr’.

  • threshold (float) – Multiplier on the IQR (default 1.5), or the z-score cutoff. Must be >= 0; a negative value inverts the keep-band and is refused, because under ‘iqr’ it silently emptied every group with a nonzero IQR. Note 0 under ‘zscore’ still keeps only rows sitting exactly on the group mean, which is what a zero cutoff means. Under ‘zscore’ an outlier inflates its own group’s standard deviation, so the usual cutoffs keep far more than ‘iqr’ does on the same data – that is the statistic, not a defect.

Returns:

pd.DataFrame – A DataFrame with outliers removed.

spacr.utils.rename_columns_in_db(db_path)[source]

Rename legacy column spellings across every table in a SQLite database.

Applies DB_COLUMN_RENAMES — the plate-metadata names — and then DB_COLUMN_RENAME_PATTERNS — the two feature families that were spelled inconsistently — to every user table. A rename is skipped when the target name already exists in that table, which gives three properties worth relying on:

  • Idempotent. After a rename the legacy name is gone, so a second run finds nothing to do. Running it on every read is therefore free after the first.

  • Never destructive. A table that somehow carries both spellings — say time_id and timeID — keeps both, untouched. Neither column is dropped and nothing raises; the readers accept either spelling, so the data stays reachable and a human can decide which one is authoritative. Dropping or overwriting one of them here would destroy data to tidy a name, which is never the right trade.

  • All or nothing. SQLite’s DDL is transactional, but Python’s sqlite3 driver only opens an implicit transaction for DML (INSERT/UPDATE/DELETE/ REPLACE) — an ALTER TABLE runs in autocommit and lands immediately. So the previous version, which relied on a trailing con.commit(), left a database half-migrated when a later rename raised. The transaction is opened explicitly here and rolled back on any error, and the connection is closed in a finally.

A partial migration would not corrupt anything — each rename is independently valid and the next read finishes the job — but “the schema changed and then the call raised” is not a state a user should have to reason about.

Parameters:

db_path – Path to the SQLite database file to update in place.

Returns:

The list of (table, old, new) renames performed.

spacr.utils.reset_cellpose_model_reports()[source]

Forget which Cellpose model notices have already been printed.

Call this at the start of a run so a second run in the same process (a GUI session segmenting a second plate) reports its model choice again instead of inheriting the first run’s silence.

spacr.utils.reset_mp()[source]

Set the multiprocessing start method appropriate for the current OS.

Uses spawn on Windows and fork on Linux/macOS.

Returns:

None.

spacr.utils.resize_images_and_labels(images, labels, target_height, target_width, show_example=True)[source]

Resize aligned image/label lists to target_height x target_width.

Parameters:
  • images – iterable of source images (2-D or 3-D).

  • labels – matching iterable of label masks, or None.

  • target_height – output height in pixels.

  • target_width – output width in pixels.

  • show_example – display an example of the resized pair when True.

Returns:

(resized_images, resized_labels) lists.

spacr.utils.resize_labels_back(labels, orig_dims)[source]

Resize a list of label masks back to their original (width, height).

Parameters:
  • labels – iterable of label masks.

  • orig_dims – matching iterable of (width, height) tuples.

Returns:

list of resized label masks.

Raises:

ValueError – if lengths differ or orig_dims entries are malformed.

spacr.utils.save_file_lists(dst, data_set, ls)[source]

Write ls as a single-column CSV named <data_set>.csv under dst.

Parameters:
  • dst – destination directory.

  • data_set – column name and file stem.

  • ls – iterable of values to persist.

Returns:

None.

spacr.utils.save_settings(settings, name='settings', show=False)[source]

Persist a settings dict to <src>/settings/<name>.csv so a spacr run can be reproduced later.

Called by every pipeline entry point to snapshot the resolved settings before real work starts. The saved copy has test_mode and plot forced to False so that a downstream load_settings() -> re-run produces a full, headless run.

Parameters:
  • settings – settings dict; must contain src.

  • name – base filename (no extension); _list is appended when src is a list. Default 'settings'.

  • show – display the DataFrame before writing. Default False.

Returns:

None. Writes <src>/settings/<name>.csv.

Example

from spacr.utils import save_settings
save_settings(my_settings, name='my_experiment', show=True)

See also

load_settings() — inverse operation.

spacr.utils.search_reduction_and_clustering(numeric_data, n_neighbors, min_dist, metric, eps, min_samples, clustering, reduction_method, verbose, reduction_param=None, embedding=None, n_jobs=-1)[source]

Variant of reduction_and_clustering() accepting extra reducer kwargs via reduction_param.

Parameters:
  • numeric_data – numeric data matrix.

  • n_neighbors – UMAP n_neighbors or t-SNE perplexity (int or fraction).

  • min_dist – UMAP min_dist.

  • metric – distance metric.

  • eps – DBSCAN eps.

  • min_samples – DBSCAN min_samples or KMeans cluster count.

  • clustering – 'dbscan' or 'kmeans'.

  • reduction_method – 'umap' or 'tsne'.

  • verbose – print progress.

  • reduction_param – extra kwargs forwarded to the reducer.

  • embedding – precomputed embedding to skip fitting.

  • n_jobs – parallel worker count.

Returns:

(embedding, labels).

Raises:

ValueError – on unsupported reduction_method or clustering.

spacr.utils.setup_plot(figuresize, black_background, theme_colors=None)[source]

Create a square Matplotlib figure using scoped theme colors.

Parameters:
  • figuresize (float) – Figure width and height in inches.

  • black_background (bool) – Use the legacy dark or light fallback when theme_colors is not supplied.

  • theme_colors (mapping, optional) – background, foreground, and border colors. Missing or invalid entries use the fallback palette.

Returns:

tuple – The (figure, axes) pair.

Notes

Theme values are applied inside matplotlib.rc_context() and then to the created artists. Global Matplotlib settings are not modified.

spacr.utils.show_cam_on_image(img, mask)[source]

Return img overlaid with a jet colormap of mask as an 8-bit RGB image.

The sum of heatmap and image is renormalized by its own peak, so the output brightness is relative to the single hottest pixel – two images overlaid separately are not comparable to each other on absolute intensity.

Parameters:
  • img – 3-channel (H, W, 3) image already scaled to [0, 1]. It is added to the colormap rather than blended, so a [0, 255] image swamps the heatmap and the result is a near-uniform wash. A 2-D grayscale array fails to broadcast against the 3-channel heatmap. An image negative enough that the blend has no positive pixel left raises rather than returning a black frame – a black attribution map is indistinguishable from “the model looked nowhere”, which is a claim this function must never make on the strength of bad input.

  • mask – (H, W) activation map in [0, 1], matching img in height and width. Values outside that range are CLIPPED to it, with a RuntimeWarning, so an un-normalized CAM saturates at the hot end instead of wrapping the np.uint8 cast: before this was clipped, 1.1 landed at the cold end of jet, 1.5 in the middle and 2.0 back at the top, which could render the hottest region of a map as the coldest colour. An all-zero mask does not produce a black overlay: jet maps 0 to a non-zero color, so a zero mask over a zero image renormalizes to a saturated flat field.

Raises:

ValueError – img or mask contains NaN or infinity, or the blend has no positive pixel to normalize against.

spacr.utils.smooth_hull_lines(cluster_data)[source]

Return the x, y coordinates of a smoothed convex-hull outline of a 2-D point set.

Parameters:

cluster_data – 2-D array of point coordinates.

Returns:

tuple (x, y) of spline-interpolated hull coordinates (100 samples).

spacr.utils.split_my_dataset(dataset, split_ratio=0.1)[source]

Randomly split dataset into (train, val) subsets.

Parameters:
  • dataset – source dataset.

  • split_ratio – fraction of samples reserved for validation.

Returns:

(train_subset, val_subset).

spacr.utils.suggest_training_changes(dst, train_csv=None, val_csv=None, last_k=25, min_epochs=10, gap_threshold_acc=0.05, plateau_eps=0.001, noisy_var_ratio=0.03)[source]

Inspect saved training/validation progress CSVs and propose concrete training changes.

Parameters:
  • dst – folder where progress CSVs were saved.

  • train_csv – explicit train-CSV path; auto-detected in dst if None.

  • val_csv – explicit val-CSV path; auto-detected in dst if None.

  • last_k – number of recent epochs used for trend and plateau checks.

  • min_epochs – minimum epochs before most suggestions are issued.

  • gap_threshold_acc – accuracy generalization-gap threshold (train - val).

  • plateau_eps – absolute slope threshold used to declare a plateau.

  • noisy_var_ratio – instability flag threshold on stdev/mean of recent val loss.

Returns:

dict with summary (key scalars), flags (short codes), and suggestions (ordered suggestion strings).

Nested helpers

GradCAM.__call__.hook(module, input, output)

Forward hook: append the target layer’s output to features.

retain_grad() is required: PyTorch only populates .grad on leaf tensors, so without it features[0].grad is None below and GradCAM died with “‘NoneType’ object has no attribute ‘cpu’”.

spacr/utils.py:7084

GradCAMGenerator.hook_layers.backward_hook(module, grad_input, grad_output)

Backward hook: cache the gradient flowing into the target layer’s output.

spacr/utils.py:6782

GradCAMGenerator.hook_layers.forward_hook(module, input, output)

Forward hook: cache the target layer’s output activations.

spacr/utils.py:6778

TorchModel._run_backbone_raw.forward_fn(t)

Run the underlying backbone on t (used as the checkpoint target).

spacr/utils.py:3978

_checkpoint_module.contexts()

Return forward and recomputation contexts for non-reentrant checkpointing.

spacr/utils.py:191

_filter_objects._failed_in(low, high)

Labels failing an entry whose index is in [low, high).

spacr/utils.py:687

_generate_representative_images._compartment_column(compartment)

Return the selected compartment series, or raise for a missing column.

spacr/utils.py:2091

_outline_and_overlay.process_dim(mask_dim)

Return a dilated outline image of the labeled mask at image[..., mask_dim].

spacr/utils.py:1846

_pivot_counts_table._pivot_dataframe(df)

Pivot count-type rows into one column per object type, NaNs filled with 0.

spacr/utils.py:3372

_pivot_counts_table._read_table_to_dataframe(db_path, table_name='object_counts')

Return the given SQLite table as a DataFrame.

spacr/utils.py:3367

_save_settings_json.plain(value)

The value as something JSON can hold, or its repr.

spacr/utils.py:1635

annotate_conditions._get_type(val)

Determine if a value maps to ‘rowID’ or ‘columnID’.

spacr/utils.py:3483

annotate_conditions._map_or_default(column_name, values, loc, df)

Assign or map values into column_name based on optional row/column loc.

spacr/utils.py:3491

annotate_predictions.assign_condition(row)

Return the condition label ('screen'/'pc'/'nc' or '') for a metadata row.

spacr/utils.py:5246

build_loss._asl(logits, y, gpos, gneg, clip)

Return mean asymmetric multilabel loss for logits and targets.

spacr/utils.py:5059

build_loss._auto_choice() → str

Return the default loss name from class count and imbalance.

spacr/utils.py:5071

build_loss._focal_bce(logits, y, alpha, gamma)

Return mean focal binary cross-entropy for logits and y.

spacr/utils.py:5032

build_loss._focal_ce(logits, y_idx, alpha, gamma)

Return mean focal cross-entropy for logits and class indices.

spacr/utils.py:5042

build_loss._infer_indices(target: torch.Tensor, C: int) → torch.Tensor

Return class indices from an index vector or 2-D target matrix.

spacr/utils.py:5015

build_loss.loss_fn(logits, target)

Closure: compute the selected per-batch loss from (logits, target).

spacr/utils.py:5087 spacr/utils.py:5092 spacr/utils.py:5101 spacr/utils.py:5107 spacr/utils.py:5115 spacr/utils.py:5125 spacr/utils.py:5133 spacr/utils.py:5139

calculate_loss._focal_bce_with_logits(logits, y, alpha=1.0, gamma=2.0, reduction='mean')

Return focal binary cross-entropy for logits and targets y.

spacr/utils.py:4516

calculate_loss._focal_cross_entropy(logits, y_idx, alpha=1.0, gamma=2.0, reduction='mean')

Return focal cross-entropy for logits and class indices y_idx.

spacr/utils.py:4528

class_visualization.blur_image(img, sigma=1)

In-place Gaussian blur of each channel of img with standard deviation sigma.

spacr/utils.py:6974

class_visualization.deprocess(img_tensor)

Undo ImageNet normalization and return an (H, W, 3) numpy image in [0, 1].

spacr/utils.py:6981

class_visualization.jitter(img, ox, oy)

Return img shifted (rolled) by ox and oy pixels along the spatial axes.

spacr/utils.py:6970

debug.decorator(func)

Inner decorator that binds the logger for func and returns the wrapper.

spacr/utils.py:1033

debug.decorator.wrapper(*args, **kwargs)

Temporarily bump the logger to DEBUG while func runs, then restore its level.

spacr/utils.py:1038

feature_folder_name._one(member)

Return one filesystem-safe selection-member slug.

spacr/utils.py:9546

group_feature_class.find_feature_class(feature, compartments)

Return the group label(s) matched in feature — joined with ‘-’ when more than one hits.

spacr/utils.py:10128

load_settings.parse_value(value)

Parse the string value into the appropriate Python data type.

spacr/utils.py:1396

merge_regression_res_with_metadata.extract_and_clean_gene(feature)

Return the gene ID parsed from a feature string like C(gene)[T.<id>_...], or None.

spacr/utils.py:9458

pick_best_model.sort_key(x)

Return (role, accuracy, epoch) from metadata or legacy name.

spacr/utils.py:4585

plot_grid.cluster_color(cluster_label)

Resolve one cluster’s color once for panels and legend.

spacr/utils.py:7987

pretty_print_settings._fmt(v)

Return v as text truncated to the table’s value width.

spacr/utils.py:1532

pretty_print_settings._row(k, v)

Return one padded key-and-formatted-value table row.

spacr/utils.py:1537

pretty_print_settings._say(line)

Print a console-safe form of line and return None.

spacr/utils.py:1528

process_masks.cluster_objects(properties, n_clusters=2)

Return a fitted KMeans object clustering the property dicts into n_clusters groups.

spacr/utils.py:9386

process_masks.measure_morphology_and_intensity(mask, image)

Return a list of dicts with area/mean_intensity/perimeter/eccentricity per labeled region.

spacr/utils.py:9380

process_masks.plot_clusters(properties, labels)

Show a 2-D PCA scatter of the property vectors colored by cluster label.

spacr/utils.py:9400

process_masks.read_files_in_batches(folder, batch_size=50)

Yield sorted lists of .npy filenames from folder in chunks of batch_size.

spacr/utils.py:9373

process_masks.remove_objects_not_in_largest_cluster(mask, labels, largest_cluster_label)

Return mask with all labeled regions removed except those in largest_cluster_label.

spacr/utils.py:9392

suggest_training_changes._find_csv(root, hint)

Return the lexically last matching CSV in root, if any.

spacr/utils.py:4748

suggest_training_changes._last_seq(series, k)

Return at most the final k values as a floating-point array.

spacr/utils.py:4789

suggest_training_changes._normalize_cols(df)

Return df with normalized, aliased, first-occurrence columns.

spacr/utils.py:4753

suggest_training_changes._poly_slope(y)

Return the finite linear slope of y, or zero when undefined.

spacr/utils.py:4778

suggest_training_changes._scalar(val)

Ensure a single float even if a Series sneaks through.

The Series branch is currently unreachable: every call site passes <Series>.iloc[<int>], which yields a numpy scalar. It is kept as a deliberate guard because the label-based .loc lookups this function used to rely on returned a Series whenever the progress CSV had a duplicated index, and that is an easy regression to reintroduce.

spacr/utils.py:4735