spacr.utils¶
Shared image, model, database, statistics, and pipeline utilities.
Exceptions¶
The table being appended to holds an import's copy of the same field. |
|
A measurement frame's units differ from the ones already in the table. |
|
An installed optional dependency is too old for spaCR's API contract. |
Classes¶
LRU cache with a fixed maximum size. |
|
Small classifier stacking |
|
1x1 convolution that fuses input channels down to 64 feature maps. |
|
Focal loss for binary, multiclass, and multilabel targets. |
|
Named-hook Grad-CAM implementation for arbitrary target layers. |
|
Grad-CAM (and variants) map generator for binary classifiers. |
|
Compute integrated-gradients attributions for a classifier. |
|
Dilated conv block followed by a 1x1 attention convolution. |
|
ResNet backbone with a two-layer spaCR binary-classification head. |
|
Generate saliency maps and predictions for a binary classifier. |
|
Standard scaled dot-product attention layer. |
|
Callable transform that zeroes out image channels not present in |
|
Linear-projected self-attention layer. |
|
Spatial attention gate that reweights features by pooled channel statistics. |
|
Thin wrapper around TorchVision classification backbones that: |
|
TorchVision backbone with a spaCR linear head (streamlined variant of |
Functions¶
|
Fit a multiple-linear regression on gene:grna interactions plus plate/row/column terms. |
|
Merge per-image correlation stats with parsed well IDs and insert into the dataset DB. |
|
Insert activation-map PNG paths and parsed well IDs into the dataset DB. |
|
Adds a new column to the database table by matching on a common column from the DataFrame. |
|
Add |
|
Run |
|
Return |
|
Annotate |
|
Read prediction CSV and add plate/well/field/object columns plus a |
|
Zero out (or set to |
|
Return colors and their positional mapping for the unique labels. |
|
Augment negative and positive class images and split them into train/test folders. |
|
Expand |
|
Return a list of PIL images covering 4 rotations x 2 horizontal reflections of |
|
Run |
|
Save six augmentations of one image (original, 90/180/270 rotations, H/V flips). |
|
Return the boundary F1 score between two masks with tolerance |
|
Return a closure |
|
Compute per-image Pearson and Manders correlations between input and activation channels. |
|
Return the intersection-over-union of two binary masks after zero-padding to a common shape. |
|
Auto-select and return a loss for binary, multiclass, or multilabel problems. |
|
Calculate the shortest edge-to-edge distance between two objects (e.g., pathogen and nucleus). |
Rename legacy column spellings on an in-memory measurement frame. |
|
|
Validate that every index label in |
|
Return |
Checks multicollinearity of the predictors by computing the VIF. |
|
|
Helper function to check if a feature is normally distributed. |
|
Return |
|
Instantiate a classification model by name for binary or multiclass problems. |
|
Synthesize an input image that maximizes the classifier score for |
|
Return a one-row DataFrame of accuracy, PR-AUC, and optimal-threshold stats. |
|
Delete the intermediate mask-pipeline folders once |
Close file descriptors from 3 up to the soft NOFILE limit. |
|
Terminate all detected multiprocessing child processes and close file descriptors. |
|
|
Perform Random Forest feature importance, ANOVA for normally distributed features, |
|
Combine the results into a single DataFrame. |
|
Return the area under the precision-recall curve swept over |
|
Return |
|
Return the IRM penalty as the sum of squared gradient dot-products across environments. |
|
Return the COCO-style segmentation AP by matching connected components across IoU thresholds. |
|
Return |
|
Return the codec text printed to |
|
Return |
|
Return filenames in |
|
Converts all int64 npy masks in a folder to uint16 with relabeling to ensure all labels are retained. |
|
Copies images from their original locations to a 'consolidated' folder, |
|
Convert cell masks under |
|
Normalize a metadata DataFrame to the canonical spaCR names and plate ids. |
Renamed legacy metadata columns. Defined in |
|
|
Rewrite PNG paths (in a DataFrame or list) so they live under |
|
Return the number of reads in a gzipped FASTQ file. |
|
Return a boolean circular mask of shape |
|
Decorator that temporarily sets the given logger to DEBUG for the wrapped call. |
|
Recursively delete |
|
Remove intermediate per-channel and stack folders under |
|
Map each RAW channel index to its position on the merged stack's axis. |
|
Return the Dice similarity of two masks, treating any nonzero value as foreground. |
|
Do nothing: IPython is unavailable, so there is nowhere to display to. |
|
Downloads all model files from Hugging Face and stores them in the |
|
Return per-class sample counts as a |
|
Return the boundary of a binary mask via morphological dilation minus erosion. |
|
Extract features from images using a pre-trained ResNet model. |
|
Extracts all .tar.bz2 files in the given folder into subfolders with the same name as the tar file. |
|
Which of |
|
A folder name for one feature selection. Safe on every filesystem. |
|
Insert cropped PNG filepaths and parsed well/object IDs into the measurements DB. |
|
Fill the holes inside each object of a label mask, keeping every id. |
|
Reads a CSV into a DataFrame, keeps the rows whose column value falls OUTSIDE |
|
Return |
|
Restrict a features DataFrame to a channel of interest and clean up correlated/low-variance columns. |
|
Return a nearby |
|
Fisher's exact test per mutant column against a binarized phenotype label. |
|
Takes a file path and reformats it to be compatible with the current operating system. |
|
Return a deterministic Viridis RGBA palette for cluster points. |
|
Generates a cytoplasm mask from nucleus and cell masks. |
|
Return a wells-by-genes fraction matrix, dropping columns below |
|
Recursively scans a folder and its subfolders for images, then creates a mapping of: |
|
Return all |
Return the installed CUDA toolkit version as a digit-only string, or |
|
|
Return the standard |
|
Return glob matches for |
|
Return the standard set of ML output paths for the given model and channel selection. |
|
Return rows of |
|
Return the standard |
|
Return all dotted submodule names of |
|
Add a column tagging each feature with its compartment (or other group) label. |
|
Initialize shared multiprocessing |
|
Return the intensity-inverted image, reflected through the dtype range. |
|
Return |
|
Return |
|
Return the Jaccard/IoU index of two binary masks. |
|
Fit Lasso or Ridge on one-hot-encoded gene/grna/plate/row/column predictors. |
|
Load and preprocess an image. |
|
Load the |
|
Reload a spacr settings CSV (written by |
|
Map a column-ID value to one of |
|
Return the number of nonzero labeled objects in |
|
Greedy match each predicted mask to a still-unmatched true mask above |
|
Copy a random subset of source files into a |
|
Merge |
|
Merge regression outputs with gene metadata on the parsed |
|
Merge by perimeter and filter labeled objects across a directory of masks. |
|
Merge touching labeled objects whose shared boundary exceeds |
|
Print RMSE/MAE/Durbin-Watson and show residual/QQ/scale-location diagnostic plots. |
|
Normalize text representations of an unfiltered feature selection. |
|
Ensures that the 'src' value is properly formatted as either a list of strings or a single string. |
|
Percentile-normalize each channel of an image stack into the target dtype range. |
|
Migrate |
|
Zero-pad |
|
Perform ANOVA or Kruskal-Wallis tests depending on normality of features. |
|
Return the strongest checkpoint anywhere below |
|
Draw cluster outlines, points, and centroid labels onto |
|
Plot a grid of example images per cluster label discovered in |
|
Plot a 2-D embedding with cluster outlines, points, and optional image overlays. |
|
Render one column per cluster of representative images with colored borders and labels. |
|
Place a zoomed thumbnail of |
|
Overlay up to |
|
Overlay sample images from |
Cast a batch to |
|
|
Prepare a feature matrix by filtering, decorrelating, log-transforming, and scaling |
|
Load and preprocess |
|
Print a settings dict to the console as a tidy, aligned table. |
|
Print a one-line progress report with an ETA derived from mean step time. |
|
Load one triple of parasite/cell/nuclei masks, merge cells in place, and return the elapsed time. |
|
Cluster object morphology/intensity across a mask folder and keep the largest cluster in place. |
|
Split image paths into well identifiers and binarize the |
|
Rank features by how well they predict the cluster label. |
|
Return |
|
Reduce |
|
Return |
|
Drop numeric columns whose absolute correlation with a prior column exceeds |
|
Drop labeled objects whose mean intensity is on the wrong side of |
|
Drop numeric columns whose variance is below |
|
Drop rows of |
|
Removes outliers from |
|
Rename legacy column spellings across every table in a SQLite database. |
Forget which Cellpose model notices have already been printed. |
|
|
Set the multiprocessing start method appropriate for the current OS. |
|
Resize aligned image/label lists to |
|
Resize a list of label masks back to their original |
|
Write |
|
Persist a settings dict to |
|
Variant of |
|
Create a square Matplotlib figure using scoped theme colors. |
|
Return |
|
Return the x, y coordinates of a smoothed convex-hull outline of a 2-D point set. |
|
Randomly split |
|
Inspect saved training/validation progress CSVs and propose concrete training changes. |
Module Contents¶
- exception spacr.utils.ImportedCopyNotReleased[source]¶
Bases:
ValueErrorThe table being appended to holds an import’s copy of the same field.
foreign.run_importcopies the imported frame into the canonicalcell/nucleus/pathogentable when the destination is empty, so a project built purely by import is readable by every spaCR tool. That copy is a convenience and it stops being one the moment spaCR measures the same field:_merge_and_save_to_database()appends, so its rows land beside theirs, in different columns, with nothing in the row marking the seam – and everycount_celldownstream becomes the sum of two populations.A resume supersedes the copy before measuring (see
spacr.resume.supersede_imported_copies()) or refuses and says so, which covers every path through a resume. This covers the path around one: a directmeasure_cropwithresumeoff.Raised only when the copy cannot be handed back provably losslessly – no
foreign_<object>to check the rows against, a row in the copy with no twin in it, a timelapse whose frames the importer never keyed, or a delete that did not act on the rows the count cleared. In every one of those cases the field’s rows are not written, because a refused write can be re-run and a mixed table cannot be un-mixed.Initialize self. See help(type(self)) for accurate signature.
- exception spacr.utils.MeasurementUnitsMismatch[source]¶
Bases:
ValueErrorA measurement frame’s units differ from the ones already in the table.
A 2-D field measures areas in px^2; a 3-D field measures volumes, in voxels or um^3, and writes them into the same
<object>_areacolumn, because that column is read by name by every downstream selector, model and threshold ever written against a spaCR database and renaming it would break all of them silently. Appending both into one table would therefore leave a numeric column that mixes two incompatible quantities with nothing in the row to tell them apart, which no amount of downstream care could recover from. So it is refused here instead.Initialize self. See help(type(self)) for accurate signature.
- exception spacr.utils.OptionalDependencyCompatibilityError[source]¶
Bases:
ImportErrorAn installed optional dependency is too old for spaCR’s API contract.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.utils.Cache(max_size)[source]¶
LRU cache with a fixed maximum size.
- Parameters:
max_size – maximum number of entries retained; oldest is evicted on overflow.
Store the size limit and initialize an empty
OrderedDict.
- class spacr.utils.CustomCellClassifier(num_classes, pathogen_channel, use_attention, use_checkpoint, dropout_rate)[source]¶
Bases:
torch.nn.ModuleSmall classifier stacking
EarlyFusionand a multi-scale attention block.- Parameters:
num_classes – output class count.
pathogen_channel – reserved for downstream use; kept for API compatibility.
use_attention – reserved for downstream use; kept for API compatibility.
use_checkpoint – run the forward pass through
torch.utils.checkpoint.dropout_rate – reserved for downstream use; kept for API compatibility.
Build the fusion, multi-scale, and linear classifier submodules.
- class spacr.utils.EarlyFusion(in_channels)[source]¶
Bases:
torch.nn.Module1x1 convolution that fuses input channels down to 64 feature maps.
- Parameters:
in_channels – number of input channels.
Create the 1x1 fusion convolution.
- class spacr.utils.FocalLossWithLogits(alpha=1.0, gamma=2.0, reduction='mean')[source]¶
Bases:
torch.nn.ModuleFocal loss for binary, multiclass, and multilabel targets.
Auto-selects the BCE or cross-entropy branch based on the shapes of
logitsandtarget:binary: logits
(N,)or(N,1); target float(N,)in{0,1}.multiclass: logits
(N,C); target long(N,)in[0..C-1].multilabel: logits
(N,C); target float(N,C)in{0,1}.
- Parameters:
alpha – class-balancing factor (float or 1-D tensor of shape
(C,)).gamma – focusing parameter.
reduction – one of
'mean','sum','none'.
Store the focal-loss hyperparameters.
- class spacr.utils.GradCAM(model, target_layers=None, use_cuda=True)[source]¶
Named-hook Grad-CAM implementation for arbitrary target layers.
- Parameters:
model – trained model to inspect.
target_layers – list of dotted layer names to hook.
use_cuda – if true, move the model and inputs to CUDA unconditionally; the caller must ensure CUDA is available.
Store the model and move it to CUDA if requested.
- class spacr.utils.GradCAMGenerator(model, target_layer, cam_type='gradcam')[source]¶
Grad-CAM (and variants) map generator for binary classifiers.
- Parameters:
model – trained model to inspect.
target_layer – dotted attribute path to the convolutional layer to probe.
cam_type – variant identifier, e.g.
'gradcam'.
Store the model, resolve the target layer, and register activation/gradient hooks.
- compute_gradcam_and_predictions(X)[source]¶
Return Grad-CAM maps and predictions for every sample.
- Parameters:
X – differentiable input image batch to classify and probe.
- compute_gradcam_maps(X, y)[source]¶
Return a normalized Grad-CAM map for one sample.
- Parameters:
X – single-sample differentiable input batch to probe.
y – binary label selecting the signed output score.
- get_layer(model, target_layer)[source]¶
Resolve a dotted attribute path into the referenced submodule.
- Parameters:
model – root model from which attribute traversal starts.
target_layer – dot-separated submodule attribute path.
- percentile_normalize(img, lower_percentile=2, upper_percentile=98)[source]¶
Per-channel percentile-normalize
imginto[0, 1].- Parameters:
img – channels-last image array to normalize.
- plot_activation_grid(X, gradcam, predictions, overlay=True, normalize=False)[source]¶
Render a grid overlaying Grad-CAM maps on inputs with predicted-class labels.
The grid is always eight columns wide with
ceil(N / 8)rows, and unused panels in an incomplete last row are hidden. The figure is returned, never shown.- Parameters:
X – batch tensor shaped
(N, C, H, W);Nfixes the grid size, and an empty batch raisesValueErrorfromsubplots. The pixels are read only underoverlay, where the sample is permuted to(H, W, C)–Cof 1, 3 or 4 renders,Cof 2 raisesTypeErrorfromimshow.gradcam – torch tensor with at least
Nentries. Each map may be(H, W),(1, H, W)or(3, H, W), with the same shape normalization used by the saliency twin. It is indexed and moved to the CPU on every iteration, so it must be a tensor either way.predictions – sequence supporting
predictions[i].item(), whose scalar is stamped in each panel’s corner. A plain Python list of ints raisesAttributeError.overlay – true draws the input beneath a translucent map; false draws the map alone. Default
True.normalize – percentile-stretch the input image only; the Grad-CAM map is always drawn raw. Has no effect unless
overlayis true, and a channel that is flat between its 2nd and 98th percentiles divides by zero and comes outNaN. DefaultFalse.
- Returns:
the Matplotlib
Figure.
- class spacr.utils.IntegratedGradients(model)[source]¶
Compute integrated-gradients attributions for a classifier.
- Parameters:
model – trained PyTorch model.
Store the model and switch it to eval mode.
- generate_integrated_gradients(input_tensor, target_label_idx, baseline=None, num_steps=50)[source]¶
Return integrated gradients from
baselinetoinput_tensorfortarget_label_idx.- Parameters:
input_tensor – input sample tensor.
target_label_idx – target class index whose logit is attributed.
baseline – reference tensor (defaults to zeros of the same shape).
num_steps – number of Riemann-sum interpolation steps.
- Returns:
attribution ndarray with the shape of
input_tensor.
- class spacr.utils.MultiScaleBlockWithAttention(in_channels, out_channels)[source]¶
Bases:
torch.nn.ModuleDilated conv block followed by a 1x1 attention convolution.
- Parameters:
in_channels – input channel count.
out_channels – output channel count.
Build the dilated convolution and 1x1 spatial-attention convolution.
- custom_forward(x)[source]¶
Apply dilated conv + ReLU followed by the 1x1 spatial attention.
- Parameters:
x – input feature map for the convolutional block.
- forward(x)[source]¶
Forward pass; delegates to
custom_forward().- Parameters:
x – input feature map for the convolutional block.
- class spacr.utils.ResNet(resnet_type='resnet50', dropout_rate=None, use_checkpoint=False, init_weights='imagenet')[source]¶
Bases:
torch.nn.ModuleResNet backbone with a two-layer spaCR binary-classification head.
- Parameters:
resnet_type – one of
'resnet18'/'resnet34'/'resnet50'/'resnet101'/'resnet152'.dropout_rate – dropout probability before the final linear layer;
Nonedisables.use_checkpoint – enable gradient checkpointing through the ResNet backbone.
init_weights –
'imagenet'for pretrained weights or'none'for random init.
- Raises:
ValueError – if
resnet_typeis unsupported, or ifinit_weightsis neither'imagenet'nor'none'.
Select the backbone and delegate head construction to
initialize_base().- forward(x)[source]¶
Return the flattened single-logit prediction for input batch
x.- Parameters:
x – image batch for the configured ResNet backbone.
- initialize_base(base_model_dict, dropout_rate, use_checkpoint, init_weights)[source]¶
Build the backbone (with or without pretrained weights) and the two-layer head.
- Parameters:
base_model_dict – dict with keys
func(model constructor) andweights.dropout_rate – dropout probability applied between the two linear layers.
use_checkpoint – enable gradient checkpointing through the backbone.
init_weights –
'imagenet'or'none'.
- Raises:
ValueError – if
init_weightsis neither'imagenet'nor'none'.
- class spacr.utils.SaliencyMapGenerator(model)[source]¶
Generate saliency maps and predictions for a binary classifier.
- Parameters:
model – trained PyTorch model with a single-logit binary output.
Store the model to be probed.
- compute_saliency_and_predictions(X)[source]¶
Return saliency maps and the model’s own predicted classes.
- Parameters:
X – differentiable input image batch to classify and probe.
- compute_saliency_maps(X, y)[source]¶
Return absolute-gradient saliency maps for inputs
X.- Parameters:
X – differentiable input image batch to probe.
y – binary labels selecting the signed output scores.
- percentile_normalize(img, lower_percentile=2, upper_percentile=98)[source]¶
Per-channel percentile-normalize
imginto[0, 1].- Parameters:
img – channels-last image array to normalize.
- plot_activation_grid(X, saliency, predictions, overlay=True, normalize=False)[source]¶
Render a grid overlaying saliency maps on inputs with predicted-class labels.
The grid is always eight columns wide with
ceil(N / 8)rows, and unused panels in an incomplete last row are hidden. The figure is returned, never shown.- Parameters:
X – batch tensor shaped
(N, C, H, W);Nfixes the grid size, and an empty batch raisesValueErrorfromsubplots. The pixels are read only underoverlay, where the sample is permuted to(H, W, C)–Cof 1, 3 or 4 renders,Cof 2 reachesimshowas an invalid shape and raisesTypeError.saliency – torch tensor of at least
Nentries. It is indexed and moved to the CPU on every iteration even whenoverlayis false, so a numpy array raisesAttributeErroreither way. Each entry may be(H, W),(1, H, W)or(3, H, W); a leading singleton is removed and a leading RGB dimension is transposed to channels-last. Other three-dimensional shapes reachimshowand raiseTypeError.predictions – sequence supporting
predictions[i].item(), whose scalar is stamped in each panel’s corner. A plain Python list of ints raisesAttributeError.overlay – true draws the input beneath a translucent map; false draws the map alone. Default
True.normalize – percentile-stretch the input image only; the saliency map is always drawn raw. Has no effect unless
overlayis true, and a channel that is flat between its 2nd and 98th percentiles divides by zero and comes outNaN. DefaultFalse.
- Returns:
the Matplotlib
Figure.
- class spacr.utils.ScaledDotProductAttention(d_k)[source]¶
Bases:
torch.nn.ModuleStandard scaled dot-product attention layer.
- Parameters:
d_k – dimensionality of key/query vectors used in the scaling factor.
Store
d_kused to scale attention logits.
- class spacr.utils.SelectChannels(channels)[source]¶
Callable transform that zeroes out image channels not present in
channels.- Parameters:
channels – iterable of 1-based channel indices to keep (1=red, 2=green, 3=blue).
Store the list of channels to preserve.
- class spacr.utils.SelfAttention(in_channels, d_k)[source]¶
Bases:
torch.nn.ModuleLinear-projected self-attention layer.
- Parameters:
in_channels – input feature dimension.
d_k – projected key/query/value dimension.
Build the Q/K/V projections and the underlying attention layer.
- class spacr.utils.SpatialAttention(kernel_size=7)[source]¶
Bases:
torch.nn.ModuleSpatial attention gate that reweights features by pooled channel statistics.
- Parameters:
kernel_size – convolution kernel width used to fuse average+max pooled maps.
Build the fusion convolution and sigmoid gate.
- class spacr.utils.TorchModel(model_name: str = 'resnet50', pretrained: bool = True, dropout_rate: float | None = None, use_checkpoint: bool = False, num_classes: int = 2, multilabel: bool = False, image_size: int = 224)[source]¶
Bases:
torch.nn.Module- Thin wrapper around TorchVision classification backbones that:
Loads a requested backbone with (optional) pretrained weights
Strips its classification head to expose features
Adds a simple Linear ‘spacr’ classifier with
num_classesoutputsOptionally applies dropout before the final classifier
Supports gradient checkpointing
Works with most TorchVision classification models. Non-classification (detection/segmentation) models are rejected with a clear error.
Build the backbone, strip its head, and attach the spaCR linear classifier.
- Parameters:
model_name – TorchVision classification model to load.
pretrained – use ImageNet-pretrained weights when available.
dropout_rate – dropout probability applied to backbone and spaCR head;
Nonedisables.use_checkpoint – enable gradient checkpointing through the backbone.
num_classes – output class count;
1yields a BCE-style binary head.multilabel – informational flag consumed by external loss/metrics code.
image_size – square input resolution used for the dummy forward pass that infers the backbone’s feature width.
- Raises:
ValueError – if
model_nameis not a TorchVision model.
- forward(x: torch.Tensor) torch.Tensor[source]¶
Return classification logits of shape
(N, num_classes).- Parameters:
x – input image batch for the configured TorchVision backbone.
- class spacr.utils.TorchModel_v2(model_name: str = 'resnet50', pretrained: bool = True, dropout_rate: float = None, use_checkpoint: bool = False, num_classes: int = 2, multilabel: bool = False)[source]¶
Bases:
torch.nn.ModuleTorchVision backbone with a spaCR linear head (streamlined variant of
TorchModel).- Parameters:
model_name – TorchVision classification model to load.
pretrained – use ImageNet-pretrained weights when available.
dropout_rate – dropout probability applied to backbone and spaCR head;
Nonedisables.use_checkpoint – enable gradient checkpointing through the backbone.
num_classes – output class count.
multilabel – informational flag consumed by external loss/metrics code.
- Raises:
ValueError – if
model_nameis not a TorchVision model.
Build the backbone, strip its head, and attach the spaCR classifier.
- forward(x: torch.Tensor) torch.Tensor[source]¶
Return classification logits of shape
(N, num_classes).- Parameters:
x – input image batch for the configured TorchVision backbone.
- spacr.utils.MLR(merged_df, refine_model)[source]¶
Fit a multiple-linear regression on gene:grna interactions plus plate/row/column terms.
- Parameters:
merged_df – DataFrame with
gene,grna,plate,row,column,predcolumns.refine_model – refit after removing outliers by residuals and Cook’s distance.
- Returns:
tuple
(max_effects, max_effects_pvalues, model, df).
- spacr.utils.activation_correlations_to_database(df, img_paths, source_folder, settings)[source]¶
Merge per-image correlation stats with parsed well IDs and insert into the dataset DB.
- Parameters:
df – DataFrame of correlation stats indexed by
file_name.img_paths – iterable of PNG paths matching rows of
df.source_folder – experiment root; DB written to
measurements/<dataset>.db.settings – settings dict; must contain
datasetandcam_type.
- Returns:
None.
- spacr.utils.activation_maps_to_database(img_paths, source_folder, settings)[source]¶
Insert activation-map PNG paths and parsed well IDs into the dataset DB.
- Parameters:
img_paths – iterable of PNG paths for activation-map images.
source_folder – experiment root; DB written to
measurements/<dataset>.db.settings – settings dict; must contain
datasetandcam_type.
- Returns:
None.
- spacr.utils.add_column_to_database(settings)[source]¶
Adds a new column to the database table by matching on a common column from the DataFrame. If the column already exists in the database, it adds the column with a suffix. NaN values will remain as NULL in the database.
- Parameters:
settings (dict) – A dictionary containing the following keys: csv_path (str): Path to the CSV file with the data to be added. db_path (str): Path to the SQLite database (or connection string for other databases). table_name (str): The name of the table in the database. update_column (str): The name of the new column in the DataFrame to add to the database. match_column (str): The common column used to match rows.
- Returns:
None
- spacr.utils.add_images_to_tar(paths_chunk, tar_path, total_images)[source]¶
Add
paths_chunkimages totar_path, updating the shared counter for progress.- Parameters:
paths_chunk – list of image paths to add.
tar_path – destination tar archive path.
total_images – overall image count used to render progress.
- Returns:
None.
- spacr.utils.adjust_cell_masks(parasite_folder, cell_folder, nuclei_folder, organelle_folder=None, overlap_threshold=5, perimeter_threshold=30, n_jobs=None, *, output_folder=None)[source]¶
Run
process_mask_file_adjust_cell()in parallel across matching mask files.- Parameters:
parasite_folder – folder of parasite masks.
cell_folder – folder of cell masks (overwritten in place).
nuclei_folder – folder of nuclei masks.
organelle_folder – optional folder of organelle masks.
overlap_threshold – fractional overlap threshold used by the merger.
perimeter_threshold – shared-perimeter threshold used by the merger.
n_jobs – worker count;
Nonedefaults tocpu_count() - 2and values below two run inline without starting a child process.output_folder – optional separate folder for adjusted masks. None preserves the historical in-place behavior. A separate folder keeps all source masks byte-identical, and every selected field is rebuilt from its source on each invocation, including after interrupted work. In place, each adjusted mask’s SHA-256 is recorded in
ADJUSTED_CELLS_LEDGERas it lands, and a mask whose bytes still match its record is left alone on the next run: adjusting an adjusted mask merges it again, so a re-run used to change the result every time. A cell mask segmented again no longer matches and is adjusted afresh.
- Returns:
None.
- Raises:
ValueError – if the three folders contain different numbers of files or mismatched filenames, or a mask is truncated, nonnumeric, empty, not two-dimensional or has different dimensions from its partners. Available organelle masks are checked too. Header-only validation of every field finishes before any mask is changed or workers are started. An explicit output folder must differ from every source folder and must not contain masks outside the selected field set.
- spacr.utils.all_elements_match(list1, list2)[source]¶
Return
Trueif every element oflist1is contained inlist2.- Parameters:
list1 – iterable of items to test.
list2 – iterable acting as the reference set.
- Returns:
Truewhenlist1is a subset oflist2, elseFalse.
- spacr.utils.annotate_conditions(df, cells=None, cell_loc=None, pathogens=None, pathogen_loc=None, treatments=None, treatment_loc=None)[source]¶
Annotate
dfwith host cell, pathogen, treatment, and combinedconditioncolumns.- Parameters:
df – DataFrame to annotate; must contain
rowID/columnID.cells – host cell types (str or list).
cell_loc – per-cell-type list-of-lists of row/column identifiers.
pathogens – pathogens (str or list).
pathogen_loc – per-pathogen list-of-lists of row/column identifiers.
treatments – treatments (str or list).
treatment_loc – per-treatment list-of-lists of row/column identifiers.
- Returns:
annotated DataFrame with
host_cells,pathogen,treatment,conditioncolumns.
- spacr.utils.annotate_predictions(csv_loc)[source]¶
Read prediction CSV and add plate/well/field/object columns plus a
condlabel.- Parameters:
csv_loc – path to a predictions CSV with a
pathcolumn of PNG paths.- Returns:
DataFrame enriched with parsed metadata and a
condcolumn ('screen'/'pc'/'nc'from the plate/well convention).
- spacr.utils.apply_mask(image, output_value=0)[source]¶
Zero out (or set to
output_value) pixels outside a circular mask fit toimage.The circle is not a parameter:
create_circular_mask()is called with no center or radius, so it is always the largest circle inscribed in the frame. On a non-square image that means the short side sets the radius.- Parameters:
image – 2-D
(H, W)or 3-D(H, W, C)array. The mask is built from the first two axes and broadcast across every channel, so all channels are cropped identically.output_value – fill written outside the circle. It goes through
np.where, so a value the input dtype cannot hold promotes the whole result – passingnp.nanto auint16image returns a float array, not a masked integer one. Default0.
- spacr.utils.assign_colors(unique_labels, random_colors)[source]¶
Return colors and their positional mapping for the unique labels.
- Parameters:
unique_labels – cluster labels in the order assigned palette indices.
random_colors – iterable of color values converted to tuples.
- spacr.utils.augment_classes(dst, nc, pc, generate=True, move=True, group_by='well', test_size=0.1)[source]¶
Augment negative and positive class images and split them into train/test folders.
- Parameters:
dst – destination root; augmented images land under
aug_nc/aug_pcand move intoaug/{train,test}/{nc,pc}.nc – negative-class source image paths.
pc – positive-class source image paths.
generate – run augmentation before moving files.
move – split augmented images into train/test folders.
group_by – source identity held intact across train/test. Default
'well';'cell'permits sibling objects on both sides.test_size – requested test fraction; whole groups make it approximate.
- Returns:
None.
- spacr.utils.augment_dataset(dataset, is_grayscale=False)[source]¶
Expand
datasetby 8x through rotation and horizontal reflection of every image tensor.- Parameters:
dataset – iterable of
(tensor, label, filename).is_grayscale – informational flag (retained for API compatibility).
- Returns:
list of augmented
(tensor, label, filename)tuples.- Raises:
TypeError – if an image is not a
torch.Tensor.
- spacr.utils.augment_image(image)[source]¶
Return a list of PIL images covering 4 rotations x 2 horizontal reflections of
image.The 8 outputs are the dihedral group of the square and include the unmodified original as element 0, so the list is an 8x expansion, not 8 extra images. Ordering is rotation-major –
[0deg, 0deg flipped, 90deg, 90deg flipped, ...]– which matters if you are keeping a parallel list of labels.- Parameters:
image – a PIL image or a numpy array. Arrays are used as-is; PIL images are converted first. A 2-D grayscale input is expanded to 3 channels via
cv2.cvtColor, so every result is RGB even when the input was not – channel count is not preserved. The rotations go throughcv2, which expectsuint8or another OpenCV-supported dtype; the finalImage.fromarraylikewise rejects the float or 16-bit arrays typical of raw microscopy, so convert to 8-bit before calling. Because 90-degree rotations swap height and width, a non-square input yields images of two different shapes in the same list.
- spacr.utils.augment_images(file_paths, dst)[source]¶
Run
augment_single_image()in parallel overfile_paths.Spawn workers so they inherit no locks held by other threads in this process. Close and join the pool after mapping; terminating forked workers can hang while their exit handlers wait on inherited locks.
- Parameters:
file_paths – iterable of source image paths.
dst – destination folder (created if missing).
- Returns:
None.
- spacr.utils.augment_single_image(args)[source]¶
Save six augmentations of one image (original, 90/180/270 rotations, H/V flips).
- Parameters:
args –
(img_path, dst)tuple.- Returns:
None.
- spacr.utils.boundary_f1_score(mask_true, mask_pred, dilation_radius=1)[source]¶
Return the boundary F1 score between two masks with tolerance
dilation_radius.Both masks are binarized before the boundary is taken, so this scores the outline of the foreground as a whole: boundaries where two labeled objects abut are interior to that foreground and do not appear. Split/merge errors between touching cells are therefore invisible to this metric.
- Parameters:
mask_true – reference label or binary mask, reduced to its boundary by
extract_boundaries().mask_pred – predicted mask of the same shape; the two boundary images are intersected element-wise, so the masks must be pixel-registered.
dilation_radius – half-width of the square structuring element, giving a band
2 * dilation_radius + 1pixels wide. This is the matching tolerance – raising it forgives localization error but also thickens both boundaries, so scores rise for every model and stop being comparable across different radii. Default1.
- spacr.utils.build_loss(loss_type: str = 'ce', num_classes: int = 2, class_counts: torch.Tensor | None = None, label_smoothing: float = 0.0, focal_gamma: float = 2.0, focal_alpha: float | None = None, logit_adjust_tau: float = 0.0, asl_gamma_pos: float = 0.0, asl_gamma_neg: float = 4.0, asl_clip: float = 0.05)[source]¶
Return a closure
loss_fn(logits, target)implementing the requested loss.Supported
loss_typevalues:'ce','ce_smooth','ce_weighted','focal_ce','bce','focal_bce','logit_adjust_ce','asl','auto'.num_classes==1selects binary (BCE variants);>=2selects multiclass (CE variants).- Parameters:
loss_type – loss identifier (see above).
num_classes – output class count.
class_counts – per-class sample counts used to derive weights or logit adjustment.
label_smoothing – label-smoothing epsilon for
ce_smooth.focal_gamma – focal-loss focusing parameter.
focal_alpha – focal-loss class-balancing factor (float or per-class tensor).
logit_adjust_tau – strength of the Menon-et-al. logit adjustment; 0 disables.
asl_gamma_pos – asymmetric-loss gamma for positives.
asl_gamma_neg – asymmetric-loss gamma for negatives.
asl_clip – asymmetric-loss negative-probability clip.
- Returns:
loss_fn(logits, target)callable returning a scalar tensor.- Raises:
ValueError – if
loss_typeis unknown or incompatible withnum_classes.
- spacr.utils.calculate_activation_correlations(inputs, activation_maps, file_names, manders_thresholds=None)[source]¶
Compute per-image Pearson and Manders correlations between input and activation channels.
- Parameters:
inputs – input image batch, tensor of shape
(B, C, H, W).activation_maps – activation-map batch, tensor of shape
(B, C, H, W)or(B, H, W).file_names – file names corresponding to each image in the batch.
manders_thresholds – intensity percentiles used for Manders coefficients. Default
[15, 50, 75].
- Returns:
DataFrame with one row per image and one column per channel-pair statistic.
- spacr.utils.calculate_iou(mask1, mask2)[source]¶
Return the intersection-over-union of two binary masks after zero-padding to a common shape.
Unlike
jaccard_index(), this returns0rather thannanwhen both masks are empty, which is what makes it safe to call inside the matching loop.- Parameters:
mask1 – 2-D array. Any nonzero value counts as foreground, so a multi-label crop is treated as one merged object – pass a single object’s mask if you want a per-object IoU.
mask2 – 2-D array compared against
mask1. Shapes may differ;pad_to_same_shape()zero-pads both at the bottom and right, which assumes the two masks share a top-left origin. Two crops taken from different offsets in the same image will score meaninglessly low.
- spacr.utils.calculate_loss(output, target, prefer_focal=False, gamma=2.0, alpha=1.0, reduction='mean')[source]¶
Auto-select and return a loss for binary, multiclass, or multilabel problems.
- Dispatches based on the shapes/dtypes of
outputandtarget: binary: logits
(N,1), float targets in{0,1}-> BCE / focal-BCE.multiclass: logits
(N,C), long targets(N,)-> CE / focal-CE.multilabel: logits
(N,C), float targets(N,C)-> BCE / focal-BCE.
- Parameters:
output – model logits.
target – ground-truth labels.
prefer_focal – use the focal-loss variant instead of plain CE/BCE.
gamma – focal-loss focusing parameter.
alpha – focal-loss class-balancing factor.
reduction – one of
'mean','sum','none'.
- Returns:
scalar loss tensor (or per-sample tensor when
reduction='none').
- Dispatches based on the shapes/dtypes of
- spacr.utils.calculate_shortest_distance(df, object1, object2)[source]¶
Calculate the shortest edge-to-edge distance between two objects (e.g., pathogen and nucleus).
Parameters: - df: Pandas DataFrame containing measurements - object1: String, name of the first object (e.g., “pathogen”) - object2: String, name of the second object (e.g., “nucleus”)
Returns: - df: Pandas DataFrame with a new column for shortest edge-to-edge distance.
- Parameters:
df – measurement frame containing centroid and Feret-diameter columns for both objects.
object1 – prefix of the first object’s measurement columns.
object2 – prefix of the second object’s measurement columns.
- spacr.utils.canonicalize_measurement_columns(df)[source]¶
Rename legacy column spellings on an in-memory measurement frame.
The DataFrame counterpart of
rename_columns_in_db(), for frames that did not come from a spaCR database and so never passed through it — a CSV exported by an older release, or a frame a user assembled themselves.Follows the same never-destructive rule: a rename whose target is already present is skipped, so a frame carrying both spellings keeps both rather than losing one to a silently dropped duplicate. The rule itself lives in
spacr.schema.canonical_rename_plan(), which this andschema.canonicalise_columnsboth call so the two frame canonicalisers cannot drift apart again — and which folds case, because these frames are written withto_sqland SQLite compares identifiers case-insensitively.- Parameters:
df – A measurement DataFrame.
- Returns:
dfwith legacy column names replaced (a copy is not made; the frame is renamed in place and returned).
- spacr.utils.check_index(df, elements=5, split_char='_')[source]¶
Validate that every index label in
dfsplits intoelementsparts onsplit_char.- Parameters:
df – DataFrame whose index labels are compound identifiers.
elements – Expected number of parts after splitting. Default
5.split_char – Delimiter used to split each index label. Default
'_'.
- Returns:
None.
- Raises:
ValueError – if any index label does not split into
elementsparts.
- spacr.utils.check_mask_folder(src, mask_fldr, resume=False)[source]¶
Return
Trueif masks insrc/masks/mask_fldrstill need generating.- Parameters:
src – experiment root containing
masks/andstack/subfolders.mask_fldr – subfolder name under
masks/.resume – accepted for the callers that pass it. Only structurally complete mask arrays are counted whether or not it is set, so an empty or truncated array left by an interrupted run is re-queued.
- Returns:
Truewhen the mask folder is missing or any expected stack filename lacks a complete mask. Unrelated masks cannot substitute for a missing field and do not force complete fields to run again.
- spacr.utils.check_multicollinearity(x)[source]¶
Checks multicollinearity of the predictors by computing the VIF.
- Parameters:
x – DataFrame of the design matrix – one row per observation, one column per predictor, and no response column. Every column is fed to
variance_inflation_factorviax.values, so all columns must be numeric; categorical predictors have to be one-hot encoded first. Add an explicit constant column if you want the intercept accounted for, since without one the VIFs are inflated by the shared mean. Perfectly collinear columns yieldinf.
- spacr.utils.check_normality(series)[source]¶
Helper function to check if a feature is normally distributed.
This is a failure to reject at alpha 0.05, not evidence of normality: the answer is
Truewhenever the D’Agostino-Pearson test does not find significant skew or kurtosis. Small samples therefore look normal for want of power, and very large ones fail on deviations too small to matter – which is what decides whetherperform_statistical_tests()sends a feature to ANOVA or to Kruskal-Wallis.- Parameters:
series – one numeric column of observations, pooled across all groups.
NaNpropagates and makes the p-valueNaN, which comparesFalseagainst alpha and so is reported as normal – drop missing values first. Under 8 observationsscipycannot run the skew test, returnsNaN, and the feature is likewise reported as normal; a constant column behaves the same way. Values are treated as one sample, so a strongly bimodal feature whose groups are each normal is judged on the mixture.
- spacr.utils.check_overlap(current_position, other_positions, threshold)[source]¶
Return
Trueifcurrent_positionis withinthresholdof any point inother_positions.- Parameters:
current_position – candidate point as a sequence of coordinates. Any dimensionality works as long as it matches the entries of
other_positions; a genuine length mismatch raisesValueErrorfrom the subtraction, while a length-1 entry broadcasts silently and yields a meaningless distance.other_positions – already-placed points to test against. Scanned linearly with an early return, so cost grows with the number of placed items – this is the inner loop of the image-scatter layout. An empty sequence returns
False, so the first placement always succeeds.threshold – minimum center-to-center Euclidean separation, in the same units as the coordinates (data units for an embedding, not pixels or points). The comparison is strict
<, so a distance exactly equal tothresholdcounts as not overlapping. Because it measures centers, set it to roughly the thumbnail width; half of that still lets images overlap visually.
- spacr.utils.choose_model(model_type: str, device: torch.device, init_weights: bool = True, dropout_rate: float = 0.0, use_checkpoint: bool = False, channels: int = 3, height: int = 224, width: int = 224, chan_dict: dict[str, Any] | None = None, num_classes: int = 2, verbose: bool = False) torch.nn.Module | None[source]¶
Instantiate a classification model by name for binary or multiclass problems.
- Parameters:
model_type – TorchVision model name (e.g.
'resnet50','vit_b_16').'custom'passes the name check but then raisesNotImplementedError, as no custom builder is wired up.device – unused; the model is built on the CPU and the caller moves it.
init_weights – load pretrained weights when available.
dropout_rate – dropout probability before the classifier head (
None/0disables).use_checkpoint – enable gradient checkpointing for the backbone.
channels – unused; the forward sanity check always feeds 3 channels.
height – square input resolution, forwarded as
TorchModel(image_size=...); it therefore fixes the dummy-forward size used to infer the backbone feature dimension (and so the size of the classifier head) as well as both dimensions of the square forward sanity check (falls back to224when falsy).width – unused;
heightsets both dimensions.chan_dict – unused; reserved for the unimplemented custom branch.
num_classes – output class count;
1yields a single-logit BCE head.verbose – print the model structure when
True.
- Returns:
The instantiated
nn.Module.- Raises:
ValueError –
model_typenames no backbone, or the built model does not produce logits of the requested shape.
Unsupported names raise immediately and include close TorchVision matches when available, so configuration errors are reported before training.
- spacr.utils.class_visualization(target_y, model_path, dtype, img_size=224, channels=None, l2_reg=0.001, learning_rate=25, num_iterations=100, blur_every=10, max_jitter=16, show_every=25, class_names=None)[source]¶
Synthesize an input image that maximizes the classifier score for
target_y.- Parameters:
target_y – target class index.
model_path – path to the trained model checkpoint.
dtype – tensor dtype; overridden internally based on CUDA availability.
img_size – square image size (pixels).
channels – input channels (defaults to
[0, 1, 2]).l2_reg – L2 regularization weight on the pixel norm.
learning_rate – gradient-ascent step size.
num_iterations – optimization iteration count.
blur_every – interval (iterations) between periodic Gaussian blurs.
max_jitter – maximum pixel jitter applied per iteration.
show_every – interval (iterations) between preview plots.
class_names – display names for the classes (defaults to
['nc', 'pc']).
- Returns:
deprocessed image as a numpy array.
- spacr.utils.classification_metrics(all_labels, prediction_pos_probs, loss, epoch)[source]¶
Return a one-row DataFrame of accuracy, PR-AUC, and optimal-threshold stats.
- Parameters:
all_labels – ground-truth binary labels.
prediction_pos_probs – predicted positive-class probabilities.
loss – loss tensor for the epoch (
.item()is called).epoch – epoch number used as the row index.
- Returns:
DataFrame indexed by epoch with accuracy, per-class accuracy, loss, PR-AUC, and optimal threshold columns.
- Raises:
ValueError – if
all_labelsandprediction_pos_probshave different lengths.
- spacr.utils.cleanup_pipeline_folders(src, keep_intermediate=False, keep_original=False, verbose=True)[source]¶
Delete the intermediate mask-pipeline folders once
merged/is built.By default spaCR keeps only
merged/(the concatenated image+mask arrays that Measure reads). This removesstack/+masks/(their data is embedded inmerged/and object labels are recorded in the database) and the raworig/backup, unless the caller opts to keep them.Heavily guarded so it never destroys un-merged data:
stack/+masks/are only removed whenmerged/is non-empty AND everystack/*.npyhas a matchingmerged/*.npy(i.e. every field of view was merged).- Parameters:
src – run root folder (holds
merged/,stack/,masks/,orig/).keep_intermediate – keep
stack/+masks/when True.keep_original – keep the raw
orig/backup when True.
- Returns:
list of folder paths that were deleted.
- spacr.utils.close_file_descriptors()[source]¶
Close file descriptors from 3 up to the soft NOFILE limit.
- spacr.utils.close_multiprocessing_processes()[source]¶
Terminate all detected multiprocessing child processes and close file descriptors.
- spacr.utils.cluster_feature_analysis(all_df, cluster_col='cluster')[source]¶
Perform Random Forest feature importance, ANOVA for normally distributed features, and Kruskal-Wallis for non-normally distributed features. Combine results into a single DataFrame.
- Parameters:
all_df – DataFrame holding the numeric feature columns and
cluster_col, one row per object. The same frame is passed to the Random Forest and to the statistical tests, so the two rankings describe the same rows. The features are selected byschema.model_feature_columns(..., allow_unknown=True), which means stray numeric bookkeeping columns (row/column indices, object IDs) are picked up as features unless you drop them first. Because each feature is routed to either ANOVA or Kruskal-Wallis, the merged output has exactly one of the two p-value pairs filled per row andNaNin the other.cluster_col – column holding the group label, excluded from the features and used as the Random Forest target and the grouping variable for the tests. It needs at least two distinct labels, and every group needs enough rows for the test to run. DBSCAN’s
-1noise label is not special-cased here, so it is analysed as if it were a real cluster – strip it withremove_noise()first if you do not want that. Default'cluster'.
- spacr.utils.combine_results(rf_df, anova_df, kruskal_df)[source]¶
Combine the results into a single DataFrame.
All three frames are keyed on
Featureand carry exactly one row per feature:rf_dfis built from the feature list, andperform_statistical_tests()sends each feature to either ANOVA or Kruskal-Wallis, never both. Henceone_to_one. A repeatedFeature– the signature of a frame with duplicated column names, or of two runs’ results concatenated by mistake – would multiply the importance rows and report the same feature several times as if independently ranked.- Parameters:
rf_df – random-forest results keyed uniquely by
Feature.anova_df – ANOVA results keyed uniquely by
Feature.kruskal_df – Kruskal-Wallis results keyed uniquely by
Feature.
- spacr.utils.compute_ap_over_iou_thresholds(true_masks, pred_masks, iou_thresholds)[source]¶
Return the area under the precision-recall curve swept over
iou_thresholds.- Parameters:
true_masks – sequence of per-object ground-truth masks, one entry per object – not a single label image. Its length is the ground-truth count used for recall, so filtering objects out changes the denominator.
pred_masks – sequence of per-object predicted masks, matched greedily against
true_masksat each threshold. Matching walks predictions in the order given and claims the first free true mask that clears the threshold, so the ordering can change which pairs form.iou_thresholds – iterable of IoU cutoffs to sweep. The curve is the trapezoid over the resulting points sorted by recall, so a single threshold gives an area of
0– pass at least two (COCO convention isnp.linspace(0.5, 0.95, 10)). Duplicate thresholds contribute zero-width segments and do not count.
- Raises:
ValueError – if a computed precision or recall falls outside
[0, 1], which indicates the mask counts disagree with the matches.
- spacr.utils.compute_average_precision(matches, num_true_masks, num_pred_masks)[source]¶
Return
(precision, recall)given match count, true count, and predicted count.Despite the name this computes a single precision/recall point, not an averaged precision;
compute_ap_over_iou_thresholds()is what integrates those points into an AP.- Parameters:
matches – the pair list from
match_masks(). Only its length is used, so any sized container works, but it must be the matched pairs rather than all candidate pairs – matching is greedy and one-to-one, so the length is the true-positive count.num_true_masks – total ground-truth objects; drives false negatives as
num_true_masks - len(matches). Passing a count smaller than the match count silently yields a recall above 1, whichcompute_ap_over_iou_thresholds()rejects withValueError.num_pred_masks – total predicted objects, used the same way for false positives. Both counts are the totals for the whole field, not per class. Zero denominators return
0instead of raising.
- spacr.utils.compute_irm_penalty(losses, dummy_w, device)[source]¶
Return the IRM penalty as the sum of squared gradient dot-products across environments.
- Parameters:
losses – per-environment loss tensors.
dummy_w – scalar dummy weight used for gradient computation.
device – torch device on which to compute the penalty.
- Returns:
scalar IRM penalty value.
- spacr.utils.compute_segmentation_ap(true_masks, pred_masks, iou_thresholds=np.linspace(0.5, 0.95, 10))[source]¶
Return the COCO-style segmentation AP by matching connected components across IoU thresholds.
This is the whole-image entry point: unlike
compute_ap_over_iou_thresholds()it takes label images and splits them into objects itself.- Parameters:
true_masks – ground-truth label or binary image for one field. It is re-run through
label(), so existing IDs are discarded and touching objects that share a border merge into one component – the AP is computed on connected components, not on the IDs you supply.pred_masks – predicted mask for the same field, treated identically. Each object is reduced to its bounding-box crop by
regionprops, so objects are compared shape-to-shape with their positions dropped; two identically shaped cells in different corners score as a perfect match.iou_thresholds – IoU cutoffs to sweep. Default
np.linspace(0.5, 0.95, 10)is the COCO sweep. This default array is evaluated once at import and shared by every call, so do not mutate it in place.
- spacr.utils.console_can_encode(text, stream=None)[source]¶
Return
Truewhentextcan be printed tostreamas-is.- Parameters:
text – the string about to be printed.
stream – text stream to test against; defaults to
sys.stdout.
- Returns:
bool.
- spacr.utils.console_encoding(stream=None)[source]¶
Return the codec text printed to
streamhas to survive.- Parameters:
stream – a text stream; defaults to
sys.stdout.- Returns:
a codec name,
'utf-8'when the stream does not declare one (a queue-backed GUI console, a StringIO, a captured pipe).
- spacr.utils.console_safe(text, stream=None)[source]¶
Return
textwith anything the console cannot encode replaced by?.Console decoration must never be able to end a run. No Windows codepage encodes spaCR’s own output set –
▸(U+25B8) is absent from cp1252, cp437, cp850, cp932 and cp936, and the box-drawing frame is absent from cp1252 – and neither does any of them encode the domain vocabulary that ends up in settings values, such as the parental strainΔku80or aµmvoxel size. Printing either to a non-UTF-8 stream raisesUnicodeEncodeError, and on Windows that is the normal case the moment stdout is redirected: a batch-queue job,spacr-run, a legacy console.- Parameters:
text – the string about to be printed.
stream – text stream to encode against; defaults to
sys.stdout.
- Returns:
textunchanged when it is printable, otherwise a lossy but printable version of it.
- spacr.utils.control_filelist(folder, mode='columnID', values=None)[source]¶
Return filenames in
folderwhose row or column ID matches one ofvalues.The filename is split on
_and the second token is inspected: characters after the first (mode='columnID') or the leading character (mode='rowID') are matched againstvalues.- Parameters:
folder – Directory to scan.
mode –
'columnID'matches trailing digits,'rowID'matches leading letter. Default'columnID'.values – Iterable of allowed ID strings. Defaults to
['01', '02'].
- Returns:
List of matching filenames.
- spacr.utils.convert_and_relabel_masks(folder_path)[source]¶
Converts all int64 npy masks in a folder to uint16 with relabeling to ensure all labels are retained.
Parameters: - folder_path (str): The path to the folder containing int64 npy mask files.
Returns: - None
- Parameters:
folder_path – directory containing
.npymasks to inspect and convert in place.
- spacr.utils.copy_images_to_consolidated(image_path_map, root_folder)[source]¶
Copies images from their original locations to a ‘consolidated’ folder, renaming them according to the generated dictionary.
- spacr.utils.correct_masks(src)[source]¶
Convert cell masks under
src/masks/cell_mask_stackto uint16 and re-stack arrays.Relabels masks so they fit in
uint16and then re-concatenates the four array folders undersrcin the layout expected downstream.- Parameters:
src – Root folder of a spacr run containing a
masks/subfolder.- Returns:
None.
- spacr.utils.correct_metadata(df)[source]¶
Normalize a metadata DataFrame to the canonical spaCR names and plate ids.
One call into
spacr.schema.canonicalise_frame(), which is the whole vocabulary in one place:every legacy spelling renamed, folding case and punctuation, so
Plate/PLATE/plate/plateid/plate_nameall arrive asplateIDand the same for the row, the column, the field and the well;one column per key. A file carrying
wellandwellIDhas two opinions about which well a row came from; they are compared row by row as stripped strings (so1,1.0and' 1 'agree and a dtype difference is not a disagreement), one is kept, and the rest are dropped. Agreement prints; disagreement warns and says how many rows differ.the
ppplate repair, over every column that embeds the plate id.
THE
ppREPAIR RUNS AFTER THE RENAMES, and over every plate-bearing column. It used to run first and only overplateID/prcfo, which meant it did nothing at all for the files that actually carry the artifact: a legacy score CSV has aplatecolumn and NOplateIDcolumn, so the guard was false, the repair was skipped, and the very next line then copied the unrepairedplatevalue intoplateID.The cost was silent and total. Score files stamped
pplate1met count files stampedplate1, everyprcdiffered by one character, the join produced ZERO rows, and the run died several steps later inside a plot withKeyError: 0– nowhere near the mismatch, and with nothing on screen naming a plate.- Parameters:
df – Metadata DataFrame that may still use legacy naming.
- Returns:
The DataFrame with canonical columns.
- spacr.utils.correct_metadata_column_names(df)[source]¶
Renamed legacy metadata columns. Defined in
spacr.schema.Re-exported here because every existing caller imports it from
utils, and moved there because importing this module costs torch, torchvision and cv2 – 6.7 seconds – for a function that needs none of them.- Parameters:
df – tabular frame whose legacy metadata names are canonicalized.
- spacr.utils.correct_paths(df, base_path, folder='data')[source]¶
Rewrite PNG paths (in a DataFrame or list) so they live under
base_path/folder.A non-string entry is passed through untouched.
png_listis LEFT-joined onto the object tables, so any object whose crop was never written arrives here withpng_path= NaN – a statespacr.io._read_and_join_tables()documents as healthy (len(merged) == len(cell) > len(png_list):save_pngoff for a field, a crop that failed to write, an interrupted run, or acell_idthat could not be migrated). Testingbase_path not in pathon that NaN raisedTypeError: argument of type 'float' is not iterableand took the whole embedding down over one missing thumbnail. There is no path to re-anchor for such a row, and it has to keep its position so the rewritten column still aligns withdf.Delegate rewriting to
spacr.crops.reanchor_path(), which handles same-platform moves, Windows paths read on Linux, and old absolute paths. It finds the rightmost anchor component after normalizing separators and checks existing roots component by component. Paths with no matchingfoldercomponent pass through unchanged; their count and one example are printed so unresolved paths remain visible.- Parameters:
df – DataFrame with a
png_pathcolumn, or a list of paths.base_path – destination root to prepend.
folder – intermediate folder name that anchors the rewrite.
- Returns:
DataFrame + list, or list, mirroring the input type.
- spacr.utils.count_reads_in_fastq(fastq_file)[source]¶
Return the number of reads in a gzipped FASTQ file.
Counts total lines and divides by four (the FASTQ record length).
- Parameters:
fastq_file – Path to a
.fastq.gzfile.- Returns:
Integer read count.
- spacr.utils.create_circular_mask(h, w, center=None, radius=None)[source]¶
Return a boolean circular mask of shape
(h, w)centered oncenter.- Parameters:
h – image height.
w – image width.
center –
(x, y)center; defaults to the image middle.radius – circle radius; defaults to the largest circle fitting inside.
- Returns:
boolean ndarray where
Truemarks pixels withinradius.
- spacr.utils.debug(enabled=True, logger_name=None)[source]¶
Decorator that temporarily sets the given logger to DEBUG for the wrapped call.
- Parameters:
enabled – no-op when
False.logger_name – logger name to tweak; defaults to the function’s module logger.
- Returns:
decorator function.
- spacr.utils.delete_folder(folder_path)[source]¶
Recursively delete
folder_pathif it exists (files and subdirectories included).- Parameters:
folder_path – directory to remove, contents and all. A missing path or a plain file is reported on stdout and ignored – the function never raises for those, so it cannot be used to confirm that a delete happened; check with
os.path.isdirafterwards if that matters. Deletion is unconditional and unprompted, with no trash or dry-run, so a wrong path is not recoverable. Symlinked subdirectories are not descended into but are still handed toos.rmdir, which raises on a symlink, so a tree containing one aborts part-way through.
- spacr.utils.delete_intermedeate_files(settings)[source]¶
Remove intermediate per-channel and stack folders under
settings['src'].Safeguarded to only run when a
merged/folder is present and theorig/backup folder exists, so raw inputs are preserved.- Parameters:
settings – Dict with an
'src'key naming the run’s root folder.- Returns:
None.
- spacr.utils.dense_mask_channel_positions(settings)[source]¶
Map each RAW channel index to its position on the merged stack’s axis.
Built the same way
io.preprocess_img_databuilds the stack, because that is the only thing that makes the answer true: walk the roles inMASK_CHANNEL_ROLE_ORDERand give each newly seen raw channel the next dense position.THE TRAP THIS EXISTS TO CLOSE. Several callers computed the position as
sorted({nucleus, cell, pathogen, organelle})instead, which agrees with role order only when the roles happen to be in ascending channel order. Withnucleus_channel=2, cell_channel=0, organelle_channel=1the stack is[2, 0, 1]– raw channel 1 sits at position 2 – while the sorted reading says position 1, which holds the CELL image. Cellpose then segments organelles on the cell plane, silently, and every count and intensity downstream is measured from the wrong masks.- Parameters:
settings – the run settings, holding the raw
*_channelkeys.- Returns:
{raw_channel: dense_position}. Channels that are None or uncoercible are absent, matching the writer’s own behaviour.
- spacr.utils.dice_coefficient(mask1, mask2)[source]¶
Return the Dice similarity of two masks, treating any nonzero value as foreground.
- Parameters:
mask1 – array binarized with
> 0, so negative values are counted as background – a signed difference image will not behave as expected.mask2 – array of the same shape as
mask1; likejaccard_index()there is no padding step, so shapes must already agree. Two empty masks return1.0here (defined as perfect agreement) rather thannan.
- spacr.utils.display(*args, **kwargs)[source]¶
Do nothing: IPython is unavailable, so there is nowhere to display to.
THE FALLBACK IS THE POINT.
IPython.display.displayis imported at module scope, and IPython can be mid-init – partially imported by another thread – which makes that import raise. Letting it propagate would make importing this module fail for a reason that has nothing to do with what the module does. spaCR only callsdisplayfrom notebook contexts; the Qt GUI ignores it.- Parameters:
args – whatever the caller would have displayed.
kwargs – likewise.
- spacr.utils.download_models(repo_id='einarolafsson/models', retries=5, delay=5)[source]¶
Downloads all model files from Hugging Face and stores them in the
resources/modelsdirectory within the installedspacrpackage.
- spacr.utils.estimate_class_counts(loader, num_classes: int, src=None, classes=None) torch.Tensor[source]¶
Return per-class sample counts as a
LongTensorof lengthnum_classes.When
srcandclassesare provided the counts are taken from the file listings undersrc/<class>, avoiding a slow DataLoader iteration on NAS.- Parameters:
loader – fallback DataLoader iterated only when folder info is missing.
num_classes – number of output classes.
src – parent folder containing per-class subfolders.
classes – ordered class-folder names matching
src.
- Returns:
LongTensorof per-class counts.
- spacr.utils.extract_boundaries(mask, dilation_radius=1)[source]¶
Return the boundary of a binary mask via morphological dilation minus erosion.
- Parameters:
mask – label or binary mask.
dilation_radius – half-width of the structuring element.
- Returns:
boolean boundary mask.
- spacr.utils.extract_features(image_paths, resnet=resnet50)[source]¶
Extract features from images using a pre-trained ResNet model.
The classification head is stripped, so each image yields its pooled backbone features.
- Parameters:
image_paths – Iterable of image paths, each loaded via
load_image().resnet – TorchVision model constructor (not an instance), called as
resnet(pretrained=True). Defaultresnet50.
- Returns:
(N, 2048)ndarray of pooled features, one row per image (2048 for ResNet-50; the width follows the chosen backbone).
- spacr.utils.extract_tar_bz2_files(folder_path)[source]¶
Extracts all .tar.bz2 files in the given folder into subfolders with the same name as the tar file.
- Parameters:
folder_path (str) – Path to the folder containing .tar.bz2 files.
- spacr.utils.feature_columns(columns, selection)[source]¶
Which of
columnsthe selection keeps. Order preserved.The union over the selection’s members, so
[1, 'morphology']is channel 1’s intensities AND the shapes rather than the empty intersection of the two.COLOCALISATION BELONGS TO BOTH CHANNELS IT MEASURES. A
cell_channel_1_channel_2_pearsonscolumn names two channels and survives a request for either – which is what makes “localization” reachable without a separate setting: ask for the channel and its relationships come with it.- Parameters:
columns – ordered column names available for selection.
selection – channel, morphology group, text filter, mixture, or
Noneas accepted byfeature_selection().
- spacr.utils.feature_folder_name(channel_of_interest) str[source]¶
A folder name for one feature selection. Safe on every filesystem.
Noneisall_features; a channel ischannel_1; several arechannels_1_2;morphologyis itself; a mixture joins them in the order given; and a free-text filter is slugified, because a user may reasonably filter onmean_intensityand a column fragment can carry anything.- Parameters:
channel_of_interest – feature selection accepted by
feature_selection().
- spacr.utils.filepaths_to_database(img_paths, settings, source_folder, crop_mode)[source]¶
Insert cropped PNG filepaths and parsed well/object IDs into the measurements DB.
- Parameters:
img_paths – iterable of PNG paths for cropped objects.
settings – settings dict;
timelapsetoggles time_id parsing.source_folder – experiment root; DB is written to
measurements/measurements.db.crop_mode – a registered object role, including any organelle slot.
- Returns:
None.
- spacr.utils.fill_holes_in_mask(mask)[source]¶
Fill the holes inside each object of a label mask, keeping every id.
Delegates to
spacr.qt.mask_engine.fill_label_holes(), the one hole filler for label images. This used to runndimage.labelover the mask first, which made every pair of touching objects one object: with Cellposefill_inon (its default in Apply), a field of 74 adjacent cells was saved as 8. Now no object is merged or renumbered, and a hole takes the id of the object that encloses it.- Parameters:
mask (np.ndarray) – A labeled mask where each object has a unique integer value. A boolean mask is labelled by connectivity first.
- Returns:
np.ndarray – The mask with holes filled and the original labels preserved, in the input’s dtype.
- spacr.utils.filter_and_save_csv(input_csv, output_csv, column_name, upper_threshold, lower_threshold)[source]¶
Reads a CSV into a DataFrame, keeps the rows whose column value falls OUTSIDE the two thresholds, and saves the filtered DataFrame to a new CSV file.
The two tests are combined with OR, not AND, so this is a two-tailed selection that keeps the extremes and discards the middle. Both comparisons are strict, so a value exactly equal to either threshold is dropped, and so is
NaN. Passing anupper_thresholdbelowlower_thresholdmakes the two conditions cover the whole line and nothing is filtered out at all.- Parameters:
input_csv (str) – Path to the input CSV file, read with
pd.read_csv.output_csv (str) – Path to save the filtered CSV file, written without the index. Its parent directory must already exist – pandas raises
OSErrorrather than creating it.column_name (str) – Column the two comparisons are applied to. A name that is not in the frame raises
KeyError, and a text column raisesTypeErrorwhen compared against a numeric threshold.upper_threshold (float) – Rows strictly greater than this are retained.
lower_threshold (float) – Rows strictly less than this are retained too; everything between the two bounds is discarded.
- Returns:
None. The filtered frame is written to
output_csv, shown withdisplayfor notebook users, and the destination is printed.
- spacr.utils.filter_columns(df, filter_by)[source]¶
Return
dfrestricted to columns matchingfilter_by(or morphology columns).- Parameters:
df – source DataFrame.
filter_by – substring required in column names, or
'morphology'to drop channel columns.
- Returns:
column-filtered DataFrame.
- spacr.utils.filter_dataframe_features(df, channel_of_interest, exclude=None, remove_low_variance_features=True, remove_highly_correlated_features=True, verbose=False)[source]¶
Restrict a features DataFrame to a channel of interest and clean up correlated/low-variance columns.
- Parameters:
df – input DataFrame.
channel_of_interest – int, str, list, or
'morphology'to select feature groups.exclude – feature(s) to drop from the final list. A single name or any number of them; the Qt ‘Exclude’ field collects a list.
remove_low_variance_features – apply
remove_low_variance_columns().remove_highly_correlated_features – apply
remove_highly_correlated_columns().verbose – print filter details.
- Returns:
(filtered_df, features).
- spacr.utils.find_non_overlapping_position(x, y, image_positions, threshold, max_attempts=100)[source]¶
Return a nearby
(x, y)jittered position that does not collide withimage_positions.- Parameters:
x – original x.
y – original y.
image_positions – previously placed points.
threshold – minimum allowed spacing.
max_attempts – retry budget before giving up.
- Returns:
(x, y)tuple; original position if no non-overlapping spot is found.
- spacr.utils.fishers_odds(df, threshold=0.5, phenotyp_col='mean_pred')[source]¶
Fisher’s exact test per mutant column against a binarized phenotype label.
- Parameters:
df – DataFrame with per-mutant presence columns plus
phenotyp_col.threshold – cutoff below which
phenotyp_colis called “high phenotype”.phenotyp_col – name of the phenotype column.
- Returns:
DataFrame with columns
Mutant,OddsRatio,PValue,AdjustedPValue.
- spacr.utils.format_path_for_system(path)[source]¶
Takes a file path and reformats it to be compatible with the current operating system.
- Parameters:
path (str) – The file path to be formatted.
- Returns:
str – The formatted path for the current operating system.
- spacr.utils.generate_colors(num_clusters, black_background)[source]¶
Return a deterministic Viridis RGBA palette for cluster points.
- Parameters:
num_clusters – how many colors to sample, evenly spaced across the Viridis range
0.08-0.92(the extremes are trimmed so the darkest cluster stays visible on black and the lightest on white). Coerced withint()and floored at1, so0or a negative still yields a one-entry palette rather than an empty one. Count the clusters you will actually plot: DBSCAN’s-1noise label is not drawn, so including it shifts every real cluster’s color.black_background – accepted for call compatibility with the plotting helpers but not used – the palette is Viridis either way. Set the background through
setup_plot/theme_colorsinstead; changing this flag will not change the colors you get back.
- spacr.utils.generate_cytoplasm_mask(nucleus_mask, cell_mask)[source]¶
Generates a cytoplasm mask from nucleus and cell masks.
Parameters: - nucleus_mask (np.array): Binary or segmented mask of the nucleus (non-zero values represent nucleus). - cell_mask (np.array): Binary or segmented mask of the whole cell (non-zero values represent cell).
Returns: - cytoplasm_mask (np.array): Copy of cell_mask with nucleus pixels set to 0, keeping the cell labels elsewhere (pathogens are not considered).
- Parameters:
nucleus_mask – nucleus mask whose nonzero pixels are excluded.
cell_mask – labeled cell mask copied into the cytoplasm result.
- spacr.utils.generate_fraction_map(df, gene_column, min_frequency=0.0)[source]¶
Return a wells-by-genes fraction matrix, dropping columns below
min_frequency.- Parameters:
df – long-format DataFrame with
prc,count,well_read_sumcolumns.gene_column – column identifying the gene/guide.
min_frequency – drop columns whose maximum fraction is below this cutoff.
- Returns:
DataFrame indexed by
prcwith per-gene fractions.
- spacr.utils.generate_image_path_map(root_folder, valid_extensions=('tif', 'tiff', 'png', 'jpg', 'jpeg', 'bmp', 'czi', 'nd2', 'lif'))[source]¶
Recursively scans a folder and its subfolders for images, then creates a mapping of: {original_image_path: new_image_path}, where the new path includes all subfolder names.
- spacr.utils.generate_path_list_from_db(db_path, file_metadata)[source]¶
Return all
png_pathvalues fromdb_pathoptionally filtered byfile_metadatasubstrings.- Parameters:
db_path – path to the measurements SQLite DB.
file_metadata – substring or list of substrings to LIKE-match against
png_path.
- Returns:
list of PNG paths.
- spacr.utils.get_cuda_version()[source]¶
Return the installed CUDA toolkit version as a digit-only string, or
None.Parses the
nvcc --versionoutput; the dots are stripped so11.8becomes"118".- Returns:
Version string without dots, or
Noneifnvccis missing or fails.
- spacr.utils.get_db_paths(src)[source]¶
Return the standard
measurements/measurements.dbpaths for one or more source roots.- Parameters:
src – plate folder, or list of plate folders for a multi-plate run. A bare string is wrapped, so the return type is always a list and a single-plate caller still has to index
[0]. These are the run roots that Measure wrote into, not themeasurementsfolder itself – themeasurements/measurements.dbsuffix is appended here. Nothing is checked for existence, so a typo produces a path that only fails later at connect time.
- spacr.utils.get_files_from_dir(dir_path, file_extension='*')[source]¶
Return glob matches for
dir_path/file_extension.- Parameters:
dir_path – directory to list. It is joined to the pattern rather than walked, so subdirectories are never searched and a nonexistent path yields an empty list instead of an error.
file_extension – despite the name this is a full glob pattern, not a suffix – pass
'*.tif', not'.tif'or'tif', or nothing matches. Matching follows the filesystem’s case sensitivity, and dotfiles are excluded byglobsemantics. Default'*'returns every non-hidden entry, directories included.
- spacr.utils.get_ml_results_paths(src, model_type='xgboost', channel_of_interest=1)[source]¶
Return the standard set of ML output paths for the given model and channel selection.
- Parameters:
src – experiment root.
model_type – model identifier (used in the results folder name).
channel_of_interest – int, list,
'morphology', orNone(aliased toall_features).
- Returns:
10-tuple of paths
(data, permutation, feature_importance, model_metrics, permutation_fig, feature_importance_fig, shap_fig, plate_heatmap, settings, ml_features).- Raises:
ValueError – if
channel_of_interesthas an unsupported type.
- spacr.utils.get_paths_from_db(df, png_df, image_type='cell_png')[source]¶
Return rows of
png_dfwhose path containsimage_typeand whoseprcfois indf.- Parameters:
df – DataFrame indexed by
prcfoidentifiers.png_df – DataFrame of PNG metadata with
png_pathandprcfocolumns.image_type – substring that must appear in
png_path.
- Returns:
filtered subset of
png_df.
- spacr.utils.get_sequencing_paths(src)[source]¶
Return the standard
sequencing/sequencing_data.csvpaths for one or more source roots.- Parameters:
src – plate folder, or list of plate folders, given in the same order as the corresponding
get_db_paths()call – the two lists are zipped positionally when measurements are joined to barcode counts, so a reordered list silently pairs a plate’s images with another plate’s reads. A bare string is wrapped, so the result is always a list, and no path is checked for existence.
- spacr.utils.get_submodules(model, prefix='')[source]¶
Return all dotted submodule names of
modelin traversal order.- Parameters:
model – PyTorch module to walk.
prefix – optional prefix prepended to returned names.
- Returns:
list of dotted submodule names.
- spacr.utils.group_feature_class(df, feature_groups=None, name='compartment')[source]¶
Add a column tagging each feature with its compartment (or other group) label.
Matches feature names against the tokens in
feature_groupsand stores the result in a new columnname. Whenname == 'channel', unmatched features are relabeled'morphology'.- Parameters:
df – DataFrame with a
featurecolumn.feature_groups – Iterable of substrings/regex tokens to look for in each feature name. Defaults to
['cell', 'cytoplasm', 'nucleus', 'pathogen'].name – Name of the column added to
df. Default'compartment'.
- Returns:
dfwith the new group column populated.
- spacr.utils.initiate_counter(counter_, lock_)[source]¶
Initialize shared multiprocessing
counterandlockglobals.- Parameters:
counter – shared
multiprocessing.Valuecounter.lock – shared
multiprocessing.Lockguarding the counter.
- Returns:
None.
- spacr.utils.invert_image(image)[source]¶
Return the intensity-inverted image, reflected through the dtype range.
The pivot is
iinfo.min + iinfo.max, which is the dtype maximum for every unsigned dtype – souint8anduint16invert exactly as they always did – and-1for a signed one, the same conventionskimage.util.invert()uses. Reflecting through the range instead of subtracting from the ceiling is what keeps a signed image in range: underint8,-100inverts to99rather than to227, which used to wrap silently to-29.- Parameters:
image – array with an integer dtype. A float or boolean image raises
ValueErrorrather than inverting – convert or rescale to an integer dtype first. The pivot is the dtype range, not the image range, so a dimuint16image inverts against 65535 and comes back near-white; normalize to the dtype range first if you want a contrast-preserving inversion.- Returns:
the inverted image, in the input dtype. Every value stays in range, so nothing wraps.
- Raises:
ValueError –
imagedoes not have an integer dtype.
- spacr.utils.is_list_of_lists(var)[source]¶
Return
Trueifvaris a list whose every element is also a list.- Parameters:
var – value to test, including an empty or nested list.
- spacr.utils.is_multiprocessing_process(process)[source]¶
Return
Trueifprocesscmdline containsmultiprocessing.- Parameters:
process – process object exposing the
psutilcmdlineAPI.
- spacr.utils.jaccard_index(mask1, mask2)[source]¶
Return the Jaccard/IoU index of two binary masks.
- Parameters:
mask1 – array of any shape; nonzero is foreground, so a multi-label mask collapses to one merged object.
mask2 – array that must already have the same shape as
mask1– there is no padding step here, so mismatched shapes either raise a broadcast error or, worse, broadcast silently against a length-1 axis. Usecalculate_iou()when the shapes can differ; it also returns0for two empty masks, whereas this divides by zero and returnsnanwith a runtime warning.
- spacr.utils.lasso_reg(merged_df, alpha_value=0.01, reg_type='lasso')[source]¶
Fit Lasso or Ridge on one-hot-encoded gene/grna/plate/row/column predictors.
- Parameters:
merged_df – DataFrame with
gene,grna,plateID,rowID,columnID,pred.alpha_value – regularization strength.
reg_type –
'lasso'or'ridge'.
- Returns:
DataFrame with
FeatureandCoefficientcolumns.
- spacr.utils.load_image(image_path)[source]¶
Load and preprocess an image.
The preprocessing is fixed to the ImageNet recipe used by
extract_features(): resize to 224x224, then normalize with the ImageNet channel means and standard deviations. None of it is configurable.- Parameters:
image_path – path to any file PIL can open. It is forced through
convert('RGB'), so a 16-bit or float microscopy TIFF is downcast to 8-bit and a single-channel image is replicated across three channels rather than rejected – the dynamic range of a raw scientific image is lost here, so rescale to 8-bit yourself if that matters. Aspect ratio is not preserved:Resize((224, 224))takes both dimensions, so non-square crops are stretched, not letterboxed. Returns a(1, 3, 224, 224)tensor with the batch axis already added, so it can be fed to a model directly but must be concatenated, not stacked, to batch several images.
- spacr.utils.load_image_paths(c, visualize)[source]¶
Load the
png_listtable into a DataFrame indexed byprcfoand optionally filter by object.- Parameters:
c – open sqlite3 cursor.
visualize – object-type prefix (
'cell'/'nucleus'/…) or falsy to keep all rows.
- Returns:
DataFrame of PNG metadata indexed by
prcfo.
- spacr.utils.load_settings(csv_file_path, show=False, setting_key='setting_key', setting_value='setting_value')[source]¶
Reload a spacr settings CSV (written by
save_settings()) back into a Python dict.Every spacr pipeline persists its resolved settings alongside its outputs so that a run can be reproduced. This helper re-parses that CSV, coercing each value into its original Python type (
bool,int,float,None,list,tuple,dict,str).- Parameters:
csv_file_path – path to the CSV file.
show – display the raw DataFrame for debugging. Default
False.setting_key – name of the key column. Default
'setting_key';'Key'is accepted too, see below.setting_value – name of the value column. Default
'setting_value';'Value'is accepted too.
- Returns:
dict of parsed settings, ready to pass back into the original pipeline entry point.
- Raises:
ValueError – if the required key / value columns are missing.
THE TWO SPELLINGS.
save_settings()writesKey/Value, while this function’s defaults ask forsetting_key/setting_value– so the documented inverse pair did not round-trip, and the example above raised. Callers had each worked around it separately (spacr/qt/dnd.pytries one spelling and catches the failure to try the other), which is how it survived: nothing that used the defaults was reading a file spacr had written.Either spelling is now read. An explicitly named column still wins, so a caller that knows its file’s header is unaffected.
Example
from spacr.utils import load_settings from spacr.core import preprocess_generate_masks settings = load_settings('/data/plate01/settings/gen_mask_settings.csv') preprocess_generate_masks(settings)
See also
save_settings()— inverse operation.
- spacr.utils.map_condition(col_value, neg='c1', pos='c2', mix='c3')[source]¶
Map a column-ID value to one of
'neg','pos','mix', or'screen'.- Parameters:
col_value – Column identifier from the plate metadata.
neg – Column ID that corresponds to negative controls. Default
'c1'.pos – Column ID that corresponds to positive controls. Default
'c2'.mix – Column ID that corresponds to mixed controls. Default
'c3'.
- Returns:
Condition label; any unlisted column returns
'screen'.
- spacr.utils.mask_object_count(mask)[source]¶
Return the number of nonzero labeled objects in
mask.- Parameters:
mask – integer label image where
0is background. The count is the number of distinct nonzero values, so label IDs need not be contiguous and gaps left by filtering are not counted. A purely binary mask therefore reports1however many blobs it contains.
- spacr.utils.match_masks(true_masks, pred_masks, iou_threshold)[source]¶
Greedy match each predicted mask to a still-unmatched true mask above
iou_threshold.- Parameters:
true_masks – iterable of ground-truth masks.
pred_masks – iterable of predicted masks.
iou_threshold – minimum IoU to count as a match.
- Returns:
list of
(true_mask, pred_mask)matched pairs.
- spacr.utils.measure_test_mode(settings)[source]¶
Copy a random subset of source files into a
test/mergedfolder whentest_modeis on.Fewer files than
test_nris not an error. test_mode is the setting a user reaches for on a SMALL plate, andrandom.sampleraisedValueError: Sample larger than population or is negativeon exactly that case – so the one folder you most want to smoke-test first was the one folder test_mode refused to run on.Only visible
.npyarrays are sampled, so a macOS._sidecar is never measured in place of a field. The folder’s.spacr_plane_layout.jsonis copied across as well when it exists: it is what says which plane is which, and atest/mergedwithout it is read as a legacy folder, against the default plane order.- Parameters:
settings – settings dict; must contain
src,test_mode,test_nr.- Returns:
settings dict with
srcoptionally redirected to the test folder.- Raises:
ValueError – if there is nothing to sample – an empty
src, or atest_nrbelow 1. Sampling zero files would pointsrcat an emptytest/mergedand the run would report “no fields found”, blaming the wrong thing.
- spacr.utils.merge_dataframes(df, image_paths_df, verbose)[source]¶
Merge
dfintoimage_paths_dfon the sharedprcfoindex.- Parameters:
df – feature DataFrame with a
prcfocolumn.image_paths_df – DataFrame indexed by
prcfo.verbose – display the merged DataFrame.
- Returns:
merged DataFrame.
- spacr.utils.merge_regression_res_with_metadata(results_file, metadata_file, name='_metadata')[source]¶
Merge regression outputs with gene metadata on the parsed
genecolumn.- Parameters:
results_file – path to a regression results CSV with a
featurecolumn.metadata_file – path to a gene metadata CSV with a
Gene IDcolumn.name – suffix appended to the output filename.
- Returns:
merged DataFrame (also written to
<results_file><name>.csv).
- spacr.utils.merge_split_objects(mask_src, intensity_img_src=None, intensity_channel=None, perimeter_fraction=0.5, min_area=0, max_area=0, remove_border_objects=False, n_jobs=1, progress_callback=None, op_name='', *, min_intensity=0, max_intensity=0, filters=None)[source]¶
Merge by perimeter and filter labeled objects across a directory of masks.
Runs the shared in-memory merge/filter pipeline on each mask file in
mask_srcin parallel, overwriting each mask in place.- Parameters:
mask_src – directory containing mask .tif/.tiff/.npy files.
intensity_img_src – directory of matched original intensity images, required when either intensity bound is enabled.
intensity_channel – explicit channel-last index for multi-channel intensity images; unnecessary for single-channel planes.
perimeter_fraction – minimum shared-boundary fraction for perimeter-based merging.
min_area – remove objects smaller than this (px); 0 disables.
max_area – remove objects larger than this (px); 0 disables.
remove_border_objects – drop objects touching the image border.
n_jobs – parallel worker count.
progress_callback – optional callback(fov_index, total, duration, op_name).
op_name – label passed to the progress callback.
min_intensity – remove objects whose own-channel mean is below this raw-image value; equality is kept and 0 disables the lower bound.
max_intensity – remove objects whose own-channel mean is above this raw-image value; equality is kept and 0 disables the upper bound.
filters – object filter entries (any scalar regionprop with a minimum and a maximum), judged with the bounds above in one pass.
- Returns:
None.
- spacr.utils.merge_touching_objects(mask, threshold=0.25)[source]¶
Merge touching labeled objects whose shared boundary exceeds
thresholdof the smaller perimeter.- Parameters:
mask – labeled mask.
threshold – fraction of the smaller perimeter required to merge.
- Returns:
merged label mask.
- spacr.utils.model_metrics(model)[source]¶
Print RMSE/MAE/Durbin-Watson and show residual/QQ/scale-location diagnostic plots.
- Parameters:
model – fitted statsmodels regression result.
- Returns:
None.
- spacr.utils.normalize_feature_filter(filter_by)[source]¶
Normalize text representations of an unfiltered feature selection.
Settings imported from CSV files and older Qt sessions can contain the literal string
"None". Treating that as a feature-name substring removes every measurement column, although the UI means “all channels”.- Parameters:
filter_by – the raw setting value. A string is stripped and, if it case-insensitively matches one of the “no filter” spellings (
"","none","null","all","all_channels","all channels","*"), collapsed toNone– otherwise the stripped string is returned as a feature-name substring. Anything that is not a string (a realNone, or a list of channels) passes through untouched, so this is safe to apply unconditionally to whatever the settings dict holds. Note"*"means no filter, not a glob: real patterns are not supported, and the surviving string is matched as a plain substring of the column name.
- spacr.utils.normalize_src_path(src)[source]¶
Ensures that the ‘src’ value is properly formatted as either a list of strings or a single string.
- spacr.utils.normalize_to_dtype(array, p1=2, p2=98, percentile_list=None, new_dtype=None)[source]¶
Percentile-normalize each channel of an image stack into the target dtype range.
- Parameters:
array – input stack of shape
(H, W, C).p1 – lower percentile. Default
2.p2 – upper percentile. Default
98.percentile_list – per-channel
(low, high)pairs; overridesp1/p2.new_dtype – target dtype (
np.uint8/np.uint16or their string forms).
- Returns:
normalized stack with the same shape as
array.
- spacr.utils.object_label_from_png_id(values)[source]¶
Migrate
png_list’s'o<N>'text ids onto the integer object label.png_liststores an object id as text ('o5') because it is the last component ofprcfo; every object table stores the same object as an integerobject_label, and the child tables store their parent as an integer (in practice a float, sincemeasurewrites NaN for “no overlapping cell”)cell_id. Two types for one identity, which is why a plain SQLpng_list.cell_id = nucleus.cell_idmatches zero rows rather than failing: SQLite compares a TEXT value with an INTEGER one by type class, and text always sorts after numbers. Measured on a database built by the real writers: 6 crops, 6 nuclei, 0 rows joined.The integer is canonical — it is what the measurement tables key on — so this is the one migration, applied on read. It replaces
series.str[1:].astype(int), which crashed on four values the real writers genuinely produce:'omulti'and'onone'—_generate_names()names a crop that overlaps several cells..._multi.pngand one that overlaps none..._none.png. Both are ordinary outcomes of a real segmentation.ValueError: invalid literal for int() with base 10: 'multi';'error'— what_map_wells_png()writes for a name it cannot parse..str[1:]turned it into'rror', so the exception did not even name the problem;NULL— every row of a different crop mode, in a database measured with more than one.TypeError: int() argument must be ... not 'NoneType';an already-integer column, from a database whose ids were migrated elsewhere:
.strraisesAttributeErroron a numeric Series.
All four now come back as
NaN, which a caller can count and drop — losing the crop’s path for those objects, never the whole read.- Parameters:
values – a
png_listobject-id column (cell_id,nucleus_id, …), of any dtype.- Returns:
a float
Seriesof object labels,NaNwhere the id holds no integer. Float rather than int becauseNaNhas no int64.
- spacr.utils.pad_to_same_shape(mask1, mask2)[source]¶
Zero-pad
mask1andmask2to their element-wise maximum shape.Padding is appended at the bottom and right only, so the two masks are aligned on their top-left corner. This is an alignment assumption, not a registration: crops taken from different offsets are not brought into correspondence by padding them.
- Parameters:
mask1 – 2-D array. Only axes 0 and 1 are considered, so a 3-D stack is padded on its first two axes and left ragged on the third.
mask2 – 2-D array padded to the same element-wise maximum shape. Each mask is padded independently, so the larger one along a given axis is returned untouched on that axis.
- spacr.utils.perform_statistical_tests(all_df, cluster_col='cluster')[source]¶
Perform ANOVA or Kruskal-Wallis tests depending on normality of features.
Each numeric feature is tested for normality and then sent to either ANOVA or Kruskal-Wallis across the groups defined by
cluster_col, never both.- Parameters:
all_df – DataFrame with the numeric feature columns and the cluster column.
cluster_col – Column holding the cluster label, excluded from the features. Default
'cluster'.
- Returns:
(anova_df, kruskal_df), with columnsFeature/ANOVA_Statistic/ANOVA_pValueandFeature/Kruskal_Statistic/Kruskal_pValuerespectively; together they cover the features exactly once.
- spacr.utils.pick_best_model(src)[source]¶
Return the strongest checkpoint anywhere below
src.Current artifacts are ranked by their stored validation metric and role; legacy files fall back to their
_acc_/_epoch_filename fields.- Parameters:
src – model directory or a checkpoint path.
- Returns:
absolute path to the top-ranked checkpoint.
- spacr.utils.plot_clusters(ax, embedding, labels, colors, cluster_centers, plot_outlines, plot_points, smooth_lines, figuresize=10, dot_size=50, verbose=False, point_color='cluster', point_alpha=0.65, outline_width=1.0)[source]¶
Draw cluster outlines, points, and centroid labels onto
axfor a 2-D embedding.- Parameters:
ax – Matplotlib axes to draw into.
embedding –
(N, 2)array of 2-D points (e.g. UMAP output).labels – length-
Ncluster labels;-1denotes noise.colors – iterable of per-cluster colors, one per unique label.
cluster_centers – iterable of
(x, y)centroids, one per unique label.plot_outlines – draw a hull/smoothed outline around each cluster.
plot_points – render the scatter points (otherwise plotted invisibly).
smooth_lines – use a smoothed hull polyline instead of the convex hull edges.
figuresize – base size in inches used to scale axis label and tick fonts. Default
10.dot_size – scatter marker size in points. Default
50.verbose – unused placeholder kept for API compatibility. Default
False.point_color –
'cluster'/'viridis'(or empty) colors points per cluster; any other Matplotlib color is applied to every point. Default'cluster'.point_alpha – scatter opacity, clamped to
[0, 1]; ignored whenplot_pointsisFalse. Default0.65.outline_width – hull line width in points, floored at
0.1. Default1.0.
- Returns:
None.
- spacr.utils.plot_clusters_grid(embedding, labels, image_nr, image_paths, colors, figuresize, black_background, verbose, theme_colors=None)[source]¶
Plot a grid of example images per cluster label discovered in
labels.- Parameters:
embedding – accepted and never read – the panels are built from
labelsandimage_pathsalone, soNoneworks.labels – cluster labels in the row order of
image_paths.-1is dropped as noise, and if nothing else remains the function printsNo clusters found.and returnsNoneinstead of a figure.image_nr – per-cluster cap on how many images are opened. A cluster no larger than this contributes all of its members.
image_paths – paths addressed positionally by the label array; a list shorter than
labelsraisesIndexError.colors – palette indexed downstream by the cluster LABEL itself rather than by its rank, so the palette has to be long enough to reach the largest label – labels
0and5against a two-color palette raiseIndexError. Entries need at least three components.figuresize – per-cluster panel size in inches, shrunk downstream so the whole row never exceeds 200 inches.
black_background – picks the white-on-black fallback theme instead of black-on-white;
theme_colorsoverrides it per role.verbose – only ever prints for STRING cluster labels; silent for the integer labels DBSCAN and KMeans produce.
theme_colors – dict with
background/foreground/bordercolors. Entries Matplotlib cannot parse are dropped silently and fall back to theblack_backgroundchoice. DefaultNone.
- Returns:
the Matplotlib
Figure, orNonewhen every label is-1.
- spacr.utils.plot_embedding(embedding, image_paths, labels, image_nr, img_zoom, colors, plot_by_cluster, plot_outlines, plot_points, plot_images, smooth_lines, black_background, figuresize, dot_size, remove_image_canvas, verbose, interactive_payload=None, theme_colors=None, point_color='cluster', point_alpha=0.65, outline_width=1.0)[source]¶
Plot a 2-D embedding with cluster outlines, points, and optional image overlays.
- Parameters:
embedding –
(N, 2)array of 2-D points (e.g. UMAP output).image_paths – length-
Nimage paths used for the overlays;Noneskips them.labels – length-
Ncluster labels;-1denotes noise.image_nr – number of images to overlay (per cluster when
plot_by_cluster).img_zoom – zoom factor applied to each overlaid thumbnail.
colors – palette of per-cluster colors, one entry per unique label.
plot_by_cluster – sample the overlaid images per cluster instead of at random.
plot_outlines – draw a hull/smoothed outline around each cluster.
plot_points – render the scatter points (otherwise plotted invisibly).
plot_images – overlay the images from
image_paths.smooth_lines – use a smoothed hull polyline instead of the convex hull edges.
black_background – use the white-on-black default theme instead of black-on-white; entries in
theme_colorsoverride it per role.figuresize – figure side length in inches; also scales label and tick fonts.
dot_size – scatter marker size in points.
remove_image_canvas – mask out zero-valued pixels of each overlaid image.
verbose – forwarded to the cluster and image helpers, which ignore it.
interactive_payload – optional object stashed on the figure as
_spacr_umap_payloadso the Qt bridge can keep point/image identities. DefaultNone.theme_colors – dict with
background/foreground/bordercolors. DefaultNone.point_color –
'cluster'/'viridis'colors points per cluster; any other Matplotlib color is applied to every point. Default'cluster'.point_alpha – scatter opacity, clamped to
[0, 1]. Default0.65.outline_width – hull line width in points, floored at
0.1. Default1.0.
- Returns:
matplotlib
Figure.
- spacr.utils.plot_grid(cluster_images, colors, figuresize, black_background, verbose, theme_colors=None)[source]¶
Render one column per cluster of representative images with colored borders and labels.
- Parameters:
cluster_images – ordered mapping of cluster label to that cluster’s list of image arrays; one column per key, and an empty mapping raises
ValueErrorfromsubplots.colors – palette used consistently for both panel borders and legend swatches. Integer labels index it by label (wrapping when necessary), while string labels use their position in
cluster_images. An empty palette falls back to neutral grey, and entries need at least three components.figuresize – figure height in inches and the label font size; the width is this times the cluster count. It is shrunk to
200 / n_clusterswhen that product would exceed 200 inches, which silently caps the font size too.black_background – picks the white-on-black fallback theme instead of black-on-white.
verbose – prints the label and its index for STRING cluster labels only; integer labels never print anything.
theme_colors – dict with
background/foreground/bordercolors overriding theblack_backgroundfallback; values Matplotlib cannot parse are ignored. DefaultNone.
- Returns:
the Matplotlib
Figure, which is also passed toplt.show.
- spacr.utils.plot_image(ax, x, y, img, img_zoom, remove_image_canvas=True)[source]¶
Place a zoomed thumbnail of
imgat(x, y)onax.- Parameters:
ax – axes the thumbnail is added to, as a frameless annotation box.
x – data-space x coordinate the thumbnail is anchored at.
y – data-space y coordinate the thumbnail is anchored at.
img – PIL image when
remove_image_canvasis true, sinceimg.modeis read; any array-like otherwise.img_zoom – scale factor handed to
OffsetImage. It sizes the thumbnail from the source’s pixel dimensions in display space, so the drawn size is unchanged by the axis limits.remove_image_canvas – true swaps the image for an RGBA array whose alpha channel hides zero-valued pixels, which accepts only PIL modes
L,IandRGB–RGBAandPraiseValueError, and a numpy array raisesAttributeErrorbecause it has nomode. An all-zeroLimage divides by its own zero maximum and comes outNaNrather than raising. False just callsnp.array. DefaultTrue.
- Returns:
None.
- spacr.utils.plot_images_by_cluster(ax, image_paths, embedding, labels, image_nr, img_zoom, colors, cluster_indices, remove_image_canvas, verbose)[source]¶
Overlay up to
image_nrimages per cluster on the embedding inax.- Parameters:
ax – axes the thumbnails are added to, as frameless annotation boxes.
image_paths – paths addressed by the indices held in
cluster_indices, so they must be in the embedding’s row order.embedding –
(N, 2)array supplying each thumbnail’s position.labels – only
np.unique(labels)is used, to decide which clusters to visit;-1is skipped as noise.image_nr – per-cluster cap. A cluster no larger than this contributes all of its members – no sampling happens.
img_zoom – scale factor handed to
OffsetImage, applied to the file’s own pixel dimensions rather than to data units.colors – accepted for caller compatibility but not read. Thumbnail overlays do not use a cluster color, and palette length no longer limits how many labels are visited.
cluster_indices – mapping of label to the row indices to draw from. Looked up with
.get(label, []), so a label present inlabelsbut absent here plots nothing instead of raising.remove_image_canvas – forwarded to
plot_image().verbose – accepted and ignored; any value at all is tolerated.
- Returns:
None.
- spacr.utils.plot_umap_images(ax, image_paths, embedding, labels, image_nr, img_zoom, colors, plot_by_cluster, remove_image_canvas, verbose)[source]¶
Overlay sample images from
image_pathson the UMAP embedding inax.- Parameters:
ax – axes the thumbnails are added to, as frameless annotation boxes.
image_paths – paths addressed by the same positional index as
embedding, so the two must share a row order; a short list raisesIndexError.embedding –
(N, 2)array whose selected rows give each thumbnail its data-space position.labels – cluster labels aligned with
embedding. Read only whenplot_by_clusteris true;Noneis accepted otherwise.image_nr – with
plot_by_clusterfalse, the exact number of rows sampled at random from the whole embedding, so a value aboveNraisesValueErrorfromrandom.sample. With it true, a per-cluster cap – a cluster no larger than this contributes every member, unsampled.img_zoom – scale factor handed to
OffsetImage. It sizes the thumbnail from the file’s own pixel dimensions in display space, so rescaling the axes does not change how big the image is drawn.colors – only zipped against
np.unique(labels)to drive the iteration; the color itself is never drawn. Its LENGTH is therefore a silent limit, and becausenp.uniqueincludes the-1noise label a palette sized to the real clusters leaves the last cluster with no images. Unused (Noneis fine) whenplot_by_clusteris false.plot_by_cluster – true samples per cluster and skips label
-1; false ignoreslabelsandcolorsentirely and samples globally.remove_image_canvas – forwarded to
plot_image(); true masks zero-valued pixels out and restricts the inputs to PIL modesL,IandRGB.verbose – accepted and ignored, here and in the helper it is passed to; any value at all is tolerated.
- Returns:
None.
- spacr.utils.prepare_batch_for_segmentation(batch)[source]¶
Cast a batch to
float32and per-image max-normalize any image whose max exceeds 1.- Parameters:
batch –
(N, ...)numpy array of images.- Returns:
The same array cast to
float32with each image scaled to[0, 1].
- spacr.utils.preprocess_data(df, filter_by, remove_highly_correlated, log_data, exclude, column_list=False, *, batch_correction='none', batch_column='plateID', batch_control_column=None, batch_control_values=None, batch_covariate_column=None, batch_combat_mean_only=False, batch_min_samples=3, batch_missing_control='error')[source]¶
Prepare a feature matrix by filtering, decorrelating, log-transforming, and scaling
df.- Parameters:
df – input DataFrame.
filter_by – channel of interest passed to
filter_dataframe_features();Noneand its text forms disable filtering.remove_highly_correlated – correlation cutoff (float) or
Trueto use0.95;Falsedisables.log_data – apply
log(x + 1e-6)to numeric columns.exclude – features to exclude from filtering.
column_list – optional explicit column subset applied before selecting numeric columns.
batch_correction –
none,center,zscore,robust_zscore,control_centerorcombat.batch_column – metadata column identifying acquisition batches.
batch_control_column – metadata column selecting reference controls.
batch_control_values – reference value(s) for
control_center.batch_covariate_column – metadata column(s) naming the biology
combatmust preserve. Required bycombat, ignored by every other method — and left blank,combatrefuses to run rather than removing the contrast along with the plate effect.batch_combat_mean_only – correct only
combat’s additive shift and leave each batch’s scale alone.batch_min_samples – minimum rows/reference controls per batch.
batch_missing_control –
errororskipwhen a batch lacks enough controls.
- Returns:
standard-scaled
ndarrayof numeric features.- Raises:
ValueError – if no numeric columns remain after filtering.
- spacr.utils.preprocess_image(image_path, normalize=True, image_size=224, channels=None)[source]¶
Load and preprocess
image_pathinto a batched tensor ready for classification.- Parameters:
image_path – path to the source image.
normalize – apply ImageNet mean/std normalization.
image_size – square resize dimension.
channels – reserved for downstream use; kept for API compatibility.
- Returns:
(pil_image, input_tensor)where the tensor has shape(1, 3, H, W).
- spacr.utils.pretty_print_settings(settings, title='Settings')[source]¶
Print a settings dict to the console as a tidy, aligned table.
Nicer than dumping a truncated pandas DataFrame: values are grouped by the spacr settings categories, keys are aligned in a column, long values are clipped, and the whole thing sits under a boxed title. Purely cosmetic – used wherever “Saving settings” is shown.
Purely cosmetic, and it stays that way: the frame degrades to ASCII and every line goes out through
console_safe(), so a console that cannot encode the decoration prints a plainer table instead of raisingUnicodeEncodeError.spacr.measure.measure_crop()calls this (throughsave_settings()) before it does any work at all, so a decoration character was enough to end a whole run before the first field was read.- Parameters:
settings – the settings dict to render.
title – heading shown in the box.
- Returns:
None.
- spacr.utils.print_progress(files_processed, files_to_process, n_jobs, time_ls=None, batch_size=None, operation_type='')[source]¶
Print a one-line progress report with an ETA derived from mean step time.
- Parameters:
files_processed – number of items done (int or list).
files_to_process – total items to do (int or list).
n_jobs – parallelism used to compute ETA.
time_ls – list of per-step durations (seconds) for ETA;
Noneskips ETA.batch_size – batch size when
time_lsis per batch rather than per image.operation_type – label printed alongside the progress line.
- Returns:
None.
- spacr.utils.process_mask_file_adjust_cell(file_name, parasite_folder, cell_folder, nuclei_folder, organelle_folder=None, overlap_threshold=5, perimeter_threshold=30, *, output_folder=None)[source]¶
Load one triple of parasite/cell/nuclei masks, merge cells in place, and return the elapsed time.
- Parameters:
file_name – mask file name (must exist in all folders).
parasite_folder – folder of parasite masks.
cell_folder – folder of cell masks (overwritten in place). The adjusted mask replaces the old one atomically, so a run killed during the write leaves the previous whole mask, never a truncated one.
nuclei_folder – folder of nuclei masks.
organelle_folder – optional folder of organelle masks.
overlap_threshold – fractional overlap threshold used by the merger.
perimeter_threshold – shared-perimeter threshold used by the merger.
output_folder – optional separate destination for adjusted masks. None retains in-place adjustment. An explicit destination must differ from every source mask folder, including through directory symlinks.
- Returns:
elapsed seconds.
- Raises:
ValueError – if the matching cell or nuclei mask file is missing, or a mask file holds pickled objects: masks are plain arrays, and nothing is unpickled.
- spacr.utils.process_masks(mask_folder, image_folder, channel, batch_size=50, n_clusters=2, plot=False)[source]¶
Cluster object morphology/intensity across a mask folder and keep the largest cluster in place.
- Parameters:
mask_folder – folder of
.npymasks.image_folder – matching folder of
.npyintensity images.channel – channel index used for intensity measurements.
batch_size – number of files to load per batch.
n_clusters – number of KMeans clusters.
plot – show a PCA scatter of the clustered objects.
- Returns:
None.
- spacr.utils.process_vision_results(df, threshold=0.5)[source]¶
Split image paths into well identifiers and binarize the
predcolumn.- Parameters:
df – DataFrame with
pathandpredcolumns.threshold – cutoff used to derive
cv_predictions.
- Returns:
enriched DataFrame with
plateID,rowID,columnID,fieldID,prc,cv_predictions.
- spacr.utils.random_forest_feature_importance(all_df, cluster_col='cluster')[source]¶
Rank features by how well they predict the cluster label.
Z-scales the numeric feature columns and fits a 100-tree
RandomForestClassifieragainstcluster_col.- Parameters:
all_df – DataFrame with the numeric feature columns and the cluster column.
cluster_col – Column holding the cluster label, excluded from the features. Default
'cluster'.
- Returns:
DataFrame with
FeatureandImportancecolumns, sorted by descending importance.
- spacr.utils.recommend_target_layers(model)[source]¶
Return
([last_conv_layer], all_conv_layers)frommodel.- Parameters:
model – PyTorch module to scan for
Conv2dlayers.- Returns:
tuple
(recommended, all)of layer-name lists.- Raises:
ValueError – if the model contains no convolutional layers.
- spacr.utils.reduction_and_clustering(numeric_data, n_neighbors, min_dist, metric, eps, min_samples, clustering, reduction_method='umap', verbose=False, embedding=None, n_jobs=-1, mode='fit', model=False, reducer_options=None, prefer_gpu=False, random_seed=42)[source]¶
Reduce
numeric_datato 2-D and cluster the embedding.Supported reducers are UMAP, t-SNE, PCA, Isomap and Spectral Embedding.
reducer_optionscarries only method-specific settings; irrelevant options are never forwarded. RAPIDS is opt-in and applies to UMAP, t-SNE and PCA, with the actual backend retained on the fitted reducer.- Parameters:
numeric_data – rows of numeric features to embed and cluster.
n_neighbors – reducer neighborhood size, or a row fraction as a float; also supplies the default t-SNE perplexity.
min_dist – minimum embedding distance used by UMAP.
metric – distance metric used by the reducer and DBSCAN.
eps – DBSCAN neighborhood radius.
min_samples – DBSCAN minimum neighborhood size, or KMeans cluster count when
clustering='kmeans'.clustering – clustering algorithm,
'dbscan'or'kmeans'.
- spacr.utils.remove_canvas(img)[source]¶
Return
imgas RGBA with zero-valued pixels made transparent.- Parameters:
img – PIL image in
L,I, orRGBmode.
Drop numeric columns whose absolute correlation with a prior column exceeds
threshold.- Parameters:
df – input DataFrame.
threshold – correlation cutoff.
verbose – print the dropped column names.
- Returns:
decorrelated DataFrame.
- spacr.utils.remove_intensity_objects(image, mask, intensity_threshold, mode)[source]¶
Drop labeled objects whose mean intensity is on the wrong side of
intensity_threshold.- Parameters:
image – intensity image.
mask – labeled mask aligned to
image.intensity_threshold – cutoff value.
mode –
'low'removes below-threshold objects,'high'removes above.
- Returns:
filtered label mask.
- spacr.utils.remove_low_variance_columns(df, threshold=0.01, verbose=False)[source]¶
Drop numeric columns whose variance is below
threshold.- Parameters:
df – input DataFrame.
threshold – variance cutoff.
verbose – print the dropped column names.
- Returns:
filtered DataFrame.
- spacr.utils.remove_noise(embedding, labels)[source]¶
Drop rows of
embedding(andlabels) whose label is DBSCAN noise (-1).Rows are removed, not renumbered, so positional indices into the original data (an
image_pathslist, a DataFrame row order) no longer line up with the returned arrays – filter those alongside, using the same mask, or keep the identities before calling.- Parameters:
embedding –
(N, D)ndarray of points. It is filtered by boolean mask, so a Python list or a DataFrame will not index correctly; pass a numpy array.labels – length-
Nndarray of cluster labels, aligned row-for-row withembedding. Only-1is treated as noise, which is the DBSCAN convention – KMeans labels contain no-1and pass through unchanged, making this a no-op rather than an error on KMeans output.
- spacr.utils.remove_outliers_by_group(df, group_col, value_col, method='iqr', threshold=1.5)[source]¶
Removes outliers from
value_colwithin each group defined bygroup_col.Rows are selected, never modified: the original index is preserved and a new frame is returned. A row whose value is
NaNfails the comparison and is always dropped, whichever method is used.- Parameters:
df (pd.DataFrame) – The input DataFrame.
group_col (str) – Column name to group by, or a list of column names. Grouping passes
observed=False, so unused categories of a Categorical are kept. Rows whose group key is missing are discarded, because pandas dropsNaNgroup keys and the per-row bound then comes backNaN.value_col (str) – Column containing values to check for outliers. A name that is not in the frame raises
KeyError.method (str) – ‘iqr’ or ‘zscore’. Anything else raises
ValueError. The two now agree on tiny groups: a one-row group has an undefined standard deviation, and since one row cannot be an outlier within its own group it is KEPT under both. It used to be dropped by ‘zscore’ and kept by ‘iqr’.threshold (float) – Multiplier on the IQR (default 1.5), or the z-score cutoff. Must be >= 0; a negative value inverts the keep-band and is refused, because under ‘iqr’ it silently emptied every group with a nonzero IQR. Note
0under ‘zscore’ still keeps only rows sitting exactly on the group mean, which is what a zero cutoff means. Under ‘zscore’ an outlier inflates its own group’s standard deviation, so the usual cutoffs keep far more than ‘iqr’ does on the same data – that is the statistic, not a defect.
- Returns:
pd.DataFrame – A DataFrame with outliers removed.
- spacr.utils.rename_columns_in_db(db_path)[source]¶
Rename legacy column spellings across every table in a SQLite database.
Applies
DB_COLUMN_RENAMES— the plate-metadata names — and thenDB_COLUMN_RENAME_PATTERNS— the two feature families that were spelled inconsistently — to every user table. A rename is skipped when the target name already exists in that table, which gives three properties worth relying on:Idempotent. After a rename the legacy name is gone, so a second run finds nothing to do. Running it on every read is therefore free after the first.
Never destructive. A table that somehow carries both spellings — say
time_idandtimeID— keeps both, untouched. Neither column is dropped and nothing raises; the readers accept either spelling, so the data stays reachable and a human can decide which one is authoritative. Dropping or overwriting one of them here would destroy data to tidy a name, which is never the right trade.All or nothing. SQLite’s DDL is transactional, but Python’s sqlite3 driver only opens an implicit transaction for DML (INSERT/UPDATE/DELETE/ REPLACE) — an
ALTER TABLEruns in autocommit and lands immediately. So the previous version, which relied on a trailingcon.commit(), left a database half-migrated when a later rename raised. The transaction is opened explicitly here and rolled back on any error, and the connection is closed in afinally.
A partial migration would not corrupt anything — each rename is independently valid and the next read finishes the job — but “the schema changed and then the call raised” is not a state a user should have to reason about.
- Parameters:
db_path – Path to the SQLite database file to update in place.
- Returns:
The list of
(table, old, new)renames performed.
- spacr.utils.reset_cellpose_model_reports()[source]¶
Forget which Cellpose model notices have already been printed.
Call this at the start of a run so a second run in the same process (a GUI session segmenting a second plate) reports its model choice again instead of inheriting the first run’s silence.
- spacr.utils.reset_mp()[source]¶
Set the multiprocessing start method appropriate for the current OS.
Uses
spawnon Windows andforkon Linux/macOS.- Returns:
None.
- spacr.utils.resize_images_and_labels(images, labels, target_height, target_width, show_example=True)[source]¶
Resize aligned image/label lists to
target_heightxtarget_width.- Parameters:
images – iterable of source images (2-D or 3-D).
labels – matching iterable of label masks, or
None.target_height – output height in pixels.
target_width – output width in pixels.
show_example – display an example of the resized pair when
True.
- Returns:
(resized_images, resized_labels)lists.
- spacr.utils.resize_labels_back(labels, orig_dims)[source]¶
Resize a list of label masks back to their original
(width, height).- Parameters:
labels – iterable of label masks.
orig_dims – matching iterable of
(width, height)tuples.
- Returns:
list of resized label masks.
- Raises:
ValueError – if lengths differ or
orig_dimsentries are malformed.
- spacr.utils.save_file_lists(dst, data_set, ls)[source]¶
Write
lsas a single-column CSV named<data_set>.csvunderdst.- Parameters:
dst – destination directory.
data_set – column name and file stem.
ls – iterable of values to persist.
- Returns:
None.
- spacr.utils.save_settings(settings, name='settings', show=False)[source]¶
Persist a settings dict to
<src>/settings/<name>.csvso a spacr run can be reproduced later.Called by every pipeline entry point to snapshot the resolved settings before real work starts. The saved copy has
test_modeandplotforced toFalseso that a downstreamload_settings()-> re-run produces a full, headless run.- Parameters:
settings – settings dict; must contain
src.name – base filename (no extension);
_listis appended whensrcis a list. Default'settings'.show – display the DataFrame before writing. Default
False.
- Returns:
None. Writes
<src>/settings/<name>.csv.
Example
from spacr.utils import save_settings save_settings(my_settings, name='my_experiment', show=True)
See also
load_settings()— inverse operation.
- spacr.utils.search_reduction_and_clustering(numeric_data, n_neighbors, min_dist, metric, eps, min_samples, clustering, reduction_method, verbose, reduction_param=None, embedding=None, n_jobs=-1)[source]¶
Variant of
reduction_and_clustering()accepting extra reducer kwargs viareduction_param.- Parameters:
numeric_data – numeric data matrix.
n_neighbors – UMAP
n_neighborsor t-SNE perplexity (int or fraction).min_dist – UMAP
min_dist.metric – distance metric.
eps – DBSCAN
eps.min_samples – DBSCAN
min_samplesor KMeans cluster count.clustering –
'dbscan'or'kmeans'.reduction_method –
'umap'or'tsne'.verbose – print progress.
reduction_param – extra kwargs forwarded to the reducer.
embedding – precomputed embedding to skip fitting.
n_jobs – parallel worker count.
- Returns:
(embedding, labels).- Raises:
ValueError – on unsupported
reduction_methodorclustering.
- spacr.utils.setup_plot(figuresize, black_background, theme_colors=None)[source]¶
Create a square Matplotlib figure using scoped theme colors.
- Parameters:
- Returns:
tuple – The
(figure, axes)pair.
Notes
Theme values are applied inside
matplotlib.rc_context()and then to the created artists. Global Matplotlib settings are not modified.
- spacr.utils.show_cam_on_image(img, mask)[source]¶
Return
imgoverlaid with a jet colormap ofmaskas an 8-bit RGB image.The sum of heatmap and image is renormalized by its own peak, so the output brightness is relative to the single hottest pixel – two images overlaid separately are not comparable to each other on absolute intensity.
- Parameters:
img – 3-channel
(H, W, 3)image already scaled to[0, 1]. It is added to the colormap rather than blended, so a[0, 255]image swamps the heatmap and the result is a near-uniform wash. A 2-D grayscale array fails to broadcast against the 3-channel heatmap. An image negative enough that the blend has no positive pixel left raises rather than returning a black frame – a black attribution map is indistinguishable from “the model looked nowhere”, which is a claim this function must never make on the strength of bad input.mask –
(H, W)activation map in[0, 1], matchingimgin height and width. Values outside that range are CLIPPED to it, with aRuntimeWarning, so an un-normalized CAM saturates at the hot end instead of wrapping thenp.uint8cast: before this was clipped,1.1landed at the cold end of jet,1.5in the middle and2.0back at the top, which could render the hottest region of a map as the coldest colour. An all-zero mask does not produce a black overlay: jet maps 0 to a non-zero color, so a zero mask over a zero image renormalizes to a saturated flat field.
- Raises:
ValueError –
imgormaskcontains NaN or infinity, or the blend has no positive pixel to normalize against.
- spacr.utils.smooth_hull_lines(cluster_data)[source]¶
Return the x, y coordinates of a smoothed convex-hull outline of a 2-D point set.
- Parameters:
cluster_data – 2-D array of point coordinates.
- Returns:
tuple
(x, y)of spline-interpolated hull coordinates (100 samples).
- spacr.utils.split_my_dataset(dataset, split_ratio=0.1)[source]¶
Randomly split
datasetinto(train, val)subsets.- Parameters:
dataset – source dataset.
split_ratio – fraction of samples reserved for validation.
- Returns:
(train_subset, val_subset).
- spacr.utils.suggest_training_changes(dst, train_csv=None, val_csv=None, last_k=25, min_epochs=10, gap_threshold_acc=0.05, plateau_eps=0.001, noisy_var_ratio=0.03)[source]¶
Inspect saved training/validation progress CSVs and propose concrete training changes.
- Parameters:
dst – folder where progress CSVs were saved.
train_csv – explicit train-CSV path; auto-detected in
dstifNone.val_csv – explicit val-CSV path; auto-detected in
dstifNone.last_k – number of recent epochs used for trend and plateau checks.
min_epochs – minimum epochs before most suggestions are issued.
gap_threshold_acc – accuracy generalization-gap threshold (train - val).
plateau_eps – absolute slope threshold used to declare a plateau.
noisy_var_ratio – instability flag threshold on
stdev/meanof recent val loss.
- Returns:
dict with
summary(key scalars),flags(short codes), andsuggestions(ordered suggestion strings).
Nested helpers¶
- GradCAM.__call__.hook(module, input, output)¶
Forward hook: append the target layer’s output to
features.retain_grad()is required: PyTorch only populates.gradon leaf tensors, so without itfeatures[0].gradis None below and GradCAM died with “‘NoneType’ object has no attribute ‘cpu’”.spacr/utils.py:7084
- GradCAMGenerator.hook_layers.backward_hook(module, grad_input, grad_output)¶
Backward hook: cache the gradient flowing into the target layer’s output.
spacr/utils.py:6782
- GradCAMGenerator.hook_layers.forward_hook(module, input, output)¶
Forward hook: cache the target layer’s output activations.
spacr/utils.py:6778
- TorchModel._run_backbone_raw.forward_fn(t)¶
Run the underlying backbone on
t(used as the checkpoint target).spacr/utils.py:3978
- _checkpoint_module.contexts()¶
Return forward and recomputation contexts for non-reentrant checkpointing.
spacr/utils.py:191
- _filter_objects._failed_in(low, high)¶
Labels failing an entry whose index is in
[low, high).spacr/utils.py:687
- _generate_representative_images._compartment_column(compartment)¶
Return the selected compartment series, or raise for a missing column.
spacr/utils.py:2091
- _outline_and_overlay.process_dim(mask_dim)¶
Return a dilated outline image of the labeled mask at
image[..., mask_dim].spacr/utils.py:1846
- _pivot_counts_table._pivot_dataframe(df)¶
Pivot count-type rows into one column per object type, NaNs filled with 0.
spacr/utils.py:3372
- _pivot_counts_table._read_table_to_dataframe(db_path, table_name='object_counts')¶
Return the given SQLite table as a DataFrame.
spacr/utils.py:3367
- _save_settings_json.plain(value)¶
The value as something JSON can hold, or its repr.
spacr/utils.py:1635
- annotate_conditions._get_type(val)¶
Determine if a value maps to ‘rowID’ or ‘columnID’.
spacr/utils.py:3483
- annotate_conditions._map_or_default(column_name, values, loc, df)¶
Assign or map
valuesintocolumn_namebased on optional row/columnloc.spacr/utils.py:3491
- annotate_predictions.assign_condition(row)¶
Return the condition label (
'screen'/'pc'/'nc'or'') for a metadata row.spacr/utils.py:5246
- build_loss._asl(logits, y, gpos, gneg, clip)¶
Return mean asymmetric multilabel loss for logits and targets.
spacr/utils.py:5059
- build_loss._auto_choice() str¶
Return the default loss name from class count and imbalance.
spacr/utils.py:5071
- build_loss._focal_bce(logits, y, alpha, gamma)¶
Return mean focal binary cross-entropy for
logitsandy.spacr/utils.py:5032
- build_loss._focal_ce(logits, y_idx, alpha, gamma)¶
Return mean focal cross-entropy for logits and class indices.
spacr/utils.py:5042
- build_loss._infer_indices(target: torch.Tensor, C: int) torch.Tensor¶
Return class indices from an index vector or 2-D target matrix.
spacr/utils.py:5015
- build_loss.loss_fn(logits, target)¶
Closure: compute the selected per-batch loss from
(logits, target).spacr/utils.py:5087spacr/utils.py:5092spacr/utils.py:5101spacr/utils.py:5107spacr/utils.py:5115spacr/utils.py:5125spacr/utils.py:5133spacr/utils.py:5139
- calculate_loss._focal_bce_with_logits(logits, y, alpha=1.0, gamma=2.0, reduction='mean')¶
Return focal binary cross-entropy for
logitsand targetsy.spacr/utils.py:4516
- calculate_loss._focal_cross_entropy(logits, y_idx, alpha=1.0, gamma=2.0, reduction='mean')¶
Return focal cross-entropy for logits and class indices
y_idx.spacr/utils.py:4528
- class_visualization.blur_image(img, sigma=1)¶
In-place Gaussian blur of each channel of
imgwith standard deviationsigma.spacr/utils.py:6974
- class_visualization.deprocess(img_tensor)¶
Undo ImageNet normalization and return an
(H, W, 3)numpy image in[0, 1].spacr/utils.py:6981
- class_visualization.jitter(img, ox, oy)¶
Return
imgshifted (rolled) byoxandoypixels along the spatial axes.spacr/utils.py:6970
- debug.decorator(func)¶
Inner decorator that binds the logger for
funcand returns the wrapper.spacr/utils.py:1033
- debug.decorator.wrapper(*args, **kwargs)¶
Temporarily bump the logger to DEBUG while
funcruns, then restore its level.spacr/utils.py:1038
- feature_folder_name._one(member)¶
Return one filesystem-safe selection-member slug.
spacr/utils.py:9546
- group_feature_class.find_feature_class(feature, compartments)¶
Return the group label(s) matched in
feature— joined with ‘-’ when more than one hits.spacr/utils.py:10128
- load_settings.parse_value(value)¶
Parse the string value into the appropriate Python data type.
spacr/utils.py:1396
- merge_regression_res_with_metadata.extract_and_clean_gene(feature)¶
Return the gene ID parsed from a
featurestring likeC(gene)[T.<id>_...], orNone.spacr/utils.py:9458
- pick_best_model.sort_key(x)¶
Return
(role, accuracy, epoch)from metadata or legacy name.spacr/utils.py:4585
- plot_grid.cluster_color(cluster_label)¶
Resolve one cluster’s color once for panels and legend.
spacr/utils.py:7987
- pretty_print_settings._fmt(v)¶
Return
vas text truncated to the table’s value width.spacr/utils.py:1532
- pretty_print_settings._row(k, v)¶
Return one padded key-and-formatted-value table row.
spacr/utils.py:1537
- pretty_print_settings._say(line)¶
Print a console-safe form of
lineand returnNone.spacr/utils.py:1528
- process_masks.cluster_objects(properties, n_clusters=2)¶
Return a fitted
KMeansobject clustering the property dicts inton_clustersgroups.spacr/utils.py:9386
- process_masks.measure_morphology_and_intensity(mask, image)¶
Return a list of dicts with area/mean_intensity/perimeter/eccentricity per labeled region.
spacr/utils.py:9380
- process_masks.plot_clusters(properties, labels)¶
Show a 2-D PCA scatter of the property vectors colored by cluster label.
spacr/utils.py:9400
- process_masks.read_files_in_batches(folder, batch_size=50)¶
Yield sorted lists of
.npyfilenames fromfolderin chunks ofbatch_size.spacr/utils.py:9373
- process_masks.remove_objects_not_in_largest_cluster(mask, labels, largest_cluster_label)¶
Return
maskwith all labeled regions removed except those inlargest_cluster_label.spacr/utils.py:9392
- suggest_training_changes._find_csv(root, hint)¶
Return the lexically last matching CSV in
root, if any.spacr/utils.py:4748
- suggest_training_changes._last_seq(series, k)¶
Return at most the final
kvalues as a floating-point array.spacr/utils.py:4789
- suggest_training_changes._normalize_cols(df)¶
Return
dfwith normalized, aliased, first-occurrence columns.spacr/utils.py:4753
- suggest_training_changes._poly_slope(y)¶
Return the finite linear slope of
y, or zero when undefined.spacr/utils.py:4778
- suggest_training_changes._scalar(val)¶
Ensure a single float even if a Series sneaks through.
The Series branch is currently unreachable: every call site passes
<Series>.iloc[<int>], which yields a numpy scalar. It is kept as a deliberate guard because the label-based.loclookups this function used to rely on returned a Series whenever the progress CSV had a duplicated index, and that is an easy regression to reintroduce.spacr/utils.py:4735