spacr.io

Image, dataset, and SQLite input/output helpers used across spaCR.

Exceptions

AcquisitionMetadataConflictError

Object tables disagree about the acquisition used to measure a row.

CropModeMismatch

png_list holds no crops of the object this join is anchored on.

JoinFanOut

A left join returned more rows than the frame it started from.

MergeCardinalityError

A database merge violated its declared relationship.

TimelapseKeyMismatch

One side of the png_list join carries a timepoint and the other does not.

Classes

CombineLoaders

Randomized interleaving of multiple live DataLoaders.

CombinedDataset

Concatenation of multiple Dataset objects behind a single index space.

LazyCropPNG

A PNG-shaped byte stream that is only produced when something opens it.

NoClassDataset

Flat directory of unlabelled images returned alongside their file paths.

TarImageDataset

Image dataset backed by a tar archive, decoded on demand.

spacrDataLoader

DataLoader that pre-fetches batches into a queue on a background thread.

spacrDataset

Image classification dataset that reads class subfolders under data_dir.

Functions

apply_augmentation(image, method)

Return image transformed by a named geometric augmentation.

class_sampling_weights(counts, mode)

Per-class sampling weight for a WeightedRandomSampler.

concatenate_and_normalize(src, channels[, save_dtype, ...])

Concatenate, optionally correct, and normalise V1 field arrays.

convert_numpy_to_tiff(folder_path[, limit])

Convert every .npy array in folder_path to a TIFF under folder_path/tiff.

convert_separate_files_to_yokogawa(folder, regex)

Rename per-slice TIFFs in folder into the Yokogawa CV filename convention.

convert_to_yokogawa(folder)

Convert every image in folder to Yokogawa-style naming with a MIP.

crop_object_type(png_type[, default])

Return the object type named by a png_type / file_metadata string.

crop_png_name(file_name, object_type, object_label[, ...])

Return the file name spacr.utils._generate_names() gives this crop.

crop_refs_for_rows(source, df[, object_type, name_column])

Return one LazyCropPNG per row of df.

crop_rows_from_object_table(db_path[, object_type, ...])

Return one crop row per object, straight off the measurement table.

dataset_filenames(dataset)

Return the source filename of every sample in dataset.

dataset_labels(dataset)

Return the integer class label of every sample in dataset.

delete_empty_subdirectories(folder_path)

Recursively delete every empty subdirectory under folder_path.

display(*args, **kwargs)

Do nothing: IPython is unavailable, so there is nowhere to display to.

expected_sampled_fractions(counts, mode)

Class frequencies the loader is expected to realise under mode.

format_class_balance_report(summary[, class_balance, ...])

Render the human-readable skew report printed on every training run.

generate_cellpose_train_test(src[, test_split])

Split image/mask pairs in src into train and test sibling folders.

generate_cv_loaders(src, n_splits[, mode, image_size, ...])

Build one (train_loader, val_loader) pair per cross-validation fold.

generate_dataset([settings])

Pack per-object PNGs referenced by one or more measurements.db files into a single tar for inference or upload.

generate_dataset_from_lists(dst, class_data, classes)

Put the crops listed per class into dst/train/<class> and dst/test/<class>.

generate_loaders(src[, mode, image_size, batch_size, ...])

Build spacrDataLoader objects for training, validation, or testing.

generate_training_dataset(settings)

Build a balanced training/testing dataset from one of:

load_images_from_paths(images_by_key)

Load images grouped by key into NumPy arrays.

make_class_balance_sampler(labels, mode[, ...])

Build the WeightedRandomSampler for a class-balance mode.

make_cv_folds(labels, n_splits[, groups, seed])

Split indices into n_splits class-stratified, optionally grouped folds.

make_validation_holdout(labels, validation_fraction, ...)

Choose one stratified group fold closest to a requested holdout size.

mark_crop_output_folder(folder[, fmt, source_folder, ...])

Stamp a folder spaCR has just filled with crop PNGs.

migrate_unescaped_plate_names(src[, dry_run])

Escape the plate component of arrays a previous release wrote raw.

open_crop_source(settings[, src, object_type, verbose])

Return the spacr.crops.CropSource a run should read crops from.

parse_gz_files(folder_path)

Group .fastq.gz files in folder_path by sample name and read direction.

prepare_cellpose_dataset(input_root[, augment_data, ...])

Aggregate image/mask pairs from sibling dataset folders into a Cellpose training split.

preprocess_img_data(settings)

Convert raw microscopy images into normalized, channel-merged .npy stacks ready for mask generation.

process_instruction(entry)

Copy one image/mask pair described by entry, applying an optional augmentation.

process_non_tif_non_2D_images(folder)

Split multi-dimensional or non-TIFF images in folder into per-channel TIFFs.

read_plot_model_stats(train_file_path, val_file_path)

Plot training vs. validation curves from a saved model's per-epoch CSVs.

read_settings_history(db_path)

Every settings snapshot ever written into db_path, oldest first.

report_class_balance(labels[, classes, class_balance, ...])

Measure class skew, decide what was done about it, and say so out loud.

report_cv_folds(labels, folds[, classes, groups, ...])

Print the fold table and every warning the split earned.

save_object_mask(output_folder, filename, mask[, ...])

Save an integer label mask as a lossless, compressed TIFF.

select_fields(names, fields)

Keep only the names whose field is in fields.

summarize_class_imbalance(labels[, classes])

Measure the class skew of a label vector.

summarize_cv_folds(labels, folds[, classes, groups])

Tabulate fold sizes and per-class validation counts.

training_dataset_from_annotation(db_path, dst[, ...])

Group per-object PNG paths by manual annotation values so they can be turned into a CNN training set.

training_dataset_from_annotation_metadata(db_path, dst)

Same as training_dataset_from_annotation() but pre-filtered by plate metadata.

Module Contents

exception spacr.io.AcquisitionMetadataConflictError[source]

Bases: ValueError

Object tables disagree about the acquisition used to measure a row.

Merging measurements with different dimensionality, units, z-depth, or voxel sizes would leave their numerical features without one physical interpretation. The source tables must be repaired or the caller must explicitly select which table’s stamp is authoritative.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.io.CropModeMismatch[source]

Bases: ValueError

png_list holds no crops of the object this join is anchored on.

measure_crop writes one object-id column per crop_mode – cell_id for crop_mode=['cell'], nucleus_id for ['nucleus'] and so on; the mapping is spacr.utils.PNG_OBJECT_ID_COLUMNS. _read_and_join_tables() anchors on the cell table, so it needs cell_id. A database measured only with nucleus crops does not have that column, and used to fail with KeyError: "['cell_id'] not in index" – which names neither the table, nor the column’s absence, nor the setting that caused it.

The cell a nucleus crop belongs to is in the crop’s file name – spacr.utils._generate_names() writes <field>_<cell>_<nucleus>.png – but it is not stored: spacr.utils.filepaths_to_database() keeps only the last token. Recovering it would mean a second file-name parser beside spacr.schema, so this is a refusal rather than a guess.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.io.JoinFanOut[source]

Bases: ValueError

A left join returned more rows than the frame it started from.

The object tables carry one row per object per field per timepoint and png_list carries one crop per the same key, so the join is one-to-one (with unmatched rows permitted by the left join) and the row count cannot grow. If it did, a join key is duplicated and every downstream measurement is multiplied.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.io.MergeCardinalityError[source]

Bases: JoinFanOut

A database merge violated its declared relationship.

This is a JoinFanOut for backwards compatibility with callers that already handle duplicate crop rows, but it is used for every explicitly validated object-table relationship.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.io.TimelapseKeyMismatch[source]

Bases: ValueError

One side of the png_list join carries a timepoint and the other does not.

A timepoint column is written only by a timelapse run — by spacr.utils.filepaths_to_database() onto png_list and by spacr.utils._merge_and_save_to_database() onto every object table — so a database where one of the two has it and the other does not was written by two runs that disagreed about whether the experiment was a timelapse. Joining them without the timepoint silently multiplies every object by the number of frames, which is precisely the failure this exception exists to stop.

Initialize self. See help(type(self)) for accurate signature.

class spacr.io.CombineLoaders(train_loaders)[source]

Randomized interleaving of multiple live DataLoaders.

Each step shuffles the loaders that have not been exhausted, probes them in that random order, and yields the first available (loader_index, batch) pair. Exhausted loaders are removed, so every batch from every input loader is yielded once even when the loaders have different lengths.

Parameters:

train_loaders – DataLoaders to combine.

Raises:

StopIteration – when every wrapped loader is exhausted.

Store loaders and initialise per-loader iterators.

__iter__()[source]

Return self — this object is its own iterator.

__next__()[source]

Return (loader_index, batch) from a randomly-chosen live loader.

class spacr.io.CombinedDataset(datasets, shuffle=True)[source]

Bases: torch.utils.data.Dataset

Concatenation of multiple Dataset objects behind a single index space.

Parameters:
  • datasets – Datasets to concatenate; their samples must be index-compatible.

  • shuffle – If True, index lookups are permuted once at construction time. Default True.

Precompute per-dataset lengths and optionally shuffle indices.

__getitem__(index)[source]

Return the sample at index from the appropriate sub-dataset.

__len__()[source]

Return the total number of samples across all sub-datasets.

class spacr.io.LazyCropPNG(source, row, name='')[source]

A PNG-shaped byte stream that is only produced when something opens it.

spacr.utils.plot_umap_images and spacr.utils.plot_clusters_grid reach for every thumbnail with PIL.Image.open(image_paths[i]). Image.open accepts a path or any seekable binary stream, so an instance of this class can sit in that list exactly where a path string used to and the plotting code needs no change – which is the point: the PNG list and the on-demand list have to be interchangeable, or the two sources are not really alternatives.

Nothing is read until something opens the object, so building one per row of a large screen costs a small dict and only the handful of thumbnails actually drawn ever touch merged/.

The bytes are always a current-format (RGB) crop PNG: the array comes from CropSource.get, which is crops.png_view for the merged source and crops.read_crop_png for the PNG one, and both of those return the corrected order. A legacy folder is therefore corrected on the way through here, the same way the Annotate screen corrects it.

Parameters:
  • source – the spacr.crops.CropSource to cut/read with.

  • row – the row mapping identifying the object.

  • name – the crop’s file name, for messages and tar members.

Record the source and row; produce nothing yet.

__enter__()[source]

Return self, so the object can be used as a context manager.

__exit__(*exc)[source]

Release the materialised bytes.

__repr__()[source]

Return a short description naming the crop.

array()[source]

Return the crop as an (H, W, 3) uint8 RGB array.

close()[source]

Drop the materialised bytes; a later read produces them again.

png_bytes()[source]

Return the crop encoded as a current-format (RGB) PNG.

read(size=-1)[source]

Read up to size bytes of the PNG.

The first read is what actually cuts and encodes the crop; every later one is served from the same buffer and continues where the previous read stopped, so re-reading the image needs an explicit seek(0).

Parameters:

size – Byte count. The default -1 (or any negative value) reads to the end of the PNG, which is what PIL does when it is handed this object as a file.

Returns:

The bytes read, empty once the buffer is exhausted.

readable()[source]

Return True – the stream is readable.

seek(offset, whence=0)[source]

Seek within the PNG.

Seeking materialises the crop if it has not been produced yet, so the cost of the first seek is the cost of encoding the whole image, not of moving a cursor.

Parameters:
  • offset – How far to move, in bytes, and it must be an integer — a float raises TypeError. Seeking past the end of the encoded PNG is allowed rather than an error: the position lands beyond the buffer and the next read returns empty.

  • whence – 0 from the start (the default), 1 from the current position, 2 from the end. Semantics are the underlying BytesIO ones, so a negative offset is a ValueError under 0 but legal under 1 and 2.

Returns:

The new absolute position.

seekable()[source]

Return True – the stream is seekable.

tell()[source]

Return the current offset.

writable()[source]

Return False – the stream is read-only.

property closed[source]

Return False – this object is never permanently closed.

class spacr.io.NoClassDataset(data_dir, transform=None, shuffle=True, load_to_memory=False, *, crop_loading_policy=DECLARED_UINT8)[source]

Bases: torch.utils.data.Dataset

Flat directory of unlabelled images returned alongside their file paths.

Parameters:
  • data_dir – Directory containing image files.

  • transform – Optional callable applied to each PIL image. If None, images are converted with ToTensor.

  • shuffle – If True, shuffle filename list at construction. Default True.

  • load_to_memory – If True, decode all images once and hold them in RAM. Default False.

  • crop_loading_policy – declared_uint8_v1 uses shared crop decoding; stored_pil_v1 preserves untagged historical checkpoints.

Enumerate files in data_dir and optionally preload them.

__getitem__(index)[source]

Return (image_tensor, filename) for the given index.

Parameters:

index – Position within the dataset.

Returns:

(tensor, path) where tensor is the transformed image and path is the source filename.

__len__()[source]

Return the number of images in the dataset.

load_image(img_path)[source]

Return the image at img_path decoded as RGB.

Parameters:

img_path – Path to the image file.

Returns:

PIL Image in RGB mode.

shuffle_dataset()[source]

Shuffle the internal filename list in place.

class spacr.io.TarImageDataset(tar_path, transform=None, *, crop_loading_policy=DECLARED_UINT8)[source]

Bases: torch.utils.data.Dataset

Image dataset backed by a tar archive, decoded on demand.

A tar written by generate_dataset() from on-demand crops carries a .spacr_crop_format.json member – the same marker spacr.crops writes into a crop folder, travelling with the bytes it describes. It is not an image, so it is excluded from the sample list and surfaced as crop_format instead; an archive without one reports None, which is every tar written before this existed.

New datasets default to declared channel order and uint8 narrowing through the shared crop decoder. Inference must pass the model’s recorded policy; untagged older checkpoints use stored_pil_v1 to preserve their pixels.

Parameters:
  • tar_path – Path to the tar archive.

  • transform – Optional callable applied to each PIL image.

  • crop_loading_policy – declared_uint8_v1 (default), or stored_pil_v1 for historical model inputs.

Enumerate archive members without extracting.

__getitem__(idx)[source]

Return (image, member_name) extracted from the tar at idx.

__len__()[source]

Return the number of image members in the archive.

class spacr.io.spacrDataLoader(*args, preload_batches=1, **kwargs)[source]

Bases: torch.utils.data.DataLoader

DataLoader that pre-fetches batches into a queue on a background thread.

Wraps torch.utils.data.DataLoader and runs a daemon thread that stays one or more batches ahead of consumption to hide I/O latency. End-of-stream is signalled with a sentinel, so the full batch stream is always delivered.

Parameters:

preload_batches – Number of batches to keep queued ahead. Default 1.

Initialise the underlying DataLoader and the preload queue.

__del__()[source]

Ensure background resources are released on garbage collection.

__iter__()[source]

Start a fresh pass over the data and return self.

Safe to call more than once (list(iter(dl)) calls it twice): any in-flight producer is stopped first, so the stream is never doubled.

__next__()[source]

Return the next queued batch, or raise StopIteration at the sentinel the preloader pushes when the stream is exhausted.

cleanup()[source]

Signal the preloader to stop and join the background thread.

class spacr.io.spacrDataset(data_dir, loader_classes, transform=None, shuffle=True, pin_memory=False, specific_files=None, specific_labels=None, *, crop_loading_policy=DECLARED_UINT8)[source]

Bases: torch.utils.data.Dataset

Image classification dataset that reads class subfolders under data_dir.

Parameters:
  • data_dir – Root directory containing one subdirectory per class.

  • loader_classes – Ordered list of class names — the index in this list becomes the integer label.

  • transform – Optional callable applied to each PIL image.

  • shuffle – If True, shuffle files+labels at construction.

  • pin_memory – If True, eagerly load every image into RAM via a multiprocessing pool.

  • specific_files – Optional explicit list of image paths. If supplied together with specific_labels, directory scanning is skipped.

  • specific_labels – Labels paired with specific_files.

  • crop_loading_policy – declared_uint8_v1 uses shared crop decoding; stored_pil_v1 preserves untagged historical checkpoints.

Raises:

ValueError – If no non-hidden image files are found for any requested class.

Build the filename/label lists and optionally preload images.

__getitem__(index)[source]

Return (image, label, filename) for the given index.

__len__()[source]

Return the number of samples in the dataset.

get_plate(filepath)[source]

Return the plate identifier parsed from a filename (leading token before _).

Parameters:

filepath – Image path.

Returns:

Plate ID string.

load_image(img_path)[source]

Return the image at img_path decoded as RGB with EXIF orientation applied.

Parameters:

img_path – Path to one crop, taken from self.filenames. The file is fully decoded and copied before the handle closes, so the returned image owns its buffer and is safe to hand to a prefetch thread. Anything PIL can open works: a greyscale crop is widened to three channels and an RGBA one loses its alpha, which is what makes every sample the same shape for the transform.

Returns:

A PIL.Image.Image in mode RGB.

shuffle_dataset()[source]

Jointly shuffle filenames and labels in place.

spacr.io.apply_augmentation(image, method)[source]

Return image transformed by a named geometric augmentation.

Parameters:
  • image – NumPy image array.

  • method – One of 'rotate90', 'rotate180', 'rotate270', 'flip_h', 'flip_v'; any other value returns the input unchanged.

Returns:

Augmented image array.

spacr.io.class_sampling_weights(counts, mode)[source]

Per-class sampling weight for a WeightedRandomSampler.

'weighted_sampler' uses 1/n_c, which makes every class equally likely to be drawn. 'sqrt_weighted_sampler' uses 1/sqrt(n_c), a partial correction that moves the realised frequencies toward balance without oversampling a tiny class so hard that the model memorises its handful of crops.

Parameters:
  • counts – per-class sample counts.

  • mode – 'weighted_sampler' or 'sqrt_weighted_sampler'.

Returns:

list of per-class weights, scaled so they sum to 1.

Raises:

ValueError – if mode does not describe a sampler.

spacr.io.concatenate_and_normalize(src, channels, save_dtype=np.float32, settings=None, illumination_session=None, psf_session=None)[source]

Concatenate, optionally correct, and normalise V1 field arrays.

Parameters:
  • src – directory containing per-field .npy channel arrays.

  • channels – channel indices retained in the output archives.

  • save_dtype – NumPy dtype for normalised arrays. Defaults to float32.

  • settings – required preprocessing settings mapping.

  • illumination_session – optional segmentation-only correction session. Corrected archives are staged privately and published as one complete set; the staging directory is removed on success or failure.

  • psf_session – optional PSF session from spacr.psf_pipeline. Uses the same complete-set publication; cancellation leaves its provenance incomplete and prevents reuse of partially processed data.

Returns:

the masks/ directory containing normalised NPZ archives.

spacr.io.convert_numpy_to_tiff(folder_path, limit=None)[source]

Convert every .npy array in folder_path to a TIFF under folder_path/tiff.

Parameters:
  • folder_path – Folder containing the .npy files.

  • limit – If set, stop after processing this many files.

Returns:

None

spacr.io.convert_separate_files_to_yokogawa(folder, regex)[source]

Rename per-slice TIFFs in folder into the Yokogawa CV filename convention.

Files are grouped by (plateID, wellID, fieldID, timeID, chanID) parsed from the regex. Groups with multiple Z-slices are max- projected before saving, and the mapping is logged to rename_log.csv.

Well naming, in full:

  1. A wellID that is a well address keeps it — spacr.convert.normalise_well() reads a1, A-01, Q01 (row 17) and AA13 (row 27) alike. Every one of those used to be thrown away and replaced with the next free synthetic id, so a real 1536-plate came out relabelled A01, A02, … with only rename_log.csv to say what had happened.

  2. Anything else — 1, well_left, a positional number — is handed a synthetic id, in _natural_key order so the same folder always converts the same way. It used to follow os.listdir order, which is the filesystem’s business and not reproducible.

  3. Each distinct source plateID gets its own plate<N> token, so well A01 of two source plates stays two wells.

Parameters:
  • folder – Folder containing the source TIFFs.

  • regex – Pattern with named groups wellID (required) plus optional plateID, fieldID, timeID, chanID, sliceID.

Returns:

None

Raises:

ValueError – when a file the regex matches carries a fieldID, timeID, chanID or sliceID that is not a whole number (before anything is written), or cannot be read and converted. The message names the file, and in the second case says how many converted files were written before the conversion stopped; rename_log.csv is written only when every file converted.

spacr.io.convert_to_yokogawa(folder)[source]

Convert every image in folder to Yokogawa-style naming with a MIP.

ND2, CZI, LIF and plain TIFF/PNG/JPEG inputs are detected by extension, max-projected over Z, and written out as plate<N>_<well>_T####F###L01C##.tif. A rename_log.csv records the original-to-new mapping.

A file that cannot be read is skipped so the rest of the folder still converts — but the skip is recorded on a spacr.errors.RunLedger, printed as a loud summary at the end, and stamped into a sibling rename_log.run_status.json. That sidecar is what lets a later reader (or spacr.errors.run_is_complete()) tell that the converted folder is missing inputs, instead of quietly analysing a subset.

Parameters:

folder – Directory of raw images, converted in place.

Returns:

the spacr.errors.RunLedger for the conversion.

Raises:

ValueError – If the folder already contains Yokogawa-named converted images or a previous rename_log.csv. The check runs before writing any image or log, including after the converted images have been moved into orig/. Read an already converted folder with metadata_type='cellvoyager', or retry raw conversion in a separate folder containing only the original inputs.

spacr.io.crop_object_type(png_type, default='cell')[source]

Return the object type named by a png_type / file_metadata string.

'cell_png' -> 'cell', '…/nucleus_png/…' -> 'nucleus'. A string that names no object type (a plate prefix, say) falls back to default, because that filter is about which rows, not which mask.

Parameters:
  • png_type – the setting value, or any path/substring containing it.

  • default – object type to assume when nothing is named.

Returns:

one of CROP_OBJECT_TYPES.

spacr.io.crop_png_name(file_name, object_type, object_label, cell_id=None)[source]

Return the file name spacr.utils._generate_names() gives this crop.

The name matters downstream: _png_group_id() parses the plate / well / field out of it for group-aware cross-validation, and spacr.utils.process_vision_results() parses the prcfo that spacr.deep_spacr.merge_predictions_into_db() merges on. A crop cut on demand has to carry the same name as the one the PNG folder would have held or those two stop lining up.

Parameters:
  • file_name – the merged array’s stem (plate1_A01_1).

  • object_type – which crop mode this is.

  • object_label – the object’s integer label.

  • cell_id – the parent cell label, for nucleus/pathogen crops.

Returns:

the crop’s file name, ending in .png.

spacr.io.crop_refs_for_rows(source, df, object_type='cell', name_column=None)[source]

Return one LazyCropPNG per row of df.

Parameters:
  • source – the crop source to cut/read with.

  • df – rows carrying whatever the source needs (png_path for the PNG source, path_name + object_label for the merged one).

  • object_type – object type stamped onto each row.

  • name_column – column holding the crop’s file name; defaults to the basename of png_path.

Returns:

list of LazyCropPNG.

spacr.io.crop_rows_from_object_table(db_path, object_type='cell', verbose=True)[source]

Return one crop row per object, straight off the measurement table.

This is the path for a project that never wrote a PNG folder at all, so there is no png_list to start from: object_label, path_name and the well keys are already on every measurement row.

Parameters:
  • db_path – the measurements.db.

  • object_type – which object table to read.

  • verbose – print what was found.

Returns:

a DataFrame with path_name, object_label, the well keys, png_name and png_path (the path the crop would have had).

spacr.io.dataset_filenames(dataset)[source]

Return the source filename of every sample in dataset.

Mirrors dataset_labels() so group ids can be derived without touching pixels.

Parameters:

dataset – dataset, Subset, or sequence of (img, label, name).

Returns:

list of filename strings, positionally aligned with the dataset.

spacr.io.dataset_labels(dataset)[source]

Return the integer class label of every sample in dataset.

Handles the three shapes that flow through the training path: a spacrDataset (labels are already a list), a torch.utils.data.Subset of one (produced by random_split and by the fold splitter), and the plain list of (image, label, filename) tuples that augment_dataset returns. Only the last shape has to be walked, and it holds tensors in memory already, so nothing here decodes an image.

Parameters:

dataset – dataset, Subset, or sequence of (img, label, name).

Returns:

list of int labels, positionally aligned with the dataset.

spacr.io.delete_empty_subdirectories(folder_path)[source]

Recursively delete every empty subdirectory under folder_path.

Parameters:

folder_path – Root directory to scan.

Returns:

None

spacr.io.display(*args, **kwargs)[source]

Do nothing: IPython is unavailable, so there is nowhere to display to.

THE FALLBACK IS THE POINT. IPython.display.display is imported at module scope, and IPython can be mid-init – partially imported by another thread – which makes that import raise. Letting it propagate would make importing this module fail for a reason that has nothing to do with what the module does. spaCR only calls display from notebook contexts; the Qt GUI ignores it.

Parameters:
  • args – whatever the caller would have displayed.

  • kwargs – likewise.

spacr.io.expected_sampled_fractions(counts, mode)[source]

Class frequencies the loader is expected to realise under mode.

This is what makes the effect visible before a single epoch runs: the report prints the observed fractions next to these.

Parameters:
  • counts – per-class sample counts.

  • mode – any value of CLASS_BALANCE_MODES.

Returns:

list of expected per-class draw probabilities.

spacr.io.format_class_balance_report(summary, class_balance='none', split_name='train')[source]

Render the human-readable skew report printed on every training run.

Parameters:
  • summary – dict from summarize_class_imbalance().

  • class_balance – the mode that was requested.

  • split_name – which split is being described, e.g. 'train'.

Returns:

multi-line report string.

spacr.io.generate_cellpose_train_test(src, test_split=0.1)[source]

Split image/mask pairs in src into train and test sibling folders.

Only images that have a corresponding mask in src/masks are considered.

Parameters:
  • src – Folder containing images and a masks subfolder.

  • test_split – Fraction of pairs to route into the test set. Default 0.1.

Returns:

None

spacr.io.generate_cv_loaders(src, n_splits, mode='train', image_size=224, batch_size=32, classes=None, n_jobs=None, pin_memory=False, normalize=False, channels=None, augment=False, verbose=False, group_by='well', class_balance='none', seed=0, crop_loading_policy=DECLARED_UINT8)[source]

Build one (train_loader, val_loader) pair per cross-validation fold.

The dataset under src/<mode> is read once and then re-split k ways, so every crop is used for validation exactly once. Folds are class-stratified and, by default, grouped by well so that crops from the same well stay on one side of the split. Class balancing is applied to the fold’s train loader only.

Parameters:
  • src – dataset root containing train/test subfolders.

  • n_splits – number of folds, must be >= 2.

  • mode – which split to fold — normally 'train'.

  • image_size – square resize target in pixels.

  • batch_size – loader batch size.

  • classes – ordered class names matching the subfolder names.

  • n_jobs – DataLoader worker count.

  • pin_memory – if True, pin batches to page-locked memory.

  • normalize – if True, apply per-channel normalisation.

  • channels – subset of RGB channels to keep.

  • augment – if True, 8-fold augment each fold’s train split.

  • verbose – log configuration to stdout.

  • group_by – fold grouping level, one of CV_GROUP_LEVELS.

  • class_balance – one of CLASS_BALANCE_MODES, train loaders only.

  • seed – seed for the deterministic fold assignment.

  • crop_loading_policy – crop decoding policy recorded on each loader; defaults to declared channel order and high-byte uint8 narrowing.

Returns:

(fold_loaders, info) where fold_loaders is a list of (train_loader, val_loader) and info holds fold_table, warnings, imbalance and groups.

Raises:

ValueError – if n_splits < 2 or a setting value is unknown.

spacr.io.generate_dataset(settings=None)[source]

Pack per-object PNGs referenced by one or more measurements.db files into a single tar for inference or upload.

Selects PNG paths (via the png_list table plus optional file_metadata filter) from each source’s measurements database, optionally random-subsamples, then bundles the images in parallel into a dated tar under the first source’s datasets/ folder. Use this to produce the tar_path consumed by spacr.deep_spacr.deep_spacr() / apply_model_to_tar.

crop_source chooses where the images come from. 'png' (and 'auto' wherever a crop folder exists) is the behaviour above, unchanged for uniform source formats: files are byte-copied into the tar with their format marker. Mixed formats are decoded into declared uint8 copies in the archive only. 'merged' (and 'auto' on a project with no crop folder) cuts every crop out of merged/*.npy through spacr.crops instead, so the tar can be built with no PNG folder on disk at all, and is rebuilt at the current crop settings rather than whatever the folder was generated with. The members are named exactly as the PNG folder would have named them, so everything that parses a crop file name downstream – fold grouping, prcfo, the prediction merge – keeps working either way.

Parameters:

settings –

Settings dict, canonicalized via spacr.settings.set_generate_dataset_defaults(). Key entries:

  • src (str or list of str) — folder(s) containing measurements/measurements.db and the PNG crops.

  • file_metadata — filter/join key applied against png_list.

  • sample — int or [int] cap on selected PNGs (random subsample); omit for all.

  • experiment — string suffix used in the tar filename.

  • crop_source — 'auto' | 'png' | 'merged'.

Returns:

Absolute path to the created …/datasets/<date>_< experiment>.tar.

Raises:

RuntimeError – if src is not a string / list of strings, no images are selected, no image could be written, or the destination folder cannot be resolved.

The tar-writing pool is closed and joined, not left to the with block’s terminate(): that sends SIGTERM to idle workers, and a worker whose SIGTERM handler needs a lock the interrupted code holds (coverage’s sigterm = True data save does) never exits, so the shutdown waited on it for ever. Workers that finished their tasks are let go by a normal exit instead.

Example

from spacr.io import generate_dataset
tar_path = generate_dataset({
    'src': ['/data/plate01', '/data/plate02'],
    'experiment': 'screen_v1',
    'sample': 100000,
})

See also

training_dataset_from_annotation() — build a labeled train/ / test/ tree instead of a flat tar. spacr.deep_spacr.deep_spacr() — consumes the tar via apply_model_to_tar.

spacr.io.generate_dataset_from_lists(dst, class_data, classes, test_split=0.1, db_path=None, random_seed=42, group_by='well')[source]

Put the crops listed per class into dst/train/<class> and dst/test/<class>.

An entry may be a path, which is copied byte for byte exactly as before, or a LazyCropPNG, which is cut out of merged/*.npy through spacr.crops and written as a current-format (RGB) crop PNG. The two are interchangeable, so a training set can be built with no crop folder on disk at all.

Each destination class folder is stamped before it is filled. Uniform source formats keep their original bytes and marker. Mixed source formats are decoded into declared uint8 crops in the destination only, so no generated folder silently loses its channel-order record.

Parameters:
  • dst – Output root; train and test subfolders are created.

  • class_data – Sequence of per-class lists of paths and/or LazyCropPNG handles.

  • classes – Class names paired positionally with class_data.

  • test_split – Fraction of each class routed to test/. Default 0.1.

  • db_path – optional measurements.db consulted for the crop format of a source folder that carries no sidecar.

  • random_seed – Reproducible global train/test split seed.

  • group_by – acquisition identity kept intact across the permanent train/test boundary. Default well. cell is the leakiest per-object choice; legacy none aliases it.

Returns:

(train_dir, test_dir) tuple of the top-level split paths.

Raises:

ValueError – if len(class_data) != len(classes).

spacr.io.generate_loaders(src, mode='train', image_size=224, batch_size=32, classes=None, n_jobs=None, validation_split=0.0, pin_memory=False, normalize=False, channels=None, augment=False, verbose=False, class_balance='none', seed=42, group_by='none', crop_loading_policy=DECLARED_UINT8)[source]

Build spacrDataLoader objects for training, validation, or testing.

Reads class subfolders under src/<mode>, applies the requested transforms (channel selection, optional normalisation, optional augmentation) and returns loaders sized to batch_size.

Parameters:
  • src – Root folder containing train/test subfolders.

  • mode – Which split to load — 'train' or 'test'.

  • image_size – Square resize target in pixels. Default 224.

  • batch_size – Loader batch size. Default 32.

  • classes – Ordered class names. Default ['nc', 'pc'].

  • n_jobs – DataLoader worker count. None (the default) means 0 workers, i.e. batches are read in the calling process.

  • validation_split – Fraction of the train split to hold out.

  • pin_memory – If True, pin batches to page-locked memory.

  • normalize – If True, apply per-channel normalisation.

  • channels – Subset of RGB channels to keep, e.g. ['r', 'g'].

  • augment – If True, apply the training augmentation pipeline.

  • verbose – If True, log configuration to stdout.

  • class_balance – One of CLASS_BALANCE_MODES. 'none' (default) leaves sampling untouched; the sampler modes attach a WeightedRandomSampler to the TRAIN loader only. The skew is reported either way.

  • seed – Reproducible train/validation split and loader order.

  • group_by – field, well or plate keeps that acquisition identity entirely on one side of the ordinary validation holdout. none retains the legacy per-object random split.

  • crop_loading_policy – crop decoding policy recorded on each loader; defaults to declared channel order and high-byte uint8 narrowing.

Returns:

For mode='train', a tuple of loaders and a plot handle; for mode='test', the test loader (plus optional metadata).

Raises:

ValueError – if class_balance is not a recognised mode.

spacr.io.generate_training_dataset(settings)[source]

Build a balanced training/testing dataset from one of:

  • metadata rules (exact matches or compound ‘where’ rules)

  • annotation columns (each <col>_<value> is a standalone class)

  • measurement rules (numeric ranges/bins; supports multiple conditions per class)

New behavior (annotation mode):

  • If a column has only one annotated value (e.g., only ‘1’s), we add a ‘<column>_random’ class using unannotated rows for that column (same size as positives).

  • Optional: persist that random selection into DB as a new INT column named ‘<column>_random’ with 1’s.

crop_source chooses where the pixels come from. 'png' (and 'auto' wherever a crop folder exists) copies the pre-generated PNGs, unchanged. 'merged' (and 'auto' with no crop folder) cuts each selected crop out of merged/*.npy through spacr.crops instead: the labels still come from png_list, but the pixels are cut fresh at the current crop settings, so the training set costs no standing disk and cannot be built out of a folder that has gone stale. A project with no png_list at all falls back to the object measurement table, which still carries everything the metadata rules select on.

Parameters:

settings –

Settings dict, first completed by spacr.settings.set_generate_training_dataset_defaults(). Keys read here:

  • src: plate root, or a list of roots merged into one dataset.

  • dataset_mode: 'metadata', 'annotation' or 'measurement'.

  • test_split: fraction of each class routed to test/.

  • cv_group_by: acquisition identity kept intact across the train/test boundary. Default 'well'.

  • path_string (legacy alias png_type): substring a crop path must contain. Default 'cell_png'.

  • crop_source: 'auto', 'png' or 'merged', as above.

  • balance_to_smallest: downsample every class to the smallest one. Default True.

  • random_seed: seeds balancing and random-class sampling. Default 42.

  • tables: object tables the project has. Only clears nuclei_limit and pathogen_limit when the matching table is absent; it does not select where the crop list comes from.

  • metadata mode: metadata_rules, or class_metadata values matched against the column the Classes editor names.

  • annotation mode: annotation_columns (legacy annotation_column), optional annotation_values filter, and write_random_annotation_column.

The resulting class_folder_names and nr_classes are written back into the dict for downstream training. A pre-split list-shaped classes entry is retired only after those folders are written; dict-shaped class definitions remain untouched.

Returns:

(train_class_dir, test_class_dir) — the train/ and test/ roots written under <src>/datasets/training (training_all when several sources are combined), suffixed to stay unique.

Raises:

ValueError – if dataset_mode is unrecognised, a rule names a missing column or unsupported operator, or a requested class selected no crops.

spacr.io.load_images_from_paths(images_by_key)[source]

Load images grouped by key into NumPy arrays.

Parameters:

images_by_key – Mapping of key -> list of image paths.

Returns:

Mapping of the same keys -> list of ndarray images. Paths that fail to load are skipped, recorded on a spacr.errors.RunLedger and reported in a loud summary, so a short list is never mistaken for a complete one.

spacr.io.make_class_balance_sampler(labels, mode, num_samples=None, generator=None)[source]

Build the WeightedRandomSampler for a class-balance mode.

Parameters:
  • labels – integer labels of the split being sampled.

  • mode – any value of CLASS_BALANCE_MODES; the non-sampler modes return (None, None).

  • num_samples – draws per epoch. Defaults to len(labels) so the epoch keeps its usual length.

  • generator – optional torch.Generator for reproducible draws.

Returns:

(sampler, per_sample_weights), or (None, None).

Raises:

ValueError – if mode is not a recognised class-balance mode.

spacr.io.make_cv_folds(labels, n_splits, groups=None, seed=0)[source]

Split indices into n_splits class-stratified, optionally grouped folds.

Every index lands in exactly one validation fold, so the k folds partition the dataset. With groups supplied, a whole group is assigned to a single fold — crops from the same well never straddle the train/val line — and groups are placed greedily into whichever fold currently leaves the per-class proportions most even, which is how stratification survives grouping.

Parameters:
  • labels – integer labels, one per sample.

  • n_splits – number of folds, must be >= 2.

  • groups – optional group id per sample (same length as labels).

  • seed – seed for the shuffle. Both branches use it: ungrouped, it shuffles each class and picks the starting fold; grouped, it orders groups of equal size and breaks ties between folds the greedy pass rates equally, so re-running with a different seed gives a different — and equally stratified — partition. Where only one partition is feasible (as many groups as folds, say) no seed can change it.

Returns:

list of (train_idx, val_idx) numpy integer arrays.

Raises:

ValueError – if n_splits < 2, if groups is the wrong length, or if there are fewer samples/groups than folds.

spacr.io.make_validation_holdout(labels, validation_fraction, groups, seed=0)[source]

Choose one stratified group fold closest to a requested holdout size.

The ordinary Classify validation split used to call random_split and could put crops from one well on both sides even though grouped CV did not. This helper uses the same group-stratified partitioner as CV and selects the candidate fold closest to the requested size and class distribution.

Parameters:
  • labels – One integer class label per sample, in dataset order; the returned indices point back into that same order.

  • validation_fraction – Target share of samples to hold out, strictly between 0 and 1; any number outside that range, nan and inf included, is a ValueError. It is only a target. The holdout is one whole fold of a split into max(2, round(1 / fraction)) folds, itself capped at the number of distinct groups, so the realised share is quantised to whole groups and can miss in either direction, by a lot. Over eight equal groups, 0.05 holds out 0.125 (no finer split is available) and 0.7 holds out 0.5 (the two-fold floor); with one dominant group — 70 of 100 samples across four groups — those same two requests instead hold out 0.10 and 0.70. A miss that big is no longer silent: whenever the realised share is further than max(0.01, 0.1 * validation_fraction) from the requested one, a UserWarning names both numbers and why they differ.

  • groups – Group id per sample — well, field, or whatever cv_group_by names — and required, not optional, because the point of this function is that a group never straddles the split. Needs the same length as labels and at least two distinct values.

  • seed – Seed passed to make_cv_folds() and used to break ties between equally suitable folds. Different seeds can produce different holdouts when several partitions satisfy the constraints. The seed has no effect when the groups permit only one partition.

Returns:

One (train_idx, val_idx) pair of numpy integer arrays.

Raises:

ValueError – if validation_fraction is outside (0, 1), if groups is missing or the wrong length, or if fewer than two distinct groups are present.

Warns UserWarning:

if whole-group quantisation makes the realised holdout share miss validation_fraction by more than the tolerance above.

spacr.io.mark_crop_output_folder(folder, fmt=None, source_folder=None, db_path=None, **extra)[source]

Stamp a folder spaCR has just filled with crop PNGs.

Called before the folder is filled, exactly as spacr.crops.stamp_crop_folder() is on the measure path, so an interrupted run leaves a marked folder holding fewer crops rather than an unmarked folder of corrected ones – the one state that is silently misread.

fmt=None inherits the format from source_folder. Byte-for-byte copies retain their source format: formats 1 and 3 use declared order, while format 2 needs channel reversal when read by a declared-order model.

Parameters:
  • folder – the folder about to be filled.

  • fmt – the format to record; None inherits from source_folder.

  • source_folder – the folder the crops are being copied from.

  • db_path – measurements.db consulted when source_folder carries no sidecar.

  • extra – extra keys recorded in the sidecar.

Returns:

the sidecar path, or None when it could not be written.

spacr.io.migrate_unescaped_plate_names(src, dry_run=False)[source]

Escape the plate component of arrays a previous release wrote raw.

A plate folder whose name holds an underscore – exp_1 – used to produce exp_1_A01_1_1.npy, five separator-delimited components for a four-component grammar, and utils._map_wells answered ('error',) * 5 for every field of it. The plate could not be measured at all.

Nothing in the measurement database needs migrating, because there is none: every frame of such a plate was refused. What DOES need moving is everything upstream of the measurement – stack/, norm_channel_stack/, merged/ and, above all, masks/, which is hours to days of segmentation. Renaming those makes the plate measurable without re-segmenting it, which is the whole reason this exists rather than a note saying “re-run the plate”.

THE NEW NAME IS NOT GUESSED. Every stem this module writes ends in three fixed tokens – well, field, timepoint – whatever timelapse is set to, so everything before them is the plate however many underscores it holds. A stem that is already escaped, or whose plate holds no separator, is left alone: the rename is a no-op for every ordinary plate, which is what makes it safe to run over a folder that does not need it.

Crops under data/ are not touched. A plate that could not be measured has none, and the nightly-only crop names that _generate_names mis-escaped came with a png_list table whose identities are wrong too, so there is nothing there to salvage by renaming – re-run measure_crop.

PUBLIC, because the person who needs it is a user with an exp_1 folder full of masks, and a recovery tool they have to reach past a leading underscore to call is a recovery tool most people will not find.

Parameters:
  • src – the plate source folder, the one holding merged/.

  • dry_run – report the renames without performing them.

Returns:

list of (old_path, new_path) pairs, renamed unless dry_run.

Raises:

FileExistsError – if a destination is already occupied, before anything is moved. A half-applied rename is worse than none.

Example

>>> from spacr.io import migrate_unescaped_plate_names
>>> migrate_unescaped_plate_names('/data/exp_1', dry_run=True)
[('/data/exp_1/merged/exp_1_A01_1_1.npy',
  '/data/exp_1/merged/exp%5F1_A01_1_1.npy')]
spacr.io.open_crop_source(settings, src=None, object_type=None, verbose=True)[source]

Return the spacr.crops.CropSource a run should read crops from.

Thin, non-raising wrapper over spacr.crops.resolve_crop_source(): it reads settings['crop_source'] ('auto' | 'png' | 'merged'), prints which source was chosen and why, and returns None when neither is available – so a caller can fall back to whatever it did before instead of failing on a project that predates merged/.

The run’s own crop-shaping settings are forwarded (see _crop_shape_overrides()), which is what makes “cut fresh at the current crop settings” true rather than a slogan: resolve_crop_source starts from the measure_crop snapshot in measurements.db and lets those override it, so a run that asks for 96 px crops gets 96 px crops out of merged/ even though the folder on disk holds 48 px ones.

Parameters:
  • settings – settings dict (or a source path) holding crop_source.

  • src – the experiment root; defaults to settings['src'] (its first entry when that is a list).

  • object_type – default object type for a merged source.

  • verbose – print the chosen source.

Returns:

a spacr.crops.CropSource, or None.

spacr.io.parse_gz_files(folder_path)[source]

Group .fastq.gz files in folder_path by sample name and read direction.

Accepts both naming conventions in the wild: <sample>_R1_... from an Illumina run, and <run>_1.fastq.gz from ENA or the SRA. See _MATE_SPELLINGS.

A file whose mate cannot be identified contributes NOTHING rather than an empty entry. The previous version created {sample: {}} for it, which turned an unrecognised filename into a KeyError: 'R1' several frames later in spacr.sequencing.generate_barecode_mapping() – a crash that named neither the file nor the problem.

Parameters:

folder_path – Directory containing gzipped FASTQ files.

Returns:

Mapping {sample_name: {"R1": path, "R2": path}}. Samples may have only one of the two.

spacr.io.prepare_cellpose_dataset(input_root, augment_data=False, train_fraction=0.8, n_jobs=None)[source]

Aggregate image/mask pairs from sibling dataset folders into a Cellpose training split.

Discovers <input_root>/*/masks layouts, balances datasets to a common size (with augmentations if requested) and copies the selected pairs into <input_root>/cellpose_dataset/train and .../test.

Parameters:
  • input_root – Directory containing one subfolder per dataset.

  • augment_data – If True, expand under-sized datasets by applying geometric augmentations. Default False.

  • train_fraction – Fraction of pairs routed to the train split. Default 0.8.

  • n_jobs – Worker count for parallel copies. Default: CPU count.

Returns:

None

Raises:

ValueError – if no valid <subdir>/masks datasets are found.

spacr.io.preprocess_img_data(settings)[source]

Convert raw microscopy images into normalized, channel-merged .npy stacks ready for mask generation.

Usually invoked internally by spacr.core.preprocess_generate_masks(), but callable directly when you only want the preprocessing half. By default it converts z-stacks to MIPs, renames files into the Yokogawa/spacr layout, merges per-channel folders into stacked .npy arrays with optional background subtraction and percentile normalization, and (in test_mode) emits example plots.

With z_stack=True and z_segmentation_mode='volumetric', a separate raw TIFF route preserves Z. Each field/channel must be one complete TIFF series explicitly labelled ZYX. Filename metadata identifies the field and channel; ambiguous axes, duplicate channels and missing companions are refused before any stack is written. This route retains the original files even if save_original_images is false, publishes canonical ZYXC stacks with a completion record, and normalises one field at a time. Reuse verifies source and stack hashes and relevant processing settings. An old projected stack cannot be reused as a volume. Individual slice-file layouts, time-series, test-mode sampling and illumination or PSF preprocessing are not supported by this raw volumetric route.

A separate Mask-only route accepts a fixed Convert map with every explicitly labelled YX plane in a dense channel-by-Z-by-time grid when z_stack and t_stack are both on. It builds native TZYXC archives for the existing 4-D segmenter; physical Z and frame spacing and TZYX axis order must be declared. The planar sources remain unchanged. Timelapse tracking, Measure and Classify are not enabled by this route.

Running it again on a plate folder it has already processed resumes rather than starting over. Raw images an earlier run moved into orig/ are read from there, and only the fields stack/ lacks are built. Every stack/*.npy and masks/*.npz an earlier run left is checked before it is reused: a file cut short is renamed to <name>.damaged, named in the log, and built again, a field stack from the raw images or channel folders and an archive from stack/. A field stack with nothing left to build it from is recorded as a failure, so the run ends incomplete instead of quietly short of that field, and an archive with nothing left to build it from stops the run with an error that names it. When no field can be built at all, the error says what the folder does hold. In test_mode a plate whose raw images are gone is sampled from its stack/ instead.

Parameters:

settings –

Preprocessing settings dict, canonicalized via spacr.settings.set_default_settings_preprocess_img_data(). Key entries:

  • src — folder of raw images (.tif/.nd2/.czi/.lif etc.).

  • metadata_type — 'cellvoyager' / 'auto'; drives filename regex.

  • custom_regex — override the built-in regex.

  • cell_channel, nucleus_channel, pathogen_channel, organelle_channel, channels — channel selection.

  • z_stack and z_segmentation_mode select the explicit volumetric TIFF route described above. Other raw layouts retain the existing per-field/channel projection behavior.

  • remove_background_cell / _nucleus / _pathogen / _organelle and each object’s *_background and *_signal_to_noise values. Every organelle slot the run enables uses its own pair, e.g. organelleb_background.

  • normalize, lower_percentile, save_dtype.

  • batch_size, randomize, test_mode, test_images, plot, cmap, figuresize.

Returns:

Tuple (settings, src) — settings with defaults applied and src pointing at the folder containing the generated stack/ / channel_stack/ outputs (the downstream mask stage reads from here).

Example

from spacr.io import preprocess_img_data
settings = {
    'src': '/data/plate01',
    'metadata_type': 'cellvoyager',
    'cell_channel': 0, 'nucleus_channel': 1, 'pathogen_channel': 2,
    'channels': [0, 1, 2, 3], 'normalize': True,
}
settings, src = preprocess_img_data(settings)

See also

spacr.core.preprocess_generate_masks() — full pipeline wrapper that calls this then generates masks.

spacr.io.process_instruction(entry)[source]

Copy one image/mask pair described by entry, applying an optional augmentation.

Parameters:

entry – Dict with keys src_img, src_msk, dst_img, dst_msk and augment (augmentation name or falsy).

Returns:

1 on success — used for progress counting.

spacr.io.process_non_tif_non_2D_images(folder)[source]

Split multi-dimensional or non-TIFF images in folder into per-channel TIFFs.

Grayscale non-TIFF images are converted to TIFF in place. Multi- dimensional images (3D/4D/5D) are split into one grayscale TIFF per (channel, Z, T) combination. Bit depth is preserved.

A file that cannot be read is recorded on a spacr.errors.RunLedger and skipped, so one corrupt image does not abort the folder — but the ledger prints a loud summary of everything that was skipped once the folder is done.

Parameters:

folder – Directory containing the input images.

Returns:

the spacr.errors.RunLedger for the conversion, so callers can check ledger.is_complete before trusting the folder’s contents.

spacr.io.read_plot_model_stats(train_file_path, val_file_path, save=False)[source]

Plot training vs. validation curves from a saved model’s per-epoch CSVs.

Parameters:
  • train_file_path – Path to the training stats CSV.

  • val_file_path – Path to the validation stats CSV.

  • save – If True, write the figures next to the training CSV instead of showing them. The file format follows the user’s figure preference via spacr.plot.save_figure(), which also corrects the extension. Default False.

Returns:

None

spacr.io.read_settings_history(db_path)[source]

Every settings snapshot ever written into db_path, oldest first.

Parameters:

db_path – path to a measurements.db.

Returns:

list of {'run_id', 'stage', 'stamped_utc', 'settings'}, one entry per recorded run, oldest first. A database that predates the history table returns [].

Example

from spacr.io import read_settings_history
for run in read_settings_history('.../measurements/measurements.db'):
    print(run['stamped_utc'], run['stage'],
          run['settings'].get('crop_mode'))
spacr.io.report_class_balance(labels, classes=None, class_balance='none', split_name='train', verbose=True)[source]

Measure class skew, decide what was done about it, and say so out loud.

A silent auto-fix is worse than none: the printed report always names the per-class counts, the imbalance ratio and the concrete action taken, and when class_balance='none' on skewed data it names the modes that would have helped instead of quietly doing nothing.

Parameters:
  • labels – integer labels of the split.

  • classes – ordered class names.

  • class_balance – requested mode, one of CLASS_BALANCE_MODES.

  • split_name – split being described ('train', 'validation', 'test').

  • verbose – print the report. The dict is returned either way.

Returns:

the summary dict, extended with mode, action, recommendation and report.

Raises:

ValueError – if class_balance is not a recognised mode.

spacr.io.report_cv_folds(labels, folds, classes=None, groups=None, group_by='none', verbose=True)[source]

Print the fold table and every warning the split earned.

Two failure modes are called out rather than allowed to surface later as mysterious metrics: a class too rare to reach every fold’s validation set (its recall is undefined there), and ungrouped folds on object crops (which leak well identity between train and validation).

Parameters:
  • labels – integer labels, one per sample.

  • folds – list of (train_idx, val_idx).

  • classes – ordered class names.

  • groups – optional group id per sample.

  • group_by – the grouping level that produced groups.

  • verbose – print the table and warnings.

Returns:

(fold_table, warnings).

spacr.io.save_object_mask(output_folder, filename, mask, compression='lzw')[source]

Save an integer label mask as a lossless, compressed TIFF.

Masks are saved as TIFF (not .npy) so they’re readable by ImageJ/other tools, with lossless compression (default LZW). Object labels are NEVER altered — the array is written verbatim as uint16, exactly as recorded in the measurements database.

Parameters:
  • output_folder – destination folder (e.g. masks/cell_mask_stack).

  • filename – reference filename (the stack basename; extension ignored).

  • mask – 2-D integer label array.

  • compression – lossless codec — 'lzw' | 'zlib' | 'none'.

Returns:

the path written.

spacr.io.select_fields(names, fields)[source]

Keep only the names whose field is in fields.

This filter allows mask generation to be rerun for selected fields without processing every field on the plate.

Parameters:
  • names – stack file names, as written by _rename_and_organize_image_files.

  • fields – what to keep. None or empty keeps everything, which is the default and the behaviour every existing run has. A list, or a comma-separated string, of field ids in any spelling the rest of spaCR accepts – 'f3', 3, 'F003' – or a glob such as 'f1*' matched against the field id.

Returns:

the kept names, in the order given.

spacr.io.summarize_class_imbalance(labels, classes=None)[source]

Measure the class skew of a label vector.

Parameters:
  • labels – iterable of integer class labels.

  • classes – ordered class names; index i names label i. Defaults to ['class_0', ...] sized to the largest label seen.

Returns:

dict with counts, fractions, imbalance_ratio (majority/minority, inf when a class is empty), minority, majority, empty_classes, skewed and severe.

spacr.io.summarize_cv_folds(labels, folds, classes=None, groups=None)[source]

Tabulate fold sizes and per-class validation counts.

Parameters:
  • labels – integer labels, one per sample.

  • folds – list of (train_idx, val_idx) from make_cv_folds().

  • classes – ordered class names.

  • groups – optional group id per sample; adds a distinct-group column.

Returns:

DataFrame with one row per fold.

spacr.io.training_dataset_from_annotation(db_path, dst, annotation_column='test', annotated_classes=(1, 2))[source]

Group per-object PNG paths by manual annotation values so they can be turned into a CNN training set.

Reads the png_list table of a spacr measurements.db, buckets PNG paths by the value found in annotation_column (typically filled by the spacr annotation GUI), and, when only one class has been annotated, samples an equal-sized “other” class from unannotated rows. The returned list-of-lists is consumed by generate_dataset_from_lists() to lay out train/<class>/*.png / test/<class>/*.png.

Parameters:
  • db_path – SQLite measurements.db containing a png_list table with png_path plus annotation_column.

  • dst – Output root (currently unused; kept for API symmetry with sister builders).

  • annotation_column – Column in png_list holding class labels. Default 'test'.

  • annotated_classes – Class values to pull from annotation_column. When length is 1, an equal-sized “other” class is sampled from rows whose annotation != that value.

Returns:

List of lists — one list of PNG paths per output class, in the same order as annotated_classes.

Example

from spacr.io import training_dataset_from_annotation, generate_dataset_from_lists
class_data = training_dataset_from_annotation(
    '/data/plate01/measurements/measurements.db',
    dst='/data/plate01/dataset',
    annotation_column='test', annotated_classes=(1, 2),
)
generate_dataset_from_lists('/data/plate01/dataset', class_data, classes=['neg','pos'])

See also

training_dataset_from_annotation_metadata() — same, but first restricts rows by plate row/column metadata. generate_dataset_from_lists() — turns the returned lists into a train/ / test/ folder tree.

spacr.io.training_dataset_from_annotation_metadata(db_path, dst, annotation_column='test', annotated_classes=(1, 2), metadata_type_by='columnID', class_metadata=None)[source]

Same as training_dataset_from_annotation() but pre-filtered by plate metadata.

Restricts source rows to those whose rowID or columnID is in class_metadata before grouping by annotation value.

Parameters:
  • db_path – SQLite database with a png_list table.

  • dst – Output root (unused; kept for API symmetry).

  • annotation_column – Column holding class labels.

  • annotated_classes – Class values to pull.

  • metadata_type_by – Which metadata column to filter on — 'rowID' or 'columnID'.

  • class_metadata – Allowed values for metadata_type_by. Default ['c1', 'c2'].

Returns:

List of lists — one list of PNG paths per output class.

Raises:

ValueError – if metadata_type_by is not 'rowID' or 'columnID'.

Nested helpers

_check_masks.needs_processing(filename)

Report whether a field still has to be generated.

Parameters:

filename (str) – Name relative to the enclosing output_folder, not a full path — it is joined onto that folder here. An existing file is validated by its header and length, so an empty or truncated .npy left behind by a killed run counts as missing, is named in the log, and is generated again.

spacr/io.py:4310

_dataset_crop_refs._filter(frame, column)

Keep the rows whose column contains any of the wanted terms.

SUBSTRING AND NOT REGEX (regex=False): the terms come from a user naming plates or wells, and a stray ( or + in one of them would otherwise raise out of pandas rather than simply matching nothing. Any term matching is enough – several terms are alternatives, which is what a user listing them means.

Parameters:
  • frame – the rows to filter.

  • column – the column to search; an absent one filters nothing, because a frame that never had it cannot contradict the request.

spacr/io.py:7147

_describe_processed_folder.count(name, suffixes)

Count the files in folder/name ending in suffixes.

spacr/io.py:3433

_load_and_concatenate_arrays.add_mask_folder(role, enabled)

Queue one object’s mask stack, if this run has that object.

EITHER the caller named a channel dimension for it OR the folder is on disk: a run that segmented an object always has the folder, and a run being re-read from settings may name the object before the folder is written. Requiring both would drop a mask stack that is sitting right there.

Parameters:
  • role – the object, e.g. 'cell'.

  • enabled – that object’s channel dimension, or None.

spacr/io.py:5344

_mask_movie_frame_geometry._points(fraction)

A font size in POINTS for a fraction of the frame’s short side.

Matplotlib sizes text in points and this geometry is in pixels, so the conversion has to use the dpi the writer will actually use – computing it against a default dpi puts the labels at the wrong size in the file while looking right on screen. Floored at 5 pt, below which a label is ink rather than text.

spacr/io.py:5007

_preprocess_mapped_volume_series.load_channel(channel)

Read one private channel from verified staged stacks.

spacr/io.py:3910

_read_and_merge_data._merge_grouped(left, right, right_name='grouped object data')

Merge grouped tables while keeping only one copy of shared acquisition metadata.

THE JOIN TYPE NOW COMES FROM THE REGISTRY. It used to be inner unconditionally – pandas’ default, since no how= was passed – and this docstring said the choice was “deliberately left alone” because the decision had not been made. It has since: object_roles. join_how records it and _read_and_join_tables already reads it, so the two readers of the same tables were disagreeing about which objects exist.

nucleus INNER a cell with no nucleus is debris png_list INNER a cell with no crop cannot be classified cytoplasm LEFT one row per cell; it makes no difference pathogen LEFT an UNINFECTED cell is usually the control organelle LEFT same reasoning

Inner for pathogen was the consequential one: it silently conditioned every result on infection, deleting the control population from the denominator without a word.

A right_name the registry does not know keeps the historical inner join rather than being guessed at – the metadata and stamp merges go through here too, and they are not object tables.

What is NOT defensible is doing it in silence, which is what this used to do. The discontinuity is brutal: on a 100-cell plate where NO crop carries a usable object id, png_list drops out before the join and all 100 cells survive; where exactly ONE does, the merge keeps that one and deletes the other 99. Nothing printed either way, and every shipped caller passes verbose=False.

So the shortfall is reported, named by table. This is the mirror of _report_fan_out for the shrinking direction.

spacr/io.py:5898

_read_and_merge_data._split_object_data(frame, group_by, object_type)

Group object data while retaining its complete provenance stamp.

spacr/io.py:6000

_read_db._quote_identifier(name)

Safely quote SQLite identifiers (e.g., table names).

spacr/io.py:5786

_save_mask_timelapse_as_gif._update(frame)

Update the frame of the animation.

Parameters: - frame (int): The frame number to update.

Returns: None

spacr/io.py:5067

_save_object_counts_to_database._count_objects(mask)

Count unique objects in a mask, assuming 0 is the background.

spacr/io.py:5129

_save_progress._save_df_to_csv(file_path, df)

Save the given DataFrame to the specified CSV file, either creating a new file or appending to an existing one.

Parameters: file_path (str): The file path where the CSV will be saved. df (pandas.DataFrame): The DataFrame to save.

spacr/io.py:5704

convert_to_yokogawa._get_next_well(used_wells)

Return the next free well, filling one plate before the next.

The well ids come from spacr.convert.well_sequence(), which builds them out of spacr.schema.PLATE_FORMATS — one definition instead of the three copies of "ABCDEFGHIJKLMNOP" and range(1, 25) this module used to carry.

The plate format stays 384 here: unlike convert_separate_files_to_yokogawa(), the inputs carry no well names at all, so nothing in them can ask for a bigger plate and the addresses are synthetic either way.

spacr/io.py:9283

crop_refs_for_rows._col(name)

One column as a plain list, or a column of None when it is absent.

PLAIN LISTS RATHER THAN itertuples, for two measured reasons: the joined UMAP frame carries a couple of hundred columns, so building a namedtuple per row of it costs more than reading the crops does, and itertuples silently RENAMES any column whose name is not a valid identifier – which is how a lookup starts missing a column that is plainly there.

Parameters:

name – the column, or a falsy value for “this frame has none”.

spacr/io.py:6808

crop_refs_for_rows._missing(value)

Whether a cell carries no answer.

NaN AS WELL AS None, because a column read out of pandas holds NaN where a row had nothing and None is not float('nan'). Testing only for None lets a NaN through as if it were a value, and it then reaches a path name or an object label.

spacr/io.py:6824

generate_training_dataset._annotation_classes_from_columns(png_df, ann_cols, ann_vals_filter=None, db_path=None)

Build classes per (column,value). If a column only has one annotated value in {1,2}, also create ‘<column>_random’ from unannotated rows (same count as positives). Optionally persist ‘<column>_random’ as a new INT column with 1’s for sampled rows.

Returns (names, lists) aligned.

spacr/io.py:8420

generate_training_dataset._apply_where(df, where)

where: list of {‘column’,’op’,’value’} AND-combined.

spacr/io.py:8354

generate_training_dataset._balance_lists(list_of_lists)

Cut every class down to the smallest one, when that was asked for.

A CLASSIFIER TRAINED ON 9,000 negatives and 300 positives learns to say “negative”, so balancing is the ordinary case rather than an exotic one – but it THROWS AWAY DATA, which is why it is a setting and not a default of this function.

Sampled rather than truncated: the first N crops of a class share a plate, a well and often a field, so taking them in order would trade a class imbalance for a batch imbalance.

The no-classes gate immediately before the call rejects an empty list, so only populated collections reach this.

Parameters:

list_of_lists – one list of crop paths per class.

spacr/io.py:8390

generate_training_dataset._class_items(frame)

Return the per-row crop entries a class list is built from.

On-demand handles when the merged source is in play, PNG paths otherwise – generate_dataset_from_lists takes either.

spacr/io.py:8318

generate_training_dataset._ensure_unique_dir(dst_base)

dst_base, or the first dst_base_N that does not exist yet.

A TRAINING SET IS NEVER WRITTEN OVER ONE THAT IS ALREADY THERE. The folder is the record of what a model was trained on, so reusing the name would leave a model whose training data cannot be reconstructed.

Parameters:

dst_base – the folder that was asked for.

Returns:

a folder path nothing occupies.

spacr/io.py:8271

generate_training_dataset._fix_path_under_src(src_root, p)

Make sure png_path lives under the current src root (portable absolute fix).

THE RULE ITSELF LIVES IN spacr.portable_paths and is shared with the montage, which needs exactly this and used to get none of it – the rule was a nested local here, reachable only from this generator, so a screen that had moved computer showed the montage 60,816 dead paths while this function resolved every one.

The /data/ rebuild is now only applied when it lands on a file that EXISTS. Rewriting to somewhere equally absent is strictly worse than leaving the recorded path alone: the copy below then fails naming a folder the user never chose.

spacr/io.py:8329

generate_training_dataset._load_png_table(db_path, object_type='cell')

The per-object crop table, or the measurements standing in for it.

png_list ALONE, deliberately: joining it against the measurement tables would drop every object those tables do not also carry, and a training set is allowed to be a subset of what was measured.

NO png_list MEANS NO PNG FOLDER WAS EVER WRITTEN, which is not the same as no data. The objects are still in the measurement table with the same well metadata the class rules select on, so that is read instead – otherwise a project holding everything it needs reports “0 classes”.

Parameters:
  • db_path – the measurements database.

  • object_type – which object’s table to fall back to.

spacr/io.py:8291

make_validation_holdout.score(candidate)

Rank one candidate fold; lower is a better holdout.

Parameters:

candidate – A (train_idx, val_idx) pair as produced by make_cv_folds(); only the validation half is looked at.

Returns:

(cost, n_validation) where cost adds the size error (as a fraction of the dataset) to the mean absolute per-class deviation from the whole dataset’s distribution. The trailing count is a tie-break, so equally good folds resolve to the smaller holdout.

spacr/io.py:7711

prepare_cellpose_dataset.find_image_mask_pairs(dataset_path)

Return (image_path, mask_path) pairs found under dataset_path.

Parameters:

dataset_path – One dataset folder holding its images at the top level and a masks subfolder in which each mask carries exactly the same file name as its image. Only .tif/.tiff images are considered, matching is by name rather than by order, and an image whose mask is absent is dropped without a message — so a short pair count here means missing or renamed masks.

spacr/io.py:9568

prepare_cellpose_dataset.get_augmentations()

Return the list of augmentation names used to expand datasets.

spacr/io.py:9564

prepare_cellpose_dataset.prepare_output_folders(base)

Create train/{images,masks} and test/{images,masks} under base.

Parameters:

base – Output root, which the caller sets to <input_root>/cellpose_dataset. Creation is exist_ok, so rerunning does not fail, but nothing is emptied first and the copies are renumbered from 00000, so a smaller second run leaves the tail of a larger earlier one mixed into the split.

spacr/io.py:9588

process_non_tif_non_2D_images.convert_grayscale_to_tiff(image, filename, folder, dtype)

Convert grayscale images that are not in TIFF format to TIFF, preserving bit depth.

Parameters:
  • image – The decoded 2D plane to write.

  • filename – Base name of the source file, not a path. Only the extension is stripped, so an absolute path here would be joined with folder and win, writing the TIFF back next to the original instead of into folder.

  • folder – Directory the .tif is written into; the original file is left in place next to it rather than replaced.

  • dtype – NumPy dtype the plane is cast to, which is what carries the source bit depth into the TIFF.

spacr/io.py:325

process_non_tif_non_2D_images.load_image(file_path)

Loads image from various formats and returns it as a numpy array along with its dtype.

Parameters:

file_path – Path to the image. Only the extension selects the reader: .tif and .tiff go to tifffile, .png, .jpg and .jpeg to PIL, .czi to czifile, .nd2 to ND2Reader. Content is never sniffed, so a mislabelled file is read with the wrong reader, and any other extension raises ValueError.

spacr/io.py:293

process_non_tif_non_2D_images.save_grayscale_images(image, base_name, folder, dtype, channel=None, z=None, t=None)

Save grayscale images with appropriate suffix based on channel, z, and t, preserving bit depth.

Parameters:
  • image – A single 2D plane already sliced out of the stack.

  • base_name – Stem of the source file. Every plane cut from one source shares it, so the _C/_Z/_T suffix built below is the only thing keeping those planes from overwriting each other.

  • folder – Directory the TIFF is written to. The converter passes the folder it is scanning, so planes land beside the multi-dimensional file they came from.

  • dtype – NumPy dtype the plane is cast to before writing. This cast is the whole of the “bit depth is preserved” promise — pass the dtype load_image reported for the source rather than a convenient default, or 16-bit data is silently rewritten.

  • channel – 1-based channel index appended as _C. None leaves that part of the suffix off.

  • z – 1-based Z index appended as _Z; None omits it.

  • t – 1-based time index appended as _T; None omits it.

spacr/io.py:232

process_non_tif_non_2D_images.split_channels(image, folder, base_name, dtype)

Splits the image into channels and handles 3D, 4D, and 5D image cases.

Parameters:
  • image – Array whose axis order is assumed to be (height, width, channel[, Z[, T]]). A 2D array returns without writing anything (the caller handles those), and anything with more than five axes falls through silently, writing nothing. The axis order is never checked, so a channel-first array from some other reader is sliced along width instead of being rejected.

  • folder – Directory the per-plane TIFFs are written into.

  • base_name – Stem shared by every plane written from this image; the _C/_Z/_T indices are appended to it, all 1-based.

  • dtype – NumPy dtype passed straight through to each write, so it should be the source image’s own dtype to keep its bit depth.

spacr/io.py:260

read_plot_model_stats._plot_and_save(train_df, val_df, column='accuracy', save=False, path=None, dpi=None)

Draw one training curve – train against validation – and write it.

One function per COLUMN rather than per figure because the caller asks for accuracy, loss and the rest by name, and every one of them is the same plot of the same two frames.

Parameters:
  • train_df – per-epoch training statistics.

  • val_df – the same for validation.

  • column – which statistic to draw.

  • save – write a PDF beside the model rather than only showing it.

  • path – the folder to write into.

  • dpi – resolution for the written file.

spacr/io.py:5555

select_fields.field_of(token)

One field’s canonical token, so ‘f3’, ‘3’ and ‘F003’ are one field.

NORMALISED THROUGH schema, which is the same rule the file names themselves were written by – matching the raw text instead would make a selection depend on which of three spellings the user typed. Anything schema cannot read is lower-cased and passed through, so a token from a convention it has not met still selects itself.

spacr/io.py:2818