spacr.submodules

Workflow inputs and outputs

Plaque Assay

Analyse plaque images or existing masks with the configured plaque model. This route need not pass through Measure; use the dedicated plaque example.

Open: Toxoplasma → Plaque Assay.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.

  • Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.

Outputs

  • Assay results — Assay-specific result tables and figures in the configured destination, preserving well and condition identities.

  • Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.

Before this module

  • Make Masks: Use plaque masks with matching source images; cell masks are not automatically plaque labels.

API reference.

Module tutorial.

Recruitment

Use compartment intensity measurements and matching host/pathogen identities to compute recruitment ratios.

Open: Toxoplasma → Recruitment.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route: cell, nucleus, pathogen, cytoplasm. Relevant columns, depending on the route: plateID, rowID, columnID, fieldID.

Outputs

  • Assay results — Assay-specific result tables and figures in the configured destination, preserving well and condition identities.

Before this module

  • Measure: Require the intended compartment intensities and identities.

API reference.

Module tutorial.

Invasion Assay

Use the required two-colour differential-staining measurements and stain-baseline controls to distinguish attachment from invasion.

Open: Toxoplasma → Invasion Assay.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route: cell, nucleus, pathogen, cytoplasm. Relevant columns, depending on the route: plateID, rowID, columnID, fieldID.

Outputs

  • Assay results — Assay-specific result tables and figures in the configured destination, preserving well and condition identities.

Before this module

  • Measure: Require two-colour stain measurements and appropriate baseline controls.

API reference.

Module tutorial.

Replication Assay

Count parasites using explicit vacuole identity and compare condition distributions; host identity alone does not define a vacuole.

Open: Toxoplasma → Replication Assay.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route: cell, nucleus, pathogen, cytoplasm. Relevant columns, depending on the route: plateID, rowID, columnID, fieldID.

Outputs

  • Assay results — Assay-specific result tables and figures in the configured destination, preserving well and condition identities.

Before this module

  • Measure: Require explicit parasite-to-vacuole identities.

API reference.

Module tutorial.

Cellpose Workbench

Open Cellpose Workbench inside Make Masks and train from verified image/mask pairs. Evaluate on separate fields before selecting the checkpoint in Mask.

Open: Make Masks → Cellpose Workbench.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Curated training fields — Separate image and integer-mask files with matching field identities; preserve original images and labels.

Outputs

  • Segmentation checkpoint — Saved Cellpose-compatible checkpoint or a compatible installed backend selected with its own configuration.

Before this module

  • Make Masks: Use independently checked image/mask pairs.

After this module

  • Mask: Select the saved compatible checkpoint in Mask.

  • Direct Cellpose mask generation: Pass the trained checkpoint as custom_model with the matching image channels and preprocessing.

API reference.

Module tutorial.

Endodyogeny size proxy

Read measured compartment areas, annotate conditions and bin area ** 1.5 into log2 size doublings. Defaults aggregate pathogen area per host cell, not per vacuole: multiple vacuoles in one cell are combined. This is an area-derived size proxy, not measured volume or a parasite count. Use Replication Assay with explicit vacuole identity for parasites-per-vacuole counts. Configure compartment, area filters, calibration, conditions and grouping before calling the API; saving is optional.

Use from Python: spacr.submodules.analyze_endodyogeny(). This API-only workflow has no Home tile or menu entry.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Host-cell-aggregated compartment areas — Each src project root/measurements/measurements.db; tables defaults to cell, nucleus, pathogen and cytoplasm, with png_list added for merging. The compartment setting selects the area column; default pathogen_area is summed per host cell. Relevant tables, depending on the route: cell, nucleus, pathogen, cytoplasm, png_list. Relevant columns, depending on the route: cell_id, pathogen_area.

Outputs

  • Area-derived size-proxy results — Returned data and chi_squared DataFrames; save=True also writes data.csv, chi_squared_results.csv, chi_squared_pairwise_results.csv and a figure under the first project root/results/analyze_endodyogeny/. Relevant columns, depending on the route: pathogen_area, pathogen_volume, pathogen_volume_bin, bin_index.

Before this module

  • Measure: Supply the measured project roots and required object/png_list tables. Verify host-cell aggregation and area units before interpreting size bins; the Mask counts database alone is insufficient.

API reference.

Run plaque, recruitment, invasion, and replication assays.

WHAT IT IS FOR

Four spaCR tiles currently share this landing page, but they answer different biological questions. analyze_plaques() segments plaque images and summarizes plaque number and area. analyze_recruitment() measures a fluorescent marker around pathogens or vacuoles relative to host cytoplasm. analyze_invasion() uses differential pre/post-permeabilization staining to classify parasites as attached outside or invaded inside a host cell. analyze_replication() counts parasites within each parasitophorous vacuole and compares the resulting replication-state distributions. Cellpose training, testing, and model-application utilities also live here, but they are not substitutes for those four assay entry points.

WHAT IT NEEDS

Plaque analysis accepts a folder of TIFF images, or existing masks beneath that folder, plus Cellpose settings and a bundled, catalogue, or local plaque model. Recruitment starts from a spaCR measurements.db containing joined cell, nucleus, pathogen, and cytoplasm features; it needs a fluorescence channel, object filters, and plate metadata that assign cell type, pathogen, and treatment. Invasion and Replication both need one row per segmented parasite in a measurement table and condition metadata. Invasion additionally needs the outside- and total-stain channels and preferably known control wells; Replication needs a defensible vacuole_key or spatial-linking distance. The Cellpose utilities require paired images and masks for training/testing, or an image folder and model path for inference.

WHAT IT PRODUCES

Plaque analysis writes <src>/masks/plaques_analysis.db with summary, stats, and details tables. Recruitment returns per-object and per-well DataFrames and writes their CSVs and plots. Invasion returns per-parasite classifications, per-field thresholds and QC, per-well efficiencies, condition summaries and comparisons, controls, and figures; saved runs place those artifacts under results/analyze_invasion. Replication returns per-vacuole counts, well and condition distributions, pairwise and omnibus statistics, figures, and the grouping method actually used, with saved output under results/analyze_replication. Model utilities produce trained weights, evaluation tables, masks, and object summaries as appropriate.

WHAT TO DO NEXT

For plaques, inspect the masks before interpreting counts or areas. For Recruitment, verify the object filters, condition annotation, and per-well denominators before comparing treatments. For Invasion, review field-level thresholds, control agreement, bimodality, and sensitivity flags before using the efficiency table. For Replication, inspect the vacuole grouping and the reported non-power-of-two fraction before comparing doubling distributions. Follow the specific function links above until the four tiles receive separate API destinations.

The analysis unit matters. Invasion is inferred from absence of outside stain, so weak staining can only inflate the invaded fraction; thresholds are therefore recorded per field and statistics use wells rather than treating parasites from one well as independent replicates. Replication groups by vacuole, not by host cell, because one cell can contain several vacuoles; 3, 5, 6, and 7 parasites remain in an explicit non-power-of-two QC bucket instead of being rounded into a biologically expected class. Plaque area is calibrated against the well scale when available, so comparisons should retain the acquisition metadata that defines that scale.

Exceptions

Cellpose3Checkpoint

A Cellpose 3 checkpoint was handed to Cellpose 4, which cannot load it.

ModelZooMissing

A named model is not where it should be.

Classes

CellposeLazyDataset

Lazy image/label dataset for Cellpose training and inference.

Functions

analyze_class_proportion(settings)

Test whether classifier class proportions differ between experimental groups.

analyze_endodyogeny(settings)

Bin pathogen size by log2 doublings and test the bin proportions per group.

analyze_invasion(settings)

Invasion assay: score every parasite attached or invaded and report efficiency per well.

analyze_percent_positive(settings)

Annotate objects above a threshold and summarise positive fractions per well.

analyze_plaques(settings)

Segment host-cell plaques with a bundled Cellpose model and summarize per-image counts and areas.

analyze_recruitment(settings)

Measure marker recruitment with host-cell and per-well summaries.

analyze_replication(settings)

Replication assay: count parasites per vacuole and compare the distributions.

apply_cellpose_model(settings)

Run a Cellpose model over a folder of images and export per-object measurements.

compare_reads_to_scores(reads_csv, scores_csv[, ...])

Compare sequencing read fractions to classifier score fractions across wells.

count_phenotypes(settings)

Count unique phenotype annotations per plate/row/column and export to CSV.

display(*args, **kwargs)

Do nothing: IPython is unavailable, so there is nowhere to display to.

explain_cellpose3(exc, model)

Turn Cellpose 4's refusal of a Cellpose 3 checkpoint into advice.

generate_score_heatmap(settings)

Combine multiple classifier score CSVs into a per-well heatmap and MAE table.

interpret_vision_model([settings])

Explain a spacr vision-model score by ranking which morphology / intensity features drive it.

plot_cellpose_batch(images, labels)

Display a two-row grid of images and their paired label masks.

post_regression_analysis(csv_file, grna_dict, grna_list)

Compute gRNA correlation and propagate fixed effect sizes across correlated gRNAs.

split_wells(settings)

Cut every multi-well image under src into one image per well.

test_cellpose_model(settings)

Evaluate a Cellpose model on a labelled test set and report per-image metrics.

train_cellpose(settings)

Fine-tune Cellpose-SAM with native paired images and instance-label masks.

Module Contents

exception spacr.submodules.Cellpose3Checkpoint[source]

Bases: ValueError

A Cellpose 3 checkpoint was handed to Cellpose 4, which cannot load it.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.submodules.ModelZooMissing[source]

Bases: FileNotFoundError

A named model is not where it should be.

Initialize self. See help(type(self)) for accurate signature.

class spacr.submodules.CellposeLazyDataset(image_files, label_files, settings, randomize: bool = True, augment: bool = False)[source]

Bases: torch.utils.data.Dataset

Lazy image/label dataset for Cellpose training and inference.

Loads paired image and label tiffs on demand, optionally normalizing, augmenting (8-fold rotations/flips), and resizing to a target size.

Parameters:
  • image_files – paths to input image tiffs.

  • label_files – paths to matching label tiffs (same length as image_files).

  • settings – dict with keys normalize, percentiles, target_size.

  • randomize – shuffle the image/label pairing order. Default True.

  • augment – enable 8-fold augmentation (dataset length x8). Default False.

Raises:

ValueError – when image/label lists differ in length or are empty.

Pair the image and label files and fix the augmentation factor.

Mismatched lengths raise here rather than at the first bad index, so a wrongly paired dataset fails at construction instead of part-way through an epoch.

__getitem__(idx)[source]

Load one item, decoding idx into a file and an augmentation.

The file is read HERE rather than at construction, which is what makes the dataset lazy: a plate larger than memory costs one image at a time.

__len__()[source]

Files times augmentations – the dataset presents each variant as its own item.

spacr.submodules.analyze_class_proportion(settings)[source]

Test whether classifier class proportions differ between experimental groups.

Runs chi-squared and pairwise tests on the class column, plots stacked bars and a plate heatmap, and follows up with normality, Levene, and posthoc statistical tests.

Parameters:

settings – dict of settings; see set_analyze_class_proportion_defaults for keys including src, tables, class_column, group_column, level and save.

Returns:

dict with data (annotated DataFrame) and chi_squared (results DataFrame).

spacr.submodules.analyze_endodyogeny(settings)[source]

Bin pathogen size by log2 doublings and test the bin proportions per group.

This is the size-proxy replication readout, not a parasite count. Read that sentence twice before quoting a number from it:

  • The rows come from spacr.io._read_and_merge_data(), which collapses the per-object pathogen table onto the host cell (prcfo is built from cell_id). pathogen_area on each row is therefore the sum of the areas of every pathogen object inside that host cell — one host cell carrying two parasitophorous vacuoles contributes a single row holding the combined area of both.

  • area ** 1.5 is a 2-D-to-3-D size proxy, not a measured volume.

  • Nothing here counts parasites. A bin is a doubling of area-derived size, which tracks parasites-per-vacuole only while the pathogen mask segments whole vacuoles and each host cell holds exactly one.

Keep using it when the pathogen channel gives you fused rosettes that cannot be resolved into single parasites. When the individual parasites are resolvable, analyze_replication() counts them and reports the parasites-per-vacuole distribution directly, which is the readout an endodyogeny experiment is actually after.

Parameters:

settings – dict of endodyogeny settings; see set_analyze_endodyogeny_defaults for keys including src, tables, compartment, min_area_bin, max_area, max_bins, um_per_px, group_column, level and save.

Returns:

dict with data (binned DataFrame) and chi_squared (results DataFrame).

Example

from spacr.submodules import analyze_endodyogeny
out = analyze_endodyogeny({'src': '/data/plate1', 'save': True})

See also

analyze_replication() — counts parasites per vacuole instead of inferring replication from object size.

spacr.submodules.analyze_invasion(settings)[source]

Invasion assay: score every parasite attached or invaded and report efficiency per well.

The red/green invasion assay stains twice. Before permeabilisation an antibody reaches only the parasites still outside the host cell, so those are positive in both channels; the cells are then permeabilised and a second antibody stains all parasites, so a parasite positive only in the post-permeabilisation channel was inside. Hence:

  • attached / outside = present in the outside-stain channel;

  • invaded / inside = absent from the outside-stain channel.

Read that asymmetry carefully, because the whole design follows from it. “Inside” is defined by an absence, and absence is the unreliable direction. Poor antibody penetration, a focal plane off the parasite’s equator, photobleaching, a low-expressing parasite — every one of them removes outside signal from a parasite that is genuinely outside, and every one of them therefore inflates invasion efficiency. Nothing plausible pushes the error the other way. The threshold on the outside channel is the single number the assay rests on, so it is derived from the data, reported per field in fields, cross-checked against a control-derived cut when one exists, and bracketed by a sensitivity pair that says how much of the answer is the threshold.

Three design decisions worth stating outright:

  • The threshold is per field. Illumination and staining vary field to field, and a plate-wide cut turns an illumination gradient into an invasion gradient. See _invasion_field_thresholds().

  • Controls beat any automatic method. control_wells names wells whose parasites are known to carry no outside stain; the threshold is then a high quantile of that honest negative distribution (control_quantile), the control wells are excluded from the results, and threshold_source says 'control' so the report cannot be mistaken for an automatic run.

  • A threshold without two populations is arbitrary. Otsu will happily split a single smear of signal down the middle and return a confident number. bimodality_coefficient and qc_flag_unimodal say when that has happened, per field and per well, instead of letting it pass silently. See _bimodality_coefficient().

invasion_efficiency = n_invaded / (n_invaded + n_attached) and it is always reported next to n_total: 90% from ten parasites and 90% from four thousand are not the same result. A well that scored nothing gets NaN, not 0.0.

Statistics use the well as the unit of replication. Parasites within a well share a coverslip, an antibody bath and a focal plane, so they are not independent; the reported test is a Mann-Whitney U on the per-well efficiencies. A pooled-parasite chi-squared is reported beside it purely so its inflation is visible. See _invasion_compare_conditions().

Parameters:

settings –

dict of invasion settings; see set_analyze_invasion_defaults. Key entries:

  • src — plate directory (or list) holding measurements/measurements.db.

  • parasite_table / compartment — table and column prefix with one row per segmented parasite. Default 'pathogen'.

  • outside_channel / total_channel — the pre- and post-permeabilisation stain channels.

  • intensity_statistic — which per-object statistic of the outside channel to threshold; 'auto' prefers the boundary-restricted one. See _resolve_invasion_intensity_column().

  • background_correction — optional per-object local background.

  • outside_threshold_method / outside_threshold — automatic method, or a fixed cut that overrides it.

  • control_wells / control_quantile / min_control_objects.

  • min_objects_for_threshold / min_objects_for_bimodality / bimodality_cutoff / threshold_agreement_tolerance / threshold_sensitivity / inflation_warn / min_parasites_per_well — the QC thresholds.

  • extracellular_class — how parasites with no host cell are scored.

  • cell_types / pathogen_types / treatments and their *_plate_metadata well maps, plus group_column and level.

  • save — write the CSVs and figures under <src>/results/analyze_invasion.

Returns:

dict with parasites (per-object classification), fields (per-field thresholds and QC), wells (per-well efficiency, denominators and QC flags), summary (per condition), comparisons (per-well statistics), chi_squared / chi_squared_pairwise (the shared proportion-bar omnibus tests), controls (the control-well objects, if any), control_thresholds, intensity_column, intensity_statistic and figures.

Raises:
  • ValueError – when the parasite table holds no usable rows.

  • KeyError – when the requested statistic or group column is absent.

Example

from spacr.submodules import analyze_invasion
out = analyze_invasion({
    'src': '/data/plate1',
    'outside_channel': 1,
    'total_channel': 0,
    'stain_baseline_wells': ['c12'],
    'pathogen_types': ['dmso', 'inhibitor'],
    'pathogen_plate_metadata': [['c1'], ['c2']],
})
print(out['wells'][['prc', 'n_total', 'invasion_efficiency',
                    'qc_flags']])

See also

analyze_replication() — the parasites-per-vacuole assay, whose table reading, condition annotation and output layout this follows.

spacr.submodules.analyze_percent_positive(settings)[source]

Annotate objects above a threshold and summarise positive fractions per well.

Merges measurements from measurements.db, thresholds on a chosen feature column, then joins the resulting well-level counts against rename_log.csv to recover human-readable plate/well identifiers.

Parameters:

settings – dict of settings; see default_settings_analyze_percent_positive for keys including src, tables, value_col, threshold and filter_1.

Returns:

DataFrame of annotated per-well positive/negative counts and fractions.

spacr.submodules.analyze_plaques(settings)[source]

Segment host-cell plaques with a bundled Cellpose model and summarize per-image counts and areas.

Downloads (if needed) the bundled toxo_plaque_cyto_e25000 model, runs Cellpose over every .tif under src, then computes per-image plaque count + mean/stddev area and writes a plaques_analysis.db (tables: summary, stats, details) alongside the masks.

Parameters:

settings –

Settings dict, canonicalized via spacr.settings.get_analyze_plaque_settings(). Key entries:

  • src — folder containing plaque images.

  • masks — if truthy, run segmentation before analysis; if falsy, expect masks already in <src>/masks.

  • diameter, flow_threshold and CP_prob, read by spacr.plaque.segment_plaque_image(), the call the Plaque preview makes too.

  • plaque_mode – 'figure' hands the folder to spacr.plaque_papers.measure_figure_folder() instead.

  • colony_counting – in plaque mode, counts bacterial or fungal colonies on plate photos instead of segmenting plaques (_analyze_colony_plates()), writing <src>/colonies/colonies.db.

Returns:

None. Writes <src>/masks/plaques_analysis.db. With colony_counting it returns the per-plate colony table instead.

Example

from spacr.submodules import analyze_plaques
analyze_plaques({'src': '/data/plaque_assay', 'masks': True})

See also

analyze_recruitment() — intensity-ratio phenotype instead of plaque counts.

spacr.submodules.analyze_recruitment(settings)[source]

Measure marker recruitment with host-cell and per-well summaries.

Reads the merged cell/nucleus/pathogen/cytoplasm feature tables from a spacr measurements.db, annotates each row with cell type / pathogen / treatment based on plate metadata, filters objects by size and intensity, computes the pathogen-to-cytoplasm mean-intensity ratio for channel_of_interest, groups by well and writes both results/cells.csv and results/wells.csv alongside recruitment plots.

Each cell row combines the pathogen measurements assigned to that host cell. Pathogen mean intensities are averaged across its associated objects; these rows represent host cells rather than independently measured vacuoles. The main recruitment ratio divides that aggregate pathogen mean by the cell’s cytoplasm mean. Each well averages its retained cell ratios.

In the GUI, open Home > Toxoplasma > Recruitment. Select a measured project, map its channels and plate conditions, review the object filters, and Run. Inspect the retained counts and ratio columns before comparing conditions. Condition plots show between-well standard deviations. For measurements linked to individual vacuoles, use spacr.host_pathogen.

Parameters:

settings –

Settings dict, canonicalized via spacr.settings.get_analyze_recruitment_default_settings(). Key entries:

  • src — folder containing measurements/measurements.db and optional merged images for overlays. A database path is also accepted; a database outside a measurements folder may be moved into one, so use a project copy when reorganizing existing data.

  • cell_types / cell_plate_metadata — labels + row/col metadata that map wells to cell lines.

  • pathogen_types / pathogen_plate_metadata.

  • treatments / treatment_plate_metadata.

  • channel_of_interest — intensity channel for the ratio.

  • cell_chann_dim / nucleus_chann_dim / pathogen_chann_dim — recorded object-channel mapping used by image overlays and intensity filtering.

  • cell_size_range, nucleus_size_range, pathogen_size_range — [min, max] px area filters.

  • *_intensity_range, target_intensity_min.

  • cells_per_well — minimum well count to keep.

  • plot, plot_control, plot_nr, figuresize.

Returns:

List [cells, wells] — the host-cell and per-well recruitment DataFrames, also written to CSV under src/results.

Example

from spacr.submodules import analyze_recruitment
settings = {
    'src': '/data/plate01',
    'cell_types': ['HeLa'], 'cell_plate_metadata': ['c2-c11'],
    'pathogen_types': ['tgme49'], 'pathogen_plate_metadata': ['c2-c11'],
    'treatments': ['dmso','drug'], 'treatment_plate_metadata': [['r1'],['r2']],
    'channel_of_interest': 3,
}
cells_df, wells_df = analyze_recruitment(settings)

See also

analyze_plaques() — plaque-count/size assay. spacr.ml.generate_ml_scores() — feature-based classifier as an alternative to recruitment ratios.

spacr.submodules.analyze_replication(settings)[source]

Replication assay: count parasites per vacuole and compare the distributions.

replication_method='direct_count' is the default described below. 'size_proxy' delegates to analyze_endodyogeny() and returns its area-derived, host-aggregated readout instead. Both return the selected method in replication_method. 'deep_learning_coming_soon' raises NotImplementedError before any data are read or outputs written; the whole-vacuole classification model is not available yet.

Toxoplasma gondii replicates by endodyogeny, two daughters forming inside a mother, so a parasitophorous vacuole holds 1, 2, 4, 8 or 16 parasites — a power of two. The readout of a replication assay is therefore the distribution of parasites-per-vacuole across a well, not a mean: a mean of 3.2 cannot distinguish “everything at 3-ish”, which is biologically impossible, from a healthy mix of 2s and 4s. A drug that slows replication moves mass from the 8 and 4 buckets down into 2 and 1, and only the distribution shows that.

The counting unit is the vacuole. Not the parasite, and emphatically not the host cell — one host cell routinely carries several vacuoles, so grouping on cell_id reports their combined count as a single vacuole and produces a plausible but meaningless number. See _assign_vacuole_ids() for how the vacuole is derived and what each vacuole_key costs you.

Rosettes of 3, 5, 6 or 7 are counted into an explicit non_power_of_two bucket that is always reported and never folded into a neighbouring bucket. That bucket is the assay’s own quality control: a well where 30% of vacuoles are off the power-of-two ladder has a segmentation problem, and its replication number should not be trusted.

Statistics: the two-condition comparison is a Mann-Whitney U test on the doubling index log2(n_parasites), with a chi-squared omnibus test alongside it. _replication_compare_conditions() explains why, and why a t-test on the raw counts is the wrong instrument.

Parameters:

settings –

dict of replication settings; see set_analyze_replication_defaults. Key entries:

  • src — plate directory (or list of them) holding measurements/measurements.db.

  • parasite_table / compartment — table and column prefix holding one row per segmented parasite. Default 'pathogen'.

  • vacuole_key — how parasite rows are grouped into vacuoles ('auto', 'spatial', 'cell_id', 'object', or a column name).

  • vacuole_link_distance / vacuole_link_factor — the spatial clustering threshold, or the multiplier used to derive it.

  • min_parasite_area / max_parasite_area — debris and merged-clump filters applied before counting.

  • max_parasites_per_vacuole — largest named power-of-two bucket.

  • non_power_of_two_warn — QC flag threshold.

  • cell_types / pathogen_types / treatments and their *_plate_metadata well maps, plus group_column and level.

  • save — write the CSVs and figures under <src>/results/analyze_replication.

Returns:

dict with vacuoles (per-vacuole counts), wells (per-well distribution), summary (per-condition distribution), comparisons (pairwise ordered tests), chi_squared / chi_squared_pairwise (omnibus proportion tests), figures and vacuole_key (the grouping actually used).

Raises:

ValueError – when the parasite table holds no usable rows.

Example

from spacr.submodules import analyze_replication
out = analyze_replication({
    'src': '/data/plate1',
    'pathogen_types': ['dmso', 'pyrimethamine'],
    'pathogen_plate_metadata': [['c1'], ['c2']],
})
print(out['summary'][['condition', 'frac_1', 'frac_2', 'frac_4',
                      'frac_8', 'frac_non_power_of_two']])

See also

analyze_endodyogeny() — the size-proxy version, for fused rosettes that cannot be resolved into single parasites.

spacr.submodules.apply_cellpose_model(settings)[source]

Run a Cellpose model over a folder of images and export per-object measurements.

Optionally masks predictions to a central circle, then records per-object area to measurements.csv and a per-image summary to summary.csv.

Parameters:

settings – dict of inference settings; see get_default_apply_cellpose_model_settings for keys including src, model_path, batch_size, FT, CP_probability, circularize and save.

Returns:

None. Writes result CSVs under <src>/results.

spacr.submodules.compare_reads_to_scores(reads_csv, scores_csv, empirical_dict=None, pc_grna='TGGT1_220950_1', nc_grna='TGGT1_233460_4', y_columns=None, column='columnID', value='c3', plate=None, save_paths=None)[source]

Compare sequencing read fractions to classifier score fractions across wells.

Loads paired reads and scores tables (single files or matched lists), computes per-well class-1 and gRNA fractions, joins them with an empirical row-to-mixture dictionary, and plots the fractions against the positive- and negative-control fractions.

Parameters:
  • reads_csv – path (or list of paths) to per-gRNA read count CSVs.

  • scores_csv – path (or list of paths) to per-object classifier score CSVs.

  • empirical_dict – mapping of rowID to (pc_units, nc_units) mixture; a 16-row default is used when None.

  • pc_grna – positive-control gRNA name. Default 'TGGT1_220950_1'.

  • nc_grna – negative-control gRNA name. Default 'TGGT1_233460_4'.

  • y_columns – Columns to plot on the y axis. None uses ['class_1_fraction', 'TGGT1_220950_1_fraction', 'nc_fraction'].

  • column – column used to select a subset of wells. Default 'columnID'.

  • value – value in column to keep. Default 'c3'.

  • plate – plate ID to stamp when a single pair of CSVs is given.

  • save_paths – two-element list of PDF output paths (pc plot, nc plot).

Returns:

two matplotlib figures [fig_pc, fig_nc].

spacr.submodules.count_phenotypes(settings)[source]

Count unique phenotype annotations per plate/row/column and export to CSV.

Parameters:

settings – dict with src (pointing at a measurements folder or measurements.db) and annotation_column (the column of interest in the png_list table).

Returns:

None. Writes phenotype_counts.csv next to the database.

spacr.submodules.display(*args, **kwargs)[source]

Do nothing: IPython is unavailable, so there is nowhere to display to.

THE FALLBACK IS THE POINT. IPython.display.display is imported at module scope, and IPython can be mid-init – partially imported by another thread – which makes that import raise. Letting it propagate would make importing this module fail for a reason that has nothing to do with what the module does. spaCR only calls display from notebook contexts; the Qt GUI ignores it.

Parameters:
  • args – whatever the caller would have displayed.

  • kwargs – likewise.

spacr.submodules.explain_cellpose3(exc, model)[source]

Turn Cellpose 4’s refusal of a Cellpose 3 checkpoint into advice.

Parameters:
  • exc – the exception Cellpose raised.

  • model – what was asked for, for the message.

Returns:

a Cellpose3Checkpoint when exc is that refusal, else exc unchanged.

spacr.submodules.generate_score_heatmap(settings)[source]

Combine multiple classifier score CSVs into a per-well heatmap and MAE table.

Aggregates per-object scores across score CSVs, merges with a cross-validation score and a reads-derived fraction column, plots a multi-channel heatmap, and computes per-channel mean absolute error against the empirical fraction.

Parameters:

settings – dict of settings including folders, csv_name, data_column, csv, cv_csv, data_column_cv, plateID, columnID, control_sgrnas, fraction_grna, cmap and dst.

Returns:

merged DataFrame joining reads, classifier scores and CV scores per well.

spacr.submodules.interpret_vision_model(settings=None)[source]

Explain a spacr vision-model score by ranking which morphology / intensity features drive it.

Joins the per-object CNN predictions (score_column) with the morphology + intensity measurements from spacr.measure.measure_crop(), expands cross-compartment feature ratios (e.g. nucleus_cell_area), then runs random-forest feature importance, permutation importance and (optionally) SHAP on the top features. Also groups importance by compartment and by channel so you can answer “is my classifier looking at the pathogen or at the cell?”.

Parameters:

settings –

Settings dict. Key entries:

  • src — folder containing measurements/measurements.db with both feature and score tables.

  • tables — DB tables to merge, e.g. ['cell','nucleus','pathogen','cytoplasm'].

  • channels — intensity channels included in the feature space (e.g. [0,1,2,3]).

  • score_column — column holding per-object CNN scores.

  • top_features — cap on features shown / SHAP-explained.

  • feature_importance / permutation_importance / shap — toggle each explainer.

  • shap_sample — subsample size for SHAP.

  • nuclei_limit / pathogen_limit — object-count caps in the read/merge step.

  • n_jobs, save.

Returns:

Dict of DataFrames keyed by analysis name ('feature_importance', 'permutation_importance', 'shap', 'compartment_importance', 'channel_importance', …).

Example

from spacr.submodules import interpret_vision_model
results = interpret_vision_model({
    'src': '/data/plate01',
    'score_column': 'pred',
    'channels': [0,1,2,3],
    'top_features': 30, 'shap': True,
})

See also

spacr.deep_spacr.deep_spacr() — trains the model whose scores this function interprets.

spacr.submodules.plot_cellpose_batch(images, labels)[source]

Display a two-row grid of images and their paired label masks.

Parameters:
  • images – iterable of 2D grayscale image arrays.

  • labels – iterable of matching integer label arrays.

Returns:

None.

spacr.submodules.post_regression_analysis(csv_file, grna_dict, grna_list, save=False)[source]

Compute gRNA correlation and propagate fixed effect sizes across correlated gRNAs.

Parameters:
  • csv_file – CSV with columns grna, fraction and prc.

  • grna_dict – mapping of anchor grna names to their fixed effect sizes.

  • grna_list – gRNAs to include in the correlation matrix.

  • save – persist correlation matrix, effect sizes and plots. Default False.

Returns:

None. Displays plots and optionally writes results to <csv_dir>/post_regression_analysis_results.

spacr.submodules.split_wells(settings)[source]

Cut every multi-well image under src into one image per well.

Runs the YOLO well detector over each image, writes one crop per well into <src>/wells, and records each crop’s box so _plaque_scale_for() can turn its diameter into a scale later.

Parameters:

settings – the plaque settings dict. Reads src, well_detector_model, well_confidence and well_pad; writes _well_geometry.

Returns:

the folder holding the crops, or src unchanged when detection is off or finds nothing.

WHY THIS IS A SEPARATE PASS rather than a branch inside the segmenter: the two shapes of input differ in what a RESULT ROW MEANS. One field per image gives one row per image; a plate gives one row per well, and the well has to be named or the conditions are pooled into a single meaningless count. Splitting first makes every downstream row a well, whichever shape arrived.

AN IMAGE THE DETECTOR FINDS NOTHING IN IS COPIED INTO THE SPLIT FOLDER WHOLE, with no geometry, so it is still analysed and simply has no ruler. A warning that says “passed through whole” and then skips the image passes nothing through: when some images in a folder split and others did not, the others would contribute no crop and no row and nothing would say they had existed. The all-or-nothing case – where the function falls back to src – is the only one that would behave as the message promised.

spacr.submodules.test_cellpose_model(settings)[source]

Evaluate a Cellpose model on a labelled test set and report per-image metrics.

Computes Jaccard, object counts, mean object area, precision, recall, F1 and accuracy for each image and writes a summary CSV.

Parameters:

settings – dict of test settings; see get_default_test_cellpose_model_settings for keys including src, model_path, batch_size, FT, CP_probability, and save.

Returns:

None. Writes test_results.csv in <src>/results when save is set.

spacr.submodules.train_cellpose(settings)[source]

Fine-tune Cellpose-SAM with native paired images and instance-label masks.

Parameters:

settings – image folder src; optional mask_src (default src/masks), validation test_src/test_mask_src, base_model, model_name, AdamW schedule, channels/channel_axis, normalize/percentiles, scale_range, min_train_masks, optional image limits and checkpoint controls. Legacy project/train/images plus project/train/masks remains accepted when mask_src is blank.

Returns:

Cellpose’s checkpoint path, training losses and validation losses. Weights are written beneath save_path/models (default src/models/cellpose_model/models).

Raises:

ValueError – invalid settings, unpaired images or incompatible masks.

Nested helpers

_cellpose_training_pairs.files(folder)

Ignore metadata sidecars and Cellpose-generated flow caches.

spacr/submodules.py:339

_invasion_field_thresholds._auto(values)

Return a finite-data threshold, or NaN below the object-count floor.

spacr/submodules.py:4482

_resolve_invasion_intensity_column._candidates(name)

Yield the current and any legacy measurement column for name.

spacr/submodules.py:4137

_resolve_invasion_intensity_column._template(name)

Return the requested statistic’s formatted measurement column.

spacr/submodules.py:4132

analyze_endodyogeny._calculate_volume_bins(df, compartment='pathogen', min_area_bin=500, max_bins=None, verbose=False)

Assign each row to a log2 volume-doubling bin and return the ordered categories.

spacr/submodules.py:2749

analyze_percent_positive.annotate_and_summarize(df, value_col, condition_col, well_col, threshold, annotation_col='annotation')

Annotate rows as above/below a threshold and summarise per condition and well.

Parameters:
  • df – measurements DataFrame to annotate in place.

  • value_col – column whose values are compared to threshold.

  • condition_col – experimental condition column used for grouping.

  • well_col – well identifier column used for grouping.

  • threshold – numeric cutoff; values above become above.

  • annotation_col – name of the new annotation column. Default 'annotation'.

Returns:

tuple (df, summary_df) with the annotated rows and a per-(condition, well) counts/fractions table.

spacr/submodules.py:933

analyze_percent_positive.translate_well_in_df(csv_loc)

Return a dataframe read from csv_loc with plateID / well columns split out of Renamed TIFF.

Parameters:

csv_loc – path to a CSV containing a Renamed TIFF column.

Returns:

pandas.DataFrame with parsed plateID and well columns.

spacr/submodules.py:887

apply_cellpose_model.plot_cellpose_result(i, j, results_dir, img, pred, flow)

Render a 4-panel diagnostic (image / pred / flow) for one Cellpose apply result.

Parameters:
  • i – outer image index used in the output filename.

  • j – inner batch index used in the output filename.

  • results_dir – folder where the composite PNG is written.

  • img – source image array.

  • pred – predicted mask array.

  • flow – Cellpose flow field.

spacr/submodules.py:718

compare_reads_to_scores.calculate_grna_fraction_ratio(df, grna1='TGGT1_220950_1', grna2='TGGT1_233460_4')

Compute the per-well read-fraction ratio between two gRNAs.

Parameters:
  • df – dataframe with prc, grna_name, and count columns.

  • grna1 – numerator gRNA.

  • grna2 – denominator gRNA.

Returns:

dataframe with one ratio value per prc.

spacr/submodules.py:2214

compare_reads_to_scores.calculate_well_read_fraction(df, count_column='count')

Compute the per-well fraction of reads for each gRNA.

Parameters:
  • df – dataframe with plateID/rowID/columnID (or prc), grna_name, and a read count column.

  • count_column – name of the read-count column.

Returns:

dataframe with a fraction column per (prc, grna_name).

spacr/submodules.py:2241

compare_reads_to_scores.calculate_well_score_fractions(df, class_columns='cv_predictions')

Aggregate per-object classifier predictions into per-well class fractions.

Parameters:
  • df – measurements dataframe with a prc well id and a classifier prediction column.

  • class_columns – name of the prediction column to summarise.

Returns:

dataframe keyed by prc with one fraction column per class.

spacr/submodules.py:2097

compare_reads_to_scores.plot_line(df, x_column, y_columns, group_column=None, xlabel=None, ylabel=None, title=None, figsize=(10, 6), save_path=None, theme='deep')

Plot one line per y-column (or per group_column value) against x_column.

Parameters:
  • df – DataFrame containing the x and y columns.

  • x_column – column used for the x axis.

  • y_columns – str or list of columns to plot as lines.

  • group_column – optional hue column when y_columns is a single column.

  • xlabel – x-axis label; falls back to x_column.

  • ylabel – y-axis label; falls back to 'Value'.

  • title – plot title; falls back to 'Line Plot'.

  • figsize – figure size in inches. Default (10, 6).

  • save_path – optional PDF path to save the figure.

  • theme – Seaborn palette name. Default 'deep'.

Returns:

the created matplotlib Figure.

spacr/submodules.py:2123

compare_reads_to_scores.plot_line._set_theme(theme)

The colours the lines are drawn in, house palette first.

ONE LINE PER MEASURED COLUMN IS A CASE WHERE THE CATEGORIES REALLY ARE THE DATA, so these series keep distinct hues – but they come from the published palette in a fixed order rather than from seaborn’s 100-colour ‘deep’ ramp reordered by an index list. A hundred hues is a hundred series nobody can tell apart, and the eighth one was a pastel that vanished on the dark ground.

An explicit non-default theme still wins, for a caller who deliberately asked for a seaborn palette.

spacr/submodules.py:2140

generate_score_heatmap.calculate_fraction_mixed_condition(csv, plate=1, column='c3', control_sgrnas=None)

Return per-well read fractions restricted to the given control sgRNAs.

Parameters:
  • csv – path to the reads CSV; needs grna_name, count, rowID and columnID (a legacy column_name is renamed).

  • plate – a plate NUMBER, not a column name. Any value but None rewrites every plateID to plate<plate>; None keeps each row’s own value. Default 1.

  • column – value kept from columnID. Default 'c3'.

  • control_sgrnas – None selects the built-in pair TGGT1_220950_1 / TGGT1_233460_4. Exactly the first two entries are read, so a longer list silently ignores the rest and a shorter one (or an empty one) raises IndexError. Entries are interpolated into an anchored regex: metacharacters are live, and a name that is only a prefix of the real sgRNA matches nothing.

Returns:

the matching rows plus total_count, fraction (that sgRNA’s share of the two controls’ combined count – non-control reads never enter the denominator) and a prc key. Controls absent from the CSV yield an empty frame, not an error.

spacr/submodules.py:5705

generate_score_heatmap.calculate_mae(df)

Return the per-channel, per-row MAE between predictions and the fraction column.

spacr/submodules.py:5848

generate_score_heatmap.combine_classification_scores(folders, csv_name, data_column, plate=1, column='c3')

Merge one data_column per sub-folder into a wide per-well DataFrame.

Parameters:
  • folders – parent directory, or a list of them; a bare string is wrapped in a list. Only the immediate sub-directories are scanned, so a CSV sitting in the parent itself, or one nested two levels down, is never found. A path that does not exist raises FileNotFoundError.

  • csv_name – file name looked for inside each sub-directory. Misses are printed, not raised – finding none leaves the accumulator None and the call ends in TypeError: 'NoneType' object is not subscriptable.

  • data_column – column averaged per well. Its output column is named after the containing sub-directory (<sub-folder>_<data_column>), so two parents holding same-named sub-folders collide and pandas appends _x / _y.

  • plate – a plate NUMBER, not a column name. Any value but None rewrites every plateID to plate<plate>; None keeps each CSV’s own value. Default 1.

  • column – value kept from columnID; a CSV lacking that column raises KeyError. Default 'c3'.

Returns:

the outer-joined well-by-channel frame with a prc key.

spacr/submodules.py:5788

generate_score_heatmap.group_cv_score(csv, plate=1, column='c3', data_column='pred')

Aggregate a CV predictions CSV to a per-(plate, row, column) mean.

Parameters:
  • csv – path to the cross-validation predictions CSV.

  • plate – a plate NUMBER, not a column name. Any value but None rewrites every row’s plateID to plate<plate>, discarding the plate the CSV itself recorded; None keeps the CSV’s own value and raises KeyError if it has no plateID column. Default 1.

  • column – value kept from columnID, or from a legacy column column which is copied to columnID first. A value matching no row returns an empty frame rather than raising; a CSV carrying neither key skips the filter silently and then dies on the groupby with KeyError: 'columnID'. Default 'c3'.

  • data_column – column averaged within each well. A name absent from the CSV raises KeyError, a non-numeric one TypeError. Default 'pred'.

Returns:

one row per well, plus a prc key of plateID_rowID_columnID.

spacr/submodules.py:5676

generate_score_heatmap.plot_multi_channel_heatmap(df, column='c3', cmap='coolwarm')

Plot a per-well heatmap with each classifier channel as a column.

Parameters:
  • df – DataFrame with score columns keyed by channel.

  • column – value in columnID used to filter rows. Default 'c3'.

  • cmap – matplotlib/seaborn colormap. Default 'coolwarm'.

Returns:

the matplotlib Figure.

spacr/submodules.py:5741

interpret_vision_model.create_extended_radar_plot(values, labels, title)

Render a polar radar plot of values against labels.

Parameters:
  • values – numeric values per axis (one per label).

  • labels – axis labels.

  • title – plot title.

spacr/submodules.py:2469

interpret_vision_model.extract_compartment_channel(feature_name)

Split feature_name into (compartment, channel) by the leading underscore token.

Parameters:

feature_name – measurement feature key, e.g. "cell_ch0_mean".

Returns:

two-tuple (compartment, channel) — either may be None.

spacr/submodules.py:2494

interpret_vision_model.generate_comparison_columns(df, compartments=None)

Add cross-compartment feature ratios (e.g. nucleus/cell) as new columns.

Parameters:
  • df – measurements DataFrame; columns prefixed with each compartment.

  • compartments – compartment prefixes to compare. Defaults to ['cell', 'nucleus', 'pathogen', 'cytoplasm'].

Returns:

tuple (df, comparison_dict) with the expanded DataFrame and a mapping of source columns to their derived ratio partners.

spacr/submodules.py:2379

interpret_vision_model.group_feature_class(df, feature_groups=None, name='compartment', include_all=False)

Sum feature importance by compartment or channel group.

Parameters:
  • df – DataFrame with columns feature and importance.

  • feature_groups – substrings identifying each group (compartments or channels).

  • name – name of the grouping column to create. Default 'compartment'.

  • include_all – append an all row summing across groups. Default False.

Returns:

DataFrame of summed importance per group.

spacr/submodules.py:2429

interpret_vision_model.group_feature_class.find_feature_class(feature, compartments)

Return the compartment(s) whose name matches feature.

spacr/submodules.py:2442

interpret_vision_model.read_and_preprocess_data(settings)

Load the measurements DB pointed at by settings and return the merged dataframe.

Parameters:

settings – settings dict; must contain src (folder holding measurements/measurements.db).

Returns:

dataframe of merged object measurements.

spacr/submodules.py:2522

post_regression_analysis._analyze_and_visualize_grna_correlation(df, grna_list, save_folder, save=False)

Return and plot the pivoted per-well gRNA fraction correlation matrix.

spacr/submodules.py:5895

post_regression_analysis._compute_effect_sizes(correlation_matrix, grna_dict, save_folder, save=False)

Return per-gRNA effect sizes propagated from anchor gRNAs via the correlation matrix.

spacr/submodules.py:5928

test_cellpose_model.plot_cellpose_resilts(i, j, results_dir, img, lbl, pred, flow)

Render one 5-panel diagnostic (image / label / pred / flow) for a Cellpose result.

Parameters:
  • i – outer image index used in the output filename.

  • j – inner batch index used in the output filename.

  • results_dir – folder where the composite PNG is written.

  • img – source image array.

  • lbl – ground-truth label array.

  • pred – predicted mask array.

  • flow – Cellpose flow field.

spacr/submodules.py:514