spacr.submodules¶
Workflow inputs and outputs¶
Plaque Assay¶
Analyse plaque images or existing masks with the configured plaque model. This route need not pass through Measure; use the dedicated plaque example.
Open: Toxoplasma → Plaque Assay.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.
Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.
Outputs
Assay results — Assay-specific result tables and figures in the configured destination, preserving well and condition identities.
Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.
Before this module
Make Masks: Use plaque masks with matching source images; cell masks are not automatically plaque labels.
Recruitment¶
Use compartment intensity measurements and matching host/pathogen identities to compute recruitment ratios.
Open: Toxoplasma → Recruitment.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route:
cell,nucleus,pathogen,cytoplasm. Relevant columns, depending on the route:plateID,rowID,columnID,fieldID.
Outputs
Assay results — Assay-specific result tables and figures in the configured destination, preserving well and condition identities.
Before this module
Measure: Require the intended compartment intensities and identities.
Invasion Assay¶
Use the required two-colour differential-staining measurements and stain-baseline controls to distinguish attachment from invasion.
Open: Toxoplasma → Invasion Assay.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route:
cell,nucleus,pathogen,cytoplasm. Relevant columns, depending on the route:plateID,rowID,columnID,fieldID.
Outputs
Assay results — Assay-specific result tables and figures in the configured destination, preserving well and condition identities.
Before this module
Measure: Require two-colour stain measurements and appropriate baseline controls.
Replication Assay¶
Count parasites using explicit vacuole identity and compare condition distributions; host identity alone does not define a vacuole.
Open: Toxoplasma → Replication Assay.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route:
cell,nucleus,pathogen,cytoplasm. Relevant columns, depending on the route:plateID,rowID,columnID,fieldID.
Outputs
Assay results — Assay-specific result tables and figures in the configured destination, preserving well and condition identities.
Before this module
Measure: Require explicit parasite-to-vacuole identities.
Cellpose Workbench¶
Open Cellpose Workbench inside Make Masks and train from verified image/mask pairs. Evaluate on separate fields before selecting the checkpoint in Mask.
Open: Make Masks → Cellpose Workbench.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Curated training fields — Separate image and integer-mask files with matching field identities; preserve original images and labels.
Outputs
Segmentation checkpoint — Saved Cellpose-compatible checkpoint or a compatible installed backend selected with its own configuration.
Before this module
Make Masks: Use independently checked image/mask pairs.
After this module
Mask: Select the saved compatible checkpoint in Mask.
Direct Cellpose mask generation: Pass the trained checkpoint as custom_model with the matching image channels and preprocessing.
Endodyogeny size proxy¶
Read measured compartment areas, annotate conditions and bin area ** 1.5 into log2 size doublings. Defaults aggregate pathogen area per host cell, not per vacuole: multiple vacuoles in one cell are combined. This is an area-derived size proxy, not measured volume or a parasite count. Use Replication Assay with explicit vacuole identity for parasites-per-vacuole counts. Configure compartment, area filters, calibration, conditions and grouping before calling the API; saving is optional.
Use from Python: spacr.submodules.analyze_endodyogeny(). This API-only workflow has no Home tile or menu entry.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Host-cell-aggregated compartment areas — Each src project root/measurements/measurements.db; tables defaults to cell, nucleus, pathogen and cytoplasm, with png_list added for merging. The compartment setting selects the area column; default pathogen_area is summed per host cell. Relevant tables, depending on the route:
cell,nucleus,pathogen,cytoplasm,png_list. Relevant columns, depending on the route:cell_id,pathogen_area.
Outputs
Area-derived size-proxy results — Returned data and chi_squared DataFrames; save=True also writes data.csv, chi_squared_results.csv, chi_squared_pairwise_results.csv and a figure under the first project root/results/analyze_endodyogeny/. Relevant columns, depending on the route:
pathogen_area,pathogen_volume,pathogen_volume_bin,bin_index.
Before this module
Measure: Supply the measured project roots and required object/png_list tables. Verify host-cell aggregation and area units before interpreting size bins; the Mask counts database alone is insufficient.
Run plaque, recruitment, invasion, and replication assays.
WHAT IT IS FOR¶
Four spaCR tiles currently share this landing page, but they answer different
biological questions. analyze_plaques() segments plaque images and
summarizes plaque number and area. analyze_recruitment() measures a
fluorescent marker around pathogens or vacuoles relative to host cytoplasm.
analyze_invasion() uses differential pre/post-permeabilization staining
to classify parasites as attached outside or invaded inside a host cell.
analyze_replication() counts parasites within each parasitophorous
vacuole and compares the resulting replication-state distributions. Cellpose
training, testing, and model-application utilities also live here, but they
are not substitutes for those four assay entry points.
WHAT IT NEEDS¶
Plaque analysis accepts a folder of TIFF images, or existing masks beneath
that folder, plus Cellpose settings and a bundled, catalogue, or local plaque
model. Recruitment starts from a spaCR measurements.db containing joined
cell, nucleus, pathogen, and cytoplasm features; it needs a fluorescence
channel, object filters, and plate metadata that assign cell type, pathogen,
and treatment. Invasion and Replication both need one row per segmented
parasite in a measurement table and condition metadata. Invasion additionally
needs the outside- and total-stain channels and preferably known control wells;
Replication needs a defensible vacuole_key or spatial-linking distance.
The Cellpose utilities require paired images and masks for training/testing,
or an image folder and model path for inference.
WHAT IT PRODUCES¶
Plaque analysis writes <src>/masks/plaques_analysis.db with summary,
stats, and details tables. Recruitment returns per-object and
per-well DataFrames and writes their CSVs and plots. Invasion returns
per-parasite classifications, per-field thresholds and QC, per-well
efficiencies, condition summaries and comparisons, controls, and figures;
saved runs place those artifacts under results/analyze_invasion.
Replication returns per-vacuole counts, well and condition distributions,
pairwise and omnibus statistics, figures, and the grouping method actually
used, with saved output under results/analyze_replication. Model utilities
produce trained weights, evaluation tables, masks, and object summaries as
appropriate.
WHAT TO DO NEXT¶
For plaques, inspect the masks before interpreting counts or areas. For Recruitment, verify the object filters, condition annotation, and per-well denominators before comparing treatments. For Invasion, review field-level thresholds, control agreement, bimodality, and sensitivity flags before using the efficiency table. For Replication, inspect the vacuole grouping and the reported non-power-of-two fraction before comparing doubling distributions. Follow the specific function links above until the four tiles receive separate API destinations.
The analysis unit matters. Invasion is inferred from absence of outside stain, so weak staining can only inflate the invaded fraction; thresholds are therefore recorded per field and statistics use wells rather than treating parasites from one well as independent replicates. Replication groups by vacuole, not by host cell, because one cell can contain several vacuoles; 3, 5, 6, and 7 parasites remain in an explicit non-power-of-two QC bucket instead of being rounded into a biologically expected class. Plaque area is calibrated against the well scale when available, so comparisons should retain the acquisition metadata that defines that scale.
Exceptions¶
A Cellpose 3 checkpoint was handed to Cellpose 4, which cannot load it. |
|
A named model is not where it should be. |
Classes¶
Lazy image/label dataset for Cellpose training and inference. |
Functions¶
|
Test whether classifier class proportions differ between experimental groups. |
|
Bin pathogen size by log2 doublings and test the bin proportions per group. |
|
Invasion assay: score every parasite attached or invaded and report efficiency per well. |
|
Annotate objects above a threshold and summarise positive fractions per well. |
|
Segment host-cell plaques with a bundled Cellpose model and summarize per-image counts and areas. |
|
Measure marker recruitment with host-cell and per-well summaries. |
|
Replication assay: count parasites per vacuole and compare the distributions. |
|
Run a Cellpose model over a folder of images and export per-object measurements. |
|
Compare sequencing read fractions to classifier score fractions across wells. |
|
Count unique phenotype annotations per plate/row/column and export to CSV. |
|
Do nothing: IPython is unavailable, so there is nowhere to display to. |
|
Turn Cellpose 4's refusal of a Cellpose 3 checkpoint into advice. |
|
Combine multiple classifier score CSVs into a per-well heatmap and MAE table. |
|
Explain a spacr vision-model score by ranking which morphology / intensity features drive it. |
|
Display a two-row grid of images and their paired label masks. |
|
Compute gRNA correlation and propagate fixed effect sizes across correlated gRNAs. |
|
Cut every multi-well image under |
|
Evaluate a Cellpose model on a labelled test set and report per-image metrics. |
|
Fine-tune Cellpose-SAM with native paired images and instance-label masks. |
Module Contents¶
- exception spacr.submodules.Cellpose3Checkpoint[source]¶
Bases:
ValueErrorA Cellpose 3 checkpoint was handed to Cellpose 4, which cannot load it.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.submodules.ModelZooMissing[source]¶
Bases:
FileNotFoundErrorA named model is not where it should be.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.submodules.CellposeLazyDataset(image_files, label_files, settings, randomize: bool = True, augment: bool = False)[source]¶
Bases:
torch.utils.data.DatasetLazy image/label dataset for Cellpose training and inference.
Loads paired image and label tiffs on demand, optionally normalizing, augmenting (8-fold rotations/flips), and resizing to a target size.
- Parameters:
image_files – paths to input image tiffs.
label_files – paths to matching label tiffs (same length as
image_files).settings – dict with keys
normalize,percentiles,target_size.randomize – shuffle the image/label pairing order. Default
True.augment – enable 8-fold augmentation (dataset length x8). Default
False.
- Raises:
ValueError – when image/label lists differ in length or are empty.
Pair the image and label files and fix the augmentation factor.
Mismatched lengths raise here rather than at the first bad index, so a wrongly paired dataset fails at construction instead of part-way through an epoch.
- spacr.submodules.analyze_class_proportion(settings)[source]¶
Test whether classifier class proportions differ between experimental groups.
Runs chi-squared and pairwise tests on the class column, plots stacked bars and a plate heatmap, and follows up with normality, Levene, and posthoc statistical tests.
- Parameters:
settings – dict of settings; see
set_analyze_class_proportion_defaultsfor keys includingsrc,tables,class_column,group_column,levelandsave.- Returns:
dict with
data(annotated DataFrame) andchi_squared(results DataFrame).
- spacr.submodules.analyze_endodyogeny(settings)[source]¶
Bin pathogen size by log2 doublings and test the bin proportions per group.
This is the size-proxy replication readout, not a parasite count. Read that sentence twice before quoting a number from it:
The rows come from
spacr.io._read_and_merge_data(), which collapses the per-objectpathogentable onto the host cell (prcfois built fromcell_id).pathogen_areaon each row is therefore the sum of the areas of every pathogen object inside that host cell — one host cell carrying two parasitophorous vacuoles contributes a single row holding the combined area of both.area ** 1.5is a 2-D-to-3-D size proxy, not a measured volume.Nothing here counts parasites. A bin is a doubling of area-derived size, which tracks parasites-per-vacuole only while the pathogen mask segments whole vacuoles and each host cell holds exactly one.
Keep using it when the pathogen channel gives you fused rosettes that cannot be resolved into single parasites. When the individual parasites are resolvable,
analyze_replication()counts them and reports the parasites-per-vacuole distribution directly, which is the readout an endodyogeny experiment is actually after.- Parameters:
settings – dict of endodyogeny settings; see
set_analyze_endodyogeny_defaultsfor keys includingsrc,tables,compartment,min_area_bin,max_area,max_bins,um_per_px,group_column,levelandsave.- Returns:
dict with
data(binned DataFrame) andchi_squared(results DataFrame).
Example
from spacr.submodules import analyze_endodyogeny out = analyze_endodyogeny({'src': '/data/plate1', 'save': True})
See also
analyze_replication()— counts parasites per vacuole instead of inferring replication from object size.
- spacr.submodules.analyze_invasion(settings)[source]¶
Invasion assay: score every parasite attached or invaded and report efficiency per well.
The red/green invasion assay stains twice. Before permeabilisation an antibody reaches only the parasites still outside the host cell, so those are positive in both channels; the cells are then permeabilised and a second antibody stains all parasites, so a parasite positive only in the post-permeabilisation channel was inside. Hence:
attached / outside = present in the outside-stain channel;
invaded / inside = absent from the outside-stain channel.
Read that asymmetry carefully, because the whole design follows from it. “Inside” is defined by an absence, and absence is the unreliable direction. Poor antibody penetration, a focal plane off the parasite’s equator, photobleaching, a low-expressing parasite — every one of them removes outside signal from a parasite that is genuinely outside, and every one of them therefore inflates invasion efficiency. Nothing plausible pushes the error the other way. The threshold on the outside channel is the single number the assay rests on, so it is derived from the data, reported per field in
fields, cross-checked against a control-derived cut when one exists, and bracketed by a sensitivity pair that says how much of the answer is the threshold.Three design decisions worth stating outright:
The threshold is per field. Illumination and staining vary field to field, and a plate-wide cut turns an illumination gradient into an invasion gradient. See
_invasion_field_thresholds().Controls beat any automatic method.
control_wellsnames wells whose parasites are known to carry no outside stain; the threshold is then a high quantile of that honest negative distribution (control_quantile), the control wells are excluded from the results, andthreshold_sourcesays'control'so the report cannot be mistaken for an automatic run.A threshold without two populations is arbitrary. Otsu will happily split a single smear of signal down the middle and return a confident number.
bimodality_coefficientandqc_flag_unimodalsay when that has happened, per field and per well, instead of letting it pass silently. See_bimodality_coefficient().
invasion_efficiency = n_invaded / (n_invaded + n_attached)and it is always reported next ton_total: 90% from ten parasites and 90% from four thousand are not the same result. A well that scored nothing gets NaN, not 0.0.Statistics use the well as the unit of replication. Parasites within a well share a coverslip, an antibody bath and a focal plane, so they are not independent; the reported test is a Mann-Whitney U on the per-well efficiencies. A pooled-parasite chi-squared is reported beside it purely so its inflation is visible. See
_invasion_compare_conditions().- Parameters:
settings –
dict of invasion settings; see
set_analyze_invasion_defaults. Key entries:src— plate directory (or list) holdingmeasurements/measurements.db.parasite_table/compartment— table and column prefix with one row per segmented parasite. Default'pathogen'.outside_channel/total_channel— the pre- and post-permeabilisation stain channels.intensity_statistic— which per-object statistic of the outside channel to threshold;'auto'prefers the boundary-restricted one. See_resolve_invasion_intensity_column().background_correction— optional per-object local background.outside_threshold_method/outside_threshold— automatic method, or a fixed cut that overrides it.control_wells/control_quantile/min_control_objects.min_objects_for_threshold/min_objects_for_bimodality/bimodality_cutoff/threshold_agreement_tolerance/threshold_sensitivity/inflation_warn/min_parasites_per_well— the QC thresholds.extracellular_class— how parasites with no host cell are scored.cell_types/pathogen_types/treatmentsand their*_plate_metadatawell maps, plusgroup_columnandlevel.save— write the CSVs and figures under<src>/results/analyze_invasion.
- Returns:
dict with
parasites(per-object classification),fields(per-field thresholds and QC),wells(per-well efficiency, denominators and QC flags),summary(per condition),comparisons(per-well statistics),chi_squared/chi_squared_pairwise(the shared proportion-bar omnibus tests),controls(the control-well objects, if any),control_thresholds,intensity_column,intensity_statisticandfigures.- Raises:
ValueError – when the parasite table holds no usable rows.
KeyError – when the requested statistic or group column is absent.
Example
from spacr.submodules import analyze_invasion out = analyze_invasion({ 'src': '/data/plate1', 'outside_channel': 1, 'total_channel': 0, 'stain_baseline_wells': ['c12'], 'pathogen_types': ['dmso', 'inhibitor'], 'pathogen_plate_metadata': [['c1'], ['c2']], }) print(out['wells'][['prc', 'n_total', 'invasion_efficiency', 'qc_flags']])
See also
analyze_replication()— the parasites-per-vacuole assay, whose table reading, condition annotation and output layout this follows.
- spacr.submodules.analyze_percent_positive(settings)[source]¶
Annotate objects above a threshold and summarise positive fractions per well.
Merges measurements from
measurements.db, thresholds on a chosen feature column, then joins the resulting well-level counts againstrename_log.csvto recover human-readable plate/well identifiers.- Parameters:
settings – dict of settings; see
default_settings_analyze_percent_positivefor keys includingsrc,tables,value_col,thresholdandfilter_1.- Returns:
DataFrame of annotated per-well positive/negative counts and fractions.
- spacr.submodules.analyze_plaques(settings)[source]¶
Segment host-cell plaques with a bundled Cellpose model and summarize per-image counts and areas.
Downloads (if needed) the bundled
toxo_plaque_cyto_e25000model, runs Cellpose over every.tifundersrc, then computes per-image plaque count + mean/stddev area and writes aplaques_analysis.db(tables:summary,stats,details) alongside the masks.- Parameters:
settings –
Settings dict, canonicalized via
spacr.settings.get_analyze_plaque_settings(). Key entries:src— folder containing plaque images.masks— if truthy, run segmentation before analysis; if falsy, expect masks already in<src>/masks.diameter,flow_thresholdandCP_prob, read byspacr.plaque.segment_plaque_image(), the call the Plaque preview makes too.plaque_mode–'figure'hands the folder tospacr.plaque_papers.measure_figure_folder()instead.colony_counting– in plaque mode, counts bacterial or fungal colonies on plate photos instead of segmenting plaques (_analyze_colony_plates()), writing<src>/colonies/colonies.db.
- Returns:
None. Writes
<src>/masks/plaques_analysis.db. Withcolony_countingit returns the per-plate colony table instead.
Example
from spacr.submodules import analyze_plaques analyze_plaques({'src': '/data/plaque_assay', 'masks': True})
See also
analyze_recruitment()— intensity-ratio phenotype instead of plaque counts.
- spacr.submodules.analyze_recruitment(settings)[source]¶
Measure marker recruitment with host-cell and per-well summaries.
Reads the merged cell/nucleus/pathogen/cytoplasm feature tables from a spacr
measurements.db, annotates each row with cell type / pathogen / treatment based on plate metadata, filters objects by size and intensity, computes the pathogen-to-cytoplasm mean-intensity ratio forchannel_of_interest, groups by well and writes bothresults/cells.csvandresults/wells.csvalongside recruitment plots.Each cell row combines the pathogen measurements assigned to that host cell. Pathogen mean intensities are averaged across its associated objects; these rows represent host cells rather than independently measured vacuoles. The main recruitment ratio divides that aggregate pathogen mean by the cell’s cytoplasm mean. Each well averages its retained cell ratios.
In the GUI, open Home > Toxoplasma > Recruitment. Select a measured project, map its channels and plate conditions, review the object filters, and Run. Inspect the retained counts and ratio columns before comparing conditions. Condition plots show between-well standard deviations. For measurements linked to individual vacuoles, use
spacr.host_pathogen.- Parameters:
settings –
Settings dict, canonicalized via
spacr.settings.get_analyze_recruitment_default_settings(). Key entries:src— folder containingmeasurements/measurements.dband optionalmergedimages for overlays. A database path is also accepted; a database outside a measurements folder may be moved into one, so use a project copy when reorganizing existing data.cell_types/cell_plate_metadata— labels + row/col metadata that map wells to cell lines.pathogen_types/pathogen_plate_metadata.treatments/treatment_plate_metadata.channel_of_interest— intensity channel for the ratio.cell_chann_dim/nucleus_chann_dim/pathogen_chann_dim— recorded object-channel mapping used by image overlays and intensity filtering.cell_size_range,nucleus_size_range,pathogen_size_range—[min, max]px area filters.*_intensity_range,target_intensity_min.cells_per_well— minimum well count to keep.plot,plot_control,plot_nr,figuresize.
- Returns:
List
[cells, wells]— the host-cell and per-well recruitment DataFrames, also written to CSV undersrc/results.
Example
from spacr.submodules import analyze_recruitment settings = { 'src': '/data/plate01', 'cell_types': ['HeLa'], 'cell_plate_metadata': ['c2-c11'], 'pathogen_types': ['tgme49'], 'pathogen_plate_metadata': ['c2-c11'], 'treatments': ['dmso','drug'], 'treatment_plate_metadata': [['r1'],['r2']], 'channel_of_interest': 3, } cells_df, wells_df = analyze_recruitment(settings)
See also
analyze_plaques()— plaque-count/size assay.spacr.ml.generate_ml_scores()— feature-based classifier as an alternative to recruitment ratios.
- spacr.submodules.analyze_replication(settings)[source]¶
Replication assay: count parasites per vacuole and compare the distributions.
replication_method='direct_count'is the default described below.'size_proxy'delegates toanalyze_endodyogeny()and returns its area-derived, host-aggregated readout instead. Both return the selected method inreplication_method.'deep_learning_coming_soon'raisesNotImplementedErrorbefore any data are read or outputs written; the whole-vacuole classification model is not available yet.Toxoplasma gondii replicates by endodyogeny, two daughters forming inside a mother, so a parasitophorous vacuole holds 1, 2, 4, 8 or 16 parasites — a power of two. The readout of a replication assay is therefore the distribution of parasites-per-vacuole across a well, not a mean: a mean of 3.2 cannot distinguish “everything at 3-ish”, which is biologically impossible, from a healthy mix of 2s and 4s. A drug that slows replication moves mass from the 8 and 4 buckets down into 2 and 1, and only the distribution shows that.
The counting unit is the vacuole. Not the parasite, and emphatically not the host cell — one host cell routinely carries several vacuoles, so grouping on
cell_idreports their combined count as a single vacuole and produces a plausible but meaningless number. See_assign_vacuole_ids()for how the vacuole is derived and what eachvacuole_keycosts you.Rosettes of 3, 5, 6 or 7 are counted into an explicit
non_power_of_twobucket that is always reported and never folded into a neighbouring bucket. That bucket is the assay’s own quality control: a well where 30% of vacuoles are off the power-of-two ladder has a segmentation problem, and its replication number should not be trusted.Statistics: the two-condition comparison is a Mann-Whitney U test on the doubling index
log2(n_parasites), with a chi-squared omnibus test alongside it._replication_compare_conditions()explains why, and why a t-test on the raw counts is the wrong instrument.- Parameters:
settings –
dict of replication settings; see
set_analyze_replication_defaults. Key entries:src— plate directory (or list of them) holdingmeasurements/measurements.db.parasite_table/compartment— table and column prefix holding one row per segmented parasite. Default'pathogen'.vacuole_key— how parasite rows are grouped into vacuoles ('auto','spatial','cell_id','object', or a column name).vacuole_link_distance/vacuole_link_factor— the spatial clustering threshold, or the multiplier used to derive it.min_parasite_area/max_parasite_area— debris and merged-clump filters applied before counting.max_parasites_per_vacuole— largest named power-of-two bucket.non_power_of_two_warn— QC flag threshold.cell_types/pathogen_types/treatmentsand their*_plate_metadatawell maps, plusgroup_columnandlevel.save— write the CSVs and figures under<src>/results/analyze_replication.
- Returns:
dict with
vacuoles(per-vacuole counts),wells(per-well distribution),summary(per-condition distribution),comparisons(pairwise ordered tests),chi_squared/chi_squared_pairwise(omnibus proportion tests),figuresandvacuole_key(the grouping actually used).- Raises:
ValueError – when the parasite table holds no usable rows.
Example
from spacr.submodules import analyze_replication out = analyze_replication({ 'src': '/data/plate1', 'pathogen_types': ['dmso', 'pyrimethamine'], 'pathogen_plate_metadata': [['c1'], ['c2']], }) print(out['summary'][['condition', 'frac_1', 'frac_2', 'frac_4', 'frac_8', 'frac_non_power_of_two']])
See also
analyze_endodyogeny()— the size-proxy version, for fused rosettes that cannot be resolved into single parasites.
- spacr.submodules.apply_cellpose_model(settings)[source]¶
Run a Cellpose model over a folder of images and export per-object measurements.
Optionally masks predictions to a central circle, then records per-object area to
measurements.csvand a per-image summary tosummary.csv.- Parameters:
settings – dict of inference settings; see
get_default_apply_cellpose_model_settingsfor keys includingsrc,model_path,batch_size,FT,CP_probability,circularizeandsave.- Returns:
None. Writes result CSVs under
<src>/results.
- spacr.submodules.compare_reads_to_scores(reads_csv, scores_csv, empirical_dict=None, pc_grna='TGGT1_220950_1', nc_grna='TGGT1_233460_4', y_columns=None, column='columnID', value='c3', plate=None, save_paths=None)[source]¶
Compare sequencing read fractions to classifier score fractions across wells.
Loads paired reads and scores tables (single files or matched lists), computes per-well class-1 and gRNA fractions, joins them with an empirical row-to-mixture dictionary, and plots the fractions against the positive- and negative-control fractions.
- Parameters:
reads_csv – path (or list of paths) to per-gRNA read count CSVs.
scores_csv – path (or list of paths) to per-object classifier score CSVs.
empirical_dict – mapping of
rowIDto(pc_units, nc_units)mixture; a 16-row default is used whenNone.pc_grna – positive-control gRNA name. Default
'TGGT1_220950_1'.nc_grna – negative-control gRNA name. Default
'TGGT1_233460_4'.y_columns – Columns to plot on the y axis.
Noneuses['class_1_fraction', 'TGGT1_220950_1_fraction', 'nc_fraction'].column – column used to select a subset of wells. Default
'columnID'.value – value in
columnto keep. Default'c3'.plate – plate ID to stamp when a single pair of CSVs is given.
save_paths – two-element list of PDF output paths (pc plot, nc plot).
- Returns:
two matplotlib figures
[fig_pc, fig_nc].
- spacr.submodules.count_phenotypes(settings)[source]¶
Count unique phenotype annotations per plate/row/column and export to CSV.
- Parameters:
settings – dict with
src(pointing at a measurements folder ormeasurements.db) andannotation_column(the column of interest in thepng_listtable).- Returns:
None. Writes
phenotype_counts.csvnext to the database.
- spacr.submodules.display(*args, **kwargs)[source]¶
Do nothing: IPython is unavailable, so there is nowhere to display to.
THE FALLBACK IS THE POINT.
IPython.display.displayis imported at module scope, and IPython can be mid-init – partially imported by another thread – which makes that import raise. Letting it propagate would make importing this module fail for a reason that has nothing to do with what the module does. spaCR only callsdisplayfrom notebook contexts; the Qt GUI ignores it.- Parameters:
args – whatever the caller would have displayed.
kwargs – likewise.
- spacr.submodules.explain_cellpose3(exc, model)[source]¶
Turn Cellpose 4’s refusal of a Cellpose 3 checkpoint into advice.
- Parameters:
exc – the exception Cellpose raised.
model – what was asked for, for the message.
- Returns:
a
Cellpose3Checkpointwhenexcis that refusal, elseexcunchanged.
- spacr.submodules.generate_score_heatmap(settings)[source]¶
Combine multiple classifier score CSVs into a per-well heatmap and MAE table.
Aggregates per-object scores across score CSVs, merges with a cross-validation score and a reads-derived fraction column, plots a multi-channel heatmap, and computes per-channel mean absolute error against the empirical fraction.
- Parameters:
settings – dict of settings including
folders,csv_name,data_column,csv,cv_csv,data_column_cv,plateID,columnID,control_sgrnas,fraction_grna,cmapanddst.- Returns:
merged DataFrame joining reads, classifier scores and CV scores per well.
- spacr.submodules.interpret_vision_model(settings=None)[source]¶
Explain a spacr vision-model score by ranking which morphology / intensity features drive it.
Joins the per-object CNN predictions (
score_column) with the morphology + intensity measurements fromspacr.measure.measure_crop(), expands cross-compartment feature ratios (e.g.nucleus_cell_area), then runs random-forest feature importance, permutation importance and (optionally) SHAP on the top features. Also groups importance by compartment and by channel so you can answer “is my classifier looking at the pathogen or at the cell?”.- Parameters:
settings –
Settings dict. Key entries:
src— folder containingmeasurements/measurements.dbwith both feature and score tables.tables— DB tables to merge, e.g.['cell','nucleus','pathogen','cytoplasm'].channels— intensity channels included in the feature space (e.g.[0,1,2,3]).score_column— column holding per-object CNN scores.top_features— cap on features shown / SHAP-explained.feature_importance/permutation_importance/shap— toggle each explainer.shap_sample— subsample size for SHAP.nuclei_limit/pathogen_limit— object-count caps in the read/merge step.n_jobs,save.
- Returns:
Dict of DataFrames keyed by analysis name (
'feature_importance','permutation_importance','shap','compartment_importance','channel_importance', …).
Example
from spacr.submodules import interpret_vision_model results = interpret_vision_model({ 'src': '/data/plate01', 'score_column': 'pred', 'channels': [0,1,2,3], 'top_features': 30, 'shap': True, })
See also
spacr.deep_spacr.deep_spacr()— trains the model whose scores this function interprets.
- spacr.submodules.plot_cellpose_batch(images, labels)[source]¶
Display a two-row grid of images and their paired label masks.
- Parameters:
images – iterable of 2D grayscale image arrays.
labels – iterable of matching integer label arrays.
- Returns:
None.
- spacr.submodules.post_regression_analysis(csv_file, grna_dict, grna_list, save=False)[source]¶
Compute gRNA correlation and propagate fixed effect sizes across correlated gRNAs.
- Parameters:
csv_file – CSV with columns
grna,fractionandprc.grna_dict – mapping of anchor
grnanames to their fixed effect sizes.grna_list – gRNAs to include in the correlation matrix.
save – persist correlation matrix, effect sizes and plots. Default
False.
- Returns:
None. Displays plots and optionally writes results to
<csv_dir>/post_regression_analysis_results.
- spacr.submodules.split_wells(settings)[source]¶
Cut every multi-well image under
srcinto one image per well.Runs the YOLO well detector over each image, writes one crop per well into
<src>/wells, and records each crop’s box so_plaque_scale_for()can turn its diameter into a scale later.- Parameters:
settings – the plaque settings dict. Reads
src,well_detector_model,well_confidenceandwell_pad; writes_well_geometry.- Returns:
the folder holding the crops, or
srcunchanged when detection is off or finds nothing.
WHY THIS IS A SEPARATE PASS rather than a branch inside the segmenter: the two shapes of input differ in what a RESULT ROW MEANS. One field per image gives one row per image; a plate gives one row per well, and the well has to be named or the conditions are pooled into a single meaningless count. Splitting first makes every downstream row a well, whichever shape arrived.
AN IMAGE THE DETECTOR FINDS NOTHING IN IS COPIED INTO THE SPLIT FOLDER WHOLE, with no geometry, so it is still analysed and simply has no ruler. A warning that says “passed through whole” and then skips the image passes nothing through: when some images in a folder split and others did not, the others would contribute no crop and no row and nothing would say they had existed. The all-or-nothing case – where the function falls back to
src– is the only one that would behave as the message promised.
- spacr.submodules.test_cellpose_model(settings)[source]¶
Evaluate a Cellpose model on a labelled test set and report per-image metrics.
Computes Jaccard, object counts, mean object area, precision, recall, F1 and accuracy for each image and writes a summary CSV.
- Parameters:
settings – dict of test settings; see
get_default_test_cellpose_model_settingsfor keys includingsrc,model_path,batch_size,FT,CP_probability, andsave.- Returns:
None. Writes
test_results.csvin<src>/resultswhensaveis set.
- spacr.submodules.train_cellpose(settings)[source]¶
Fine-tune Cellpose-SAM with native paired images and instance-label masks.
- Parameters:
settings – image folder src; optional mask_src (default src/masks), validation test_src/test_mask_src, base_model, model_name, AdamW schedule, channels/channel_axis, normalize/percentiles, scale_range, min_train_masks, optional image limits and checkpoint controls. Legacy project/train/images plus project/train/masks remains accepted when mask_src is blank.
- Returns:
Cellpose’s checkpoint path, training losses and validation losses. Weights are written beneath save_path/models (default src/models/cellpose_model/models).
- Raises:
ValueError – invalid settings, unpaired images or incompatible masks.
Nested helpers¶
- _cellpose_training_pairs.files(folder)¶
Ignore metadata sidecars and Cellpose-generated flow caches.
spacr/submodules.py:339
- _invasion_field_thresholds._auto(values)¶
Return a finite-data threshold, or NaN below the object-count floor.
spacr/submodules.py:4482
- _resolve_invasion_intensity_column._candidates(name)¶
Yield the current and any legacy measurement column for
name.spacr/submodules.py:4137
- _resolve_invasion_intensity_column._template(name)¶
Return the requested statistic’s formatted measurement column.
spacr/submodules.py:4132
- analyze_endodyogeny._calculate_volume_bins(df, compartment='pathogen', min_area_bin=500, max_bins=None, verbose=False)¶
Assign each row to a log2 volume-doubling bin and return the ordered categories.
spacr/submodules.py:2749
- analyze_percent_positive.annotate_and_summarize(df, value_col, condition_col, well_col, threshold, annotation_col='annotation')¶
Annotate rows as
above/belowa threshold and summarise per condition and well.- Parameters:
df – measurements DataFrame to annotate in place.
value_col – column whose values are compared to
threshold.condition_col – experimental condition column used for grouping.
well_col – well identifier column used for grouping.
threshold – numeric cutoff; values above become
above.annotation_col – name of the new annotation column. Default
'annotation'.
- Returns:
tuple
(df, summary_df)with the annotated rows and a per-(condition, well) counts/fractions table.
spacr/submodules.py:933
- analyze_percent_positive.translate_well_in_df(csv_loc)¶
Return a dataframe read from
csv_locwithplateID/wellcolumns split out ofRenamed TIFF.- Parameters:
csv_loc – path to a CSV containing a
Renamed TIFFcolumn.- Returns:
pandas.DataFramewith parsedplateIDandwellcolumns.
spacr/submodules.py:887
- apply_cellpose_model.plot_cellpose_result(i, j, results_dir, img, pred, flow)¶
Render a 4-panel diagnostic (image / pred / flow) for one Cellpose apply result.
- Parameters:
i – outer image index used in the output filename.
j – inner batch index used in the output filename.
results_dir – folder where the composite PNG is written.
img – source image array.
pred – predicted mask array.
flow – Cellpose flow field.
spacr/submodules.py:718
- compare_reads_to_scores.calculate_grna_fraction_ratio(df, grna1='TGGT1_220950_1', grna2='TGGT1_233460_4')¶
Compute the per-well read-fraction ratio between two gRNAs.
- Parameters:
df – dataframe with
prc,grna_name, andcountcolumns.grna1 – numerator gRNA.
grna2 – denominator gRNA.
- Returns:
dataframe with one ratio value per
prc.
spacr/submodules.py:2214
- compare_reads_to_scores.calculate_well_read_fraction(df, count_column='count')¶
Compute the per-well fraction of reads for each gRNA.
- Parameters:
df – dataframe with
plateID/rowID/columnID(orprc),grna_name, and a read count column.count_column – name of the read-count column.
- Returns:
dataframe with a
fractioncolumn per(prc, grna_name).
spacr/submodules.py:2241
- compare_reads_to_scores.calculate_well_score_fractions(df, class_columns='cv_predictions')¶
Aggregate per-object classifier predictions into per-well class fractions.
- Parameters:
df – measurements dataframe with a
prcwell id and a classifier prediction column.class_columns – name of the prediction column to summarise.
- Returns:
dataframe keyed by
prcwith one fraction column per class.
spacr/submodules.py:2097
- compare_reads_to_scores.plot_line(df, x_column, y_columns, group_column=None, xlabel=None, ylabel=None, title=None, figsize=(10, 6), save_path=None, theme='deep')¶
Plot one line per y-column (or per
group_columnvalue) againstx_column.- Parameters:
df – DataFrame containing the x and y columns.
x_column – column used for the x axis.
y_columns – str or list of columns to plot as lines.
group_column – optional hue column when
y_columnsis a single column.xlabel – x-axis label; falls back to
x_column.ylabel – y-axis label; falls back to
'Value'.title – plot title; falls back to
'Line Plot'.figsize – figure size in inches. Default
(10, 6).save_path – optional PDF path to save the figure.
theme – Seaborn palette name. Default
'deep'.
- Returns:
the created matplotlib Figure.
spacr/submodules.py:2123
- compare_reads_to_scores.plot_line._set_theme(theme)¶
The colours the lines are drawn in, house palette first.
ONE LINE PER MEASURED COLUMN IS A CASE WHERE THE CATEGORIES REALLY ARE THE DATA, so these series keep distinct hues – but they come from the published palette in a fixed order rather than from seaborn’s 100-colour ‘deep’ ramp reordered by an index list. A hundred hues is a hundred series nobody can tell apart, and the eighth one was a pastel that vanished on the dark ground.
An explicit non-default
themestill wins, for a caller who deliberately asked for a seaborn palette.spacr/submodules.py:2140
- generate_score_heatmap.calculate_fraction_mixed_condition(csv, plate=1, column='c3', control_sgrnas=None)¶
Return per-well read fractions restricted to the given control sgRNAs.
- Parameters:
csv – path to the reads CSV; needs
grna_name,count,rowIDandcolumnID(a legacycolumn_nameis renamed).plate – a plate NUMBER, not a column name. Any value but
Nonerewrites everyplateIDtoplate<plate>;Nonekeeps each row’s own value. Default1.column – value kept from
columnID. Default'c3'.control_sgrnas –
Noneselects the built-in pairTGGT1_220950_1/TGGT1_233460_4. Exactly the first two entries are read, so a longer list silently ignores the rest and a shorter one (or an empty one) raisesIndexError. Entries are interpolated into an anchored regex: metacharacters are live, and a name that is only a prefix of the real sgRNA matches nothing.
- Returns:
the matching rows plus
total_count,fraction(that sgRNA’s share of the two controls’ combined count – non-control reads never enter the denominator) and aprckey. Controls absent from the CSV yield an empty frame, not an error.
spacr/submodules.py:5705
- generate_score_heatmap.calculate_mae(df)¶
Return the per-channel, per-row MAE between predictions and the
fractioncolumn.spacr/submodules.py:5848
- generate_score_heatmap.combine_classification_scores(folders, csv_name, data_column, plate=1, column='c3')¶
Merge one
data_columnper sub-folder into a wide per-well DataFrame.- Parameters:
folders – parent directory, or a list of them; a bare string is wrapped in a list. Only the immediate sub-directories are scanned, so a CSV sitting in the parent itself, or one nested two levels down, is never found. A path that does not exist raises
FileNotFoundError.csv_name – file name looked for inside each sub-directory. Misses are printed, not raised – finding none leaves the accumulator
Noneand the call ends inTypeError: 'NoneType' object is not subscriptable.data_column – column averaged per well. Its output column is named after the containing sub-directory (
<sub-folder>_<data_column>), so two parents holding same-named sub-folders collide and pandas appends_x/_y.plate – a plate NUMBER, not a column name. Any value but
Nonerewrites everyplateIDtoplate<plate>;Nonekeeps each CSV’s own value. Default1.column – value kept from
columnID; a CSV lacking that column raisesKeyError. Default'c3'.
- Returns:
the outer-joined well-by-channel frame with a
prckey.
spacr/submodules.py:5788
- generate_score_heatmap.group_cv_score(csv, plate=1, column='c3', data_column='pred')¶
Aggregate a CV predictions CSV to a per-(plate, row, column) mean.
- Parameters:
csv – path to the cross-validation predictions CSV.
plate – a plate NUMBER, not a column name. Any value but
Nonerewrites every row’splateIDtoplate<plate>, discarding the plate the CSV itself recorded;Nonekeeps the CSV’s own value and raisesKeyErrorif it has noplateIDcolumn. Default1.column – value kept from
columnID, or from a legacycolumncolumn which is copied tocolumnIDfirst. A value matching no row returns an empty frame rather than raising; a CSV carrying neither key skips the filter silently and then dies on the groupby withKeyError: 'columnID'. Default'c3'.data_column – column averaged within each well. A name absent from the CSV raises
KeyError, a non-numeric oneTypeError. Default'pred'.
- Returns:
one row per well, plus a
prckey ofplateID_rowID_columnID.
spacr/submodules.py:5676
- generate_score_heatmap.plot_multi_channel_heatmap(df, column='c3', cmap='coolwarm')¶
Plot a per-well heatmap with each classifier channel as a column.
- Parameters:
df – DataFrame with score columns keyed by channel.
column – value in
columnIDused to filter rows. Default'c3'.cmap – matplotlib/seaborn colormap. Default
'coolwarm'.
- Returns:
the matplotlib Figure.
spacr/submodules.py:5741
- interpret_vision_model.create_extended_radar_plot(values, labels, title)¶
Render a polar radar plot of
valuesagainstlabels.- Parameters:
values – numeric values per axis (one per label).
labels – axis labels.
title – plot title.
spacr/submodules.py:2469
- interpret_vision_model.extract_compartment_channel(feature_name)¶
Split
feature_nameinto(compartment, channel)by the leading underscore token.- Parameters:
feature_name – measurement feature key, e.g.
"cell_ch0_mean".- Returns:
two-tuple
(compartment, channel)— either may beNone.
spacr/submodules.py:2494
- interpret_vision_model.generate_comparison_columns(df, compartments=None)¶
Add cross-compartment feature ratios (e.g. nucleus/cell) as new columns.
- Parameters:
df – measurements DataFrame; columns prefixed with each compartment.
compartments – compartment prefixes to compare. Defaults to
['cell', 'nucleus', 'pathogen', 'cytoplasm'].
- Returns:
tuple
(df, comparison_dict)with the expanded DataFrame and a mapping of source columns to their derived ratio partners.
spacr/submodules.py:2379
- interpret_vision_model.group_feature_class(df, feature_groups=None, name='compartment', include_all=False)¶
Sum feature importance by compartment or channel group.
- Parameters:
df – DataFrame with columns
featureandimportance.feature_groups – substrings identifying each group (compartments or channels).
name – name of the grouping column to create. Default
'compartment'.include_all – append an
allrow summing across groups. DefaultFalse.
- Returns:
DataFrame of summed importance per group.
spacr/submodules.py:2429
- interpret_vision_model.group_feature_class.find_feature_class(feature, compartments)¶
Return the compartment(s) whose name matches
feature.spacr/submodules.py:2442
- interpret_vision_model.read_and_preprocess_data(settings)¶
Load the measurements DB pointed at by
settingsand return the merged dataframe.- Parameters:
settings – settings dict; must contain
src(folder holdingmeasurements/measurements.db).- Returns:
dataframe of merged object measurements.
spacr/submodules.py:2522
- post_regression_analysis._analyze_and_visualize_grna_correlation(df, grna_list, save_folder, save=False)¶
Return and plot the pivoted per-well gRNA fraction correlation matrix.
spacr/submodules.py:5895
- post_regression_analysis._compute_effect_sizes(correlation_matrix, grna_dict, save_folder, save=False)¶
Return per-gRNA effect sizes propagated from anchor gRNAs via the correlation matrix.
spacr/submodules.py:5928
- test_cellpose_model.plot_cellpose_resilts(i, j, results_dir, img, lbl, pred, flow)¶
Render one 5-panel diagnostic (image / label / pred / flow) for a Cellpose result.
- Parameters:
i – outer image index used in the output filename.
j – inner batch index used in the output filename.
results_dir – folder where the composite PNG is written.
img – source image array.
lbl – ground-truth label array.
pred – predicted mask array.
flow – Cellpose flow field.
spacr/submodules.py:514