spacr.plot

Scientific plotting and statistical-annotation helpers.

Classes

spacrGraph

Grouped plot + statistical-test helper for spacr experiment DataFrames.

Functions

create_grouped_plot(df, grouping_column, data_column)

Plot grouped observations and run assumption-aware comparisons.

create_venn_diagram(file1, file2[, gene_column, ...])

Compute a two-set gene overlap from CSVs and draw its Venn diagram.

data_colours(fig)

Every colour in fig that carries the CLAIM rather than the frame.

deliverable_dpi(fig, dpi[, path])

The DPI this figure can actually be written at, and a word if it is not

figure_output_preferences()

Return (format, dpi) from the user's preferences.

figure_path(path[, fmt])

path with the extension the figure-format preference will write.

generate_mask_random_cmap(mask)

Return a random ListedColormap sized to the labels in mask.

generate_plate_heatmap(df, plate_number, variable, ...)

Aggregate a well-level DataFrame into a plate-shaped heatmap.

graph_importance(settings)

Concatenate feature-importance CSVs and hand off to spacrGraph for plotting.

illegible_data_colours(fig, ground[, floor])

The data colours a reader will not find on ground, as hex.

jitterplot_by_annotation(src, x_column, y_column[, ...])

Read measurements + annotation from a spacr DB and plot a class-balanced jitter plot.

normalize_and_visualize(image, normalized_image[, title])

Show the original and the normalised image side by side in grayscale.

outline_palette_colours(palette)

The four outline colours for palette.

overlay_masks_on_images(img_folder[, normalize, ...])

Overlay masks/* outlines onto matching images from img_folder.

plot_arrays(src[, figuresize, cmap, nr, normalize, q1, q2])

Plot random .npy / .npz arrays from src, one channel per subplot.

plot_cellpose4_output(batch, masks, flows[, cmap, ...])

Display per-channel images, label mask and flow field for Cellpose v4 outputs.

plot_comparison_results(comparison_results)

Plot Jaccard, Dice, boundary-F1 and average-precision distributions per comparison.

plot_data_from_csv(settings)

Load per-plate CSVs, filter/outlier-clean and render a spacrGraph plot.

plot_data_from_db(settings)

Read one or more measurement DBs, annotate conditions and render a spacrGraph plot.

plot_feature_importance(feature_importance_df[, title])

Plot a horizontal bar chart of raw feature importances.

plot_histogram(df, column[, dst])

Plot a histogram of df[column] and optionally save it as PDF.

plot_image_grid(image_paths, percentiles)

Render a square grid of percentile-normalised images with a black background.

plot_image_mask_overlay(file, channels, cell_channel, ...)

Plot image and mask overlays.

plot_image_mask_overlay_magenta_outlines(file, ...[, ...])

Plot image and mask overlays, outlining each channel's own mask in magenta.

plot_images_and_arrays(folders[, lower_percentile, ...])

Show side-by-side images and arrays found across multiple folders.

plot_lorenz_curves(csv_files[, name_column, ...])

Overlay Lorenz curves from multiple gRNA count CSVs with per-plate Gini coefficients.

plot_masks(batch, masks, flows[, cmap, figuresize, ...])

Display per-channel images, label masks and flow fields for a batch.

plot_merged(src, settings)

Show multi-channel image stacks with per-object outlines overlaid.

plot_object_outlines(src[, objects, channels, max_nr])

Overlay mask outlines on the matching channel image for each object type.

plot_organelle_output(img_batch, masks, settings[, ...])

Plot organelle segmentation results: raw channel, label mask, morphology-specific diagnostic.

plot_permutation(permutation_df)

Plot a horizontal bar chart of permutation feature importances with error bars.

plot_plates(df, variable, grouping, min_max, cmap[, ...])

Render every plate of a screen as ONE panel, wells square, on one colour scale.

plot_proportion_stacked_bars(settings, df, ...[, ...])

Plot stacked proportion bars per group with chi-squared and pairwise stats.

plot_region(settings)

Render mask overlay, cropped PNG grid and activation-map grid for one FOV.

plot_resize(images, resized_images, labels, resized_labels)

Show original vs. resized image/label pairs in a 2x2 grid.

print_mask_and_flows(stack, mask, flows[, overlay, ...])

Show a single image, its label mask (optionally outlined) and flow image.

print_ready(fig[, mode, announce])

Repaint fig's chrome for paper for the length of the block.

proportion_mixed_model(df, group_column, bin_column, ...)

A binomial GLM on the per-object outcome, standard errors clustered by unit.

proportion_test_by_unit(df, group_column, bin_column, ...)

Compare conditions on their PER-UNIT proportions, one row per bin.

proportions_per_unit(df, group_column, bin_column, ...)

Each unit's share of every bin, one row per unit.

random_cmap([num_objects])

Return a random ListedColormap with num_objects + 1 colours.

read_and_plot__vision_results(base_dir[, y_axis, ...])

Aggregate vision-model test CSVs under base_dir and plot mean score per model.

save_figure(fig, path, *[, fmt, dpi, close, ...])

Write fig to path, honouring the figure preferences.

visualize_cellpose_masks(masks[, titles, filename, ...])

Display several Cellpose-style label masks side by side for a quick visual QC.

visualize_masks(mask1, mask2, mask3[, title])

Show three masks side by side with random colormaps.

volcano_plot(, title, xlim, float]] = None, ylim, ...)

Read a table (CSV/TSV/XLS/XLSX or a DataFrame) and render a volcano plot.

Module Contents

class spacr.plot.spacrGraph(df, grouping_column, data_column, graph_type='jitter_box', summary_func='mean', order=None, colors=None, output_dir='./output', save=False, y_lim=None, log_y=False, log_x=False, error_bar_type='std', remove_outliers=False, theme='pastel', representation='object', paired=False, all_to_all=True, compare_group=None, graph_name=None, annotate_stats=False)[source]

Grouped plot + statistical-test helper for spacr experiment DataFrames.

Wraps preprocessing (aggregation by object / well / plate), normality and variance testing, group-wise pairwise stats, and plot rendering (bar / jitter / box / violin / jitter_box / jitter_bar / line / line_std) in a single object whose output can optionally be persisted alongside a CSV of stats.

Parameters:
  • df – Input DataFrame.

  • grouping_column – Categorical grouping variable.

  • data_column – Metric column (or list of columns) to summarise.

  • graph_type – Plot type. Default 'jitter_box' – a box with the points over it. See create_grouped_plot() for why that default is a correction rather than a taste.

  • summary_func – Aggregator for well/plate level. Default 'mean'.

  • order – Explicit ordering of groups.

  • colors – Optional colour palette.

  • output_dir – Save location when save=True.

  • save – If True, persist plot and stats.

  • y_lim – Two-element y-axis limits.

  • log_y – Use log scale for y-axis.

  • log_x – Use log scale for x-axis.

  • error_bar_type – 'std' or 'sem'. Default 'std'.

  • remove_outliers – Drop 1.5*IQR outliers per group before plotting.

  • theme – Seaborn palette name. Default 'pastel'.

  • representation – Aggregation level — 'object', 'well' or 'plate'. Default 'object'.

  • paired – Treat groups as paired samples where applicable.

  • all_to_all – Run every pairwise comparison; False compares each group to compare_group.

  • compare_group – Reference group when all_to_all=False.

  • graph_name – Prefix for saved file names.

  • annotate_stats – Draw a bracket over each pairwise comparison with its asterisks (or ns) above it. Default False: the tests are run and written to the results table on every plot, but with all_to_all=True an N-group plot has N(N-1)/2 comparisons and a stack of that many brackets buries the data it is about. Ask for them when the comparisons are few enough to read. Only drawn for a single data_column; see _draw_comparison_lines().

Store configuration, set the theme, and preprocess the DataFrame.

create_plot(ax=None)[source]

Build the plot for the chosen graph type onto self.fig.

Nothing is displayed: retrieve the figure with get_figure() (and the statistics with get_results()), or call plt.show().

Parameters:

ax – Existing Axes to draw into, for placing this graph in a panel of a larger figure. self.fig is then set to that axes’ parent figure, so a later save=True writes the whole enclosing figure, not this panel alone. None creates a fresh figure sized from the group count and bar_width — and note that with a single data_column the standardisation pass still calls ax.figure.set_size_inches, which resizes a shared figure underneath its other panels.

get_figure()[source]

Return the generated figure.

get_results()[source]

Return the results dataframe.

perform_levene_test(unique_groups)[source]

Levene’s test for equal variance on data_column[0], MEDIAN-centred.

Delegates to spacr.figures.stats.check_equal_variance(). Two things moved when it did, and both change the number a caller writes into a CSV:

  • The centring is the median (Brown-Forsythe), not SciPy’s default mean. Median centring is less sensitive to non-normal data, and this function is called before the normality verdict is known.

  • Below spacr.figures.stats.MIN_N_FOR_ASSUMPTIONS observations in the smallest group the result is (nan, nan). On three replicates Levene has almost no power, so “p = 0.7, variances are equal” means “we could not tell”, and printing 0.7 into a results table invites exactly the reading that publishes a difference that is not there.

Parameters:

unique_groups – Groups to compare.

Returns:

(statistic, p_value), both NaN when the check had no power.

perform_normality_tests()[source]

Evaluate normality for each requested column and group.

Shapiro-Wilk results and the overall verdict come from spacr.figures.stats.check_normality(), including its Bonferroni correction across groups. Groups with fewer than three observations are reported as "Skipped". Checks below the configured information threshold report Informative=False rather than treating a failure to reject as evidence of normality.

Returns:

  • is_normal (bool) – True only when every requested column passes.

  • results (list of dict) – Per-group test statistics, sample sizes, and verdicts.

perform_posthoc_tests(is_normal, unique_groups)[source]

Perform post-hoc tests for multiple groups based on all_to_all flag.

Parameters:
  • is_normal – Outcome of the normality check, which selects the family of test: True runs Tukey HSD, False runs Dunn’s test with an automatically chosen p-adjustment. It only matters when post-hoc testing runs at all — see unique_groups. It must be the verdict perform_normality_tests() returned, which is spacr.figures.stats.check_normality()’s. A hand-computed one puts the omnibus test and the pairwise tests on different footing — Kruskal-Wallis across the groups followed by Tukey between them is two different assumptions about one dataset — and it is how the power floor gets bypassed: three replicates buy Dunn’s, not Tukey.

  • unique_groups – The distinct group labels. Only its length is read; the comparisons themselves are rebuilt from self.df[self.grouping_column], so reordering or renaming entries has no effect. Fewer than three groups returns an empty list, as does self.all_to_all being False, because pairwise correction is meaningless for a single comparison.

Returns:

A list of per-comparison dicts with Comparison, Test Statistic (always None — neither test reports one), p-value, Test Name and the n_object / n_well counts; empty when no post-hoc test was warranted. Only self.data_column[0] is tested, so extra data columns are ignored here.

perform_statistical_tests(unique_groups, is_normal)[source]

Run one supported group comparison per data column.

Parameters:
  • unique_groups (sequence) – Groups to compare. Two groups produce a pairwise test; larger sets produce an omnibus test.

  • is_normal (bool) – External normality verdict. False forces a rank test; True still requires informative engine-level assumption checks.

Returns:

list of dict – Test name, statistic, p-value, sample counts, effect size, and selection rationale for each data column. Untestable comparisons use Test Name='not testable' and include the reason.

Notes

Test selection is delegated to spacr.figures.stats.compare(). Two-group comparisons may use Student’s t, Welch’s t, or Mann-Whitney U; larger comparisons may use one-way ANOVA, Welch’s ANOVA, or Kruskal-Wallis. Paired data use the paired t-test or Wilcoxon signed-rank test.

preprocess_data()[source]

Return a new DataFrame aggregated to the configured representation.

Drops rows with NaN in the grouping or data columns, aggregates the data columns with summary_func per well ('prc') or per plate ('plateID', split out of prc when needed) — or leaves them per object — and makes the grouping column an ordered Categorical.

Returns:

The preprocessed DataFrame; __init__ assigns it back to self.df rather than the frame being modified in place.

Raises:
  • KeyError – if representation='plate' and neither a plateID nor a prc column is available.

  • ValueError – if representation is not 'object', 'well' or 'plate'.

remove_outliers_from_plot()[source]

Remove outliers from the plot but keep them in the data.

spacr.plot.create_grouped_plot(df, grouping_column, data_column, graph_type='jitter_box', summary_func='mean', order=None, colors=None, output_dir='./output', save=False, y_lim=None, error_bar_type='std')[source]

Plot grouped observations and run assumption-aware comparisons.

Pairwise tests are chosen independently by spacr.figures.stats.compare(). Student’s t, Welch’s t, or Mann-Whitney U is used according to the normality and equal-variance checks; an underpowered assumption check selects the rank test. When at least three groups jointly pass normality, Tukey HSD rows are added.

The 'jitter_box' default is a STATISTICAL CORRECTION, not a presentation preference: it shows the observations and their distribution instead of reducing each group to a mean bar. Two groups can have the same mean while having different spreads. The box summarizes the distribution; the jitter stays because the points are the evidence.

Parameters:
  • df (pandas.DataFrame) – Source observations.

  • grouping_column (str) – Categorical column defining groups.

  • data_column (str) – Numeric column to plot and compare.

  • graph_type ({'bar', 'violin', 'jitter', 'box', 'jitter_box'}, optional) – Plot representation. The default shows every observation together with median, quartiles, and whiskers.

  • summary_func (str or callable, optional) – Aggregation used by the bar representation.

  • order (sequence of str, optional) – Group order. By default, sort observed group values.

  • colors (palette-like, optional) – Colours passed to seaborn. By default, use the house data colour.

  • output_dir (path-like, optional) – Directory for saved output.

  • save (bool, optional) – Save the figure and test_results.csv when true.

  • y_lim (sequence of float, optional) – Two-element vertical-axis limits.

  • error_bar_type ({'std', 'sem'}, optional) – Error statistic for bar plots.

Returns:

  • figure (matplotlib.figure.Figure) – Displayed figure. It carries the source recipe used by the interactive representation menu.

  • results_df (pandas.DataFrame) – Normality, pairwise, and optional Tukey HSD results.

Raises:

ValueError – If a bar plot receives an unsupported error_bar_type.

spacr.plot.create_venn_diagram(file1, file2, gene_column='gene', filter_coeff=0.1, save=True, save_path=None)[source]

Compute a two-set gene overlap from CSVs and draw its Venn diagram.

Parameters:
  • file1 – First CSV file.

  • file2 – Second CSV file.

  • gene_column – Column identifying genes. Default 'gene'.

  • filter_coeff – Threshold on the coefficient column — positive filters > threshold, negative filters < threshold.

  • save – If True, save as PDF; requires save_path.

  • save_path – Output PDF path when save is True.

Returns:

{'overlap', 'unique_to_file1', 'unique_to_file2'} lists.

Raises:

ValueError – if save is True but save_path is missing.

spacr.plot.data_colours(fig)[source]

Every colour in fig that carries the CLAIM rather than the frame.

Parameters:

fig – Matplotlib figure whose data artists are inspected.

Used only to say when one of them stops working on paper (150 D), never to change one. A data line is identified the same way _chrome identifies a reference line, from the opposite side of the same test.

spacr.plot.deliverable_dpi(fig, dpi, path=None)[source]

The DPI this figure can actually be written at, and a word if it is not the one that was asked for.

A resolution preference is a request, not a guarantee. spacrGraph pins its canvas to at least 10 inches square and grows it with the number of groups, so 600 and 1200 DPI are not available for a large grouped figure – the raster would be tens of thousands of pixels on a side.

The old behaviour was to hand the number to matplotlib and find out. This returns the DPI that will be used and says, by name, when that is not the DPI that was requested. Appearing to accept a setting and then quietly delivering another one is the failure this avoids.

Parameters:
  • fig – the figure about to be written.

  • dpi – the requested dots per inch.

  • path – destination, named in the message when there is one.

Returns:

the DPI to pass to savefig.

spacr.plot.figure_output_preferences()[source]

Return (format, dpi) from the user’s preferences.

A format or DPI changed in the Preferences figure settings wins over the “Figure format” and “Resolution” preferences. Degrades to DEFAULT_FIGURE_FORMAT / DEFAULT_FIGURE_DPI rather than raising: the preference store is Qt’s, and the pipelines that call this run headless from the CLI and from notebooks, where importing PySide6 to decide a file extension would be absurd.

spacr.plot.figure_path(path, fmt=None)[source]

path with the extension the figure-format preference will write.

THE ONE PLACE A NON-MATPLOTLIB RENDERER CAN ASK. save_figure rewrites the extension itself, which is why every matplotlib save has honoured the preference for months – but pyqtgraph decides what it writes FROM THE NAME (FastPlot.export branches on .pdf / .svg / else), so a renderer that is handed volcano.pdf writes a PDF whatever the user chose. The name has to be settled before the export sees it, and settling it twice in two modules is how the two drift apart.

Parameters:
  • path – a destination, with or without an extension.

  • fmt – force a format, bypassing the preference. Unknown formats fall back to the preference rather than raising: a run must not lose a figure over a typo.

Returns:

the path as a str.

spacr.plot.generate_mask_random_cmap(mask)[source]

Return a random ListedColormap sized to the labels in mask.

Parameters:

mask – Label mask array (0 = background).

Returns:

Random colormap where index 0 is black and remaining entries are random opaque RGBA colours.

spacr.plot.generate_plate_heatmap(df, plate_number, variable, grouping, min_max, min_count)[source]

Aggregate a well-level DataFrame into a plate-shaped heatmap.

The grid is read off the data. It used to be pinned to r1..r16 by c1..c27, so every well of a 1536 plate past row P or past column 27 fell outside the Categorical, became NaN, and was dropped by the groupby — measured, in the database, and absent from the figure with nothing said. Rows and columns now go through spacr.plate_qc.parse_row_label() / parse_column_label (which is spacr.schema’s letter walk, so AA…AF and beyond are real rows), and the axes span exactly the wells present: a 96 plate is still 8x12 and a 384 still 16x24, because nothing is padded out to the largest format that exists.

A well that genuinely cannot be placed — a prc with too few parts, or a row/column token holding no position — is reported through spacr.errors.raise_if_strict() (an ERROR on spacr.errors, or a raise under SPACR_STRICT_ERRORS) naming the identifiers concerned. Replacing a silent drop with a quieter silent drop would fix nothing.

Parameters:
  • df – Long-format DataFrame with a prc (plate_row_column) identifier and the requested variable column.

  • plate_number – Plate ID selecting the subset to display.

  • variable – Column to aggregate. Ignored when grouping='count'.

  • grouping – Aggregation — 'count', 'mean' or 'sum'.

  • min_max – Colour scale spec — 'all', 'allq', or a two-element list [vmin, vmax] (floats treated as quantiles).

  • min_count – Drop wells with fewer than this many rows.

Returns:

(plate_map, (vmin, vmax)) — the pivoted matrix, indexed 'r<N>' by 'c<N>', and the colour-limit tuple.

Raises:
  • ValueError – if grouping is not one of the accepted values.

  • KeyError – if variable is missing and required.

spacr.plot.graph_importance(settings)[source]

Concatenate feature-importance CSVs and hand off to spacrGraph for plotting.

Parameters:

settings – Settings dict with csvs (single path or list), grouping_column, data_column, graph_type, save.

Returns:

None (side-effects: plot shown, artefacts saved).

spacr.plot.illegible_data_colours(fig, ground, floor=None)[source]

The data colours a reader will not find on ground, as hex.

Parameters:
  • fig – Matplotlib figure whose data colours are checked.

  • ground – background colour against which contrast is measured.

The data deliberately does NOT flip, so a palette chosen against near-black can be illegible on paper – and the honest answer is to NAME the colour, because a substitution the user did not ask for changes what the picture says. Deduplicated and sorted so the same sentence comes out of the same figure twice.

spacr.plot.jitterplot_by_annotation(src, x_column, y_column, plot_title='Jitter Plot', output_path=None, filter_column=None, filter_values=None)[source]

Read measurements + annotation from a spacr DB and plot a class-balanced jitter plot.

Parameters:
  • src – Path to a spacr experiment directory containing measurements/measurements.db.

  • x_column – Column used as grouping variable (x-axis).

  • y_column – Numeric column plotted on the y-axis.

  • plot_title – Title for the plot. Default 'Jitter Plot'.

  • output_path – If set, save the figure to this path; otherwise show it.

  • filter_column – Optional column (or list of columns) to filter rows on before plotting.

  • filter_values – Values (or list of value lists) accepted per filter_column.

Returns:

Balanced DataFrame used for the plot.

Raises:

KeyError – if required plate/row/col columns are missing.

spacr.plot.normalize_and_visualize(image, normalized_image, title='')[source]

Show the original and the normalised image side by side in grayscale.

Multi-channel inputs are averaged over their channels for display.

Parameters:
  • image – Original image, 2D or (H, W, C).

  • normalized_image – Normalised counterpart to compare against.

  • title – Suffix appended to both panel titles. Default "".

Returns:

None

spacr.plot.outline_palette_colours(palette)[source]

The four outline colours for palette.

Parameters:

palette – a key of OUTLINE_PALETTES. Anything unknown – including None – falls back to default rather than raising: a figure drawn in the historic colours is a far smaller problem than a pipeline that stops at the plotting step.

Returns:

{object_name: colour}.

spacr.plot.overlay_masks_on_images(img_folder, normalize=True, resize=True, save=False, plot=False, thickness=2)[source]

Overlay masks/* outlines onto matching images from img_folder.

Parameters:
  • img_folder – Folder containing images; masks live in img_folder/masks with matching filenames.

  • normalize – If True, percentile-normalise images before blending. Default True.

  • resize – If True, resize the blended overlay to 1000x1000. Default True.

  • save – If True, write PNGs to img_folder/overlay/. Default False.

  • plot – If True, show each overlay via matplotlib. Default False.

  • thickness – Contour line thickness in pixels. Default 2.

Returns:

{'written': int, 'failed': [(filename, reason)]}. A field that cannot be read is named and skipped rather than ending the run, so a folder holding one truncated TIFF still produces every other overlay – and the caller can tell which ones are missing.

spacr.plot.plot_arrays(src, figuresize=10, cmap='inferno', nr=1, normalize=True, q1=1, q2=99)[source]

Plot random .npy / .npz arrays from src, one channel per subplot.

Parameters:
  • src – Directory or single .npy/.npz path.

  • figuresize – Base figure size. Default 10.

  • cmap – Matplotlib colormap. Default 'inferno'.

  • nr – Maximum number of arrays to plot. Default 1.

  • normalize – If True, percentile-normalise before display. Default True.

  • q1 – Lower percentile for normalisation. Default 1.

  • q2 – Upper percentile for normalisation. Default 99.

Returns:

None

spacr.plot.plot_cellpose4_output(batch, masks, flows, cmap='inferno', figuresize=10, nr=1, print_object_number=True)[source]

Display per-channel images, label mask and flow field for Cellpose v4 outputs.

Parameters:
  • batch – Image batch of shape (N, H, W, C).

  • masks – Label masks, one per image.

  • flows – Flow arrays, one per image.

  • cmap – Colormap for image channels. Default 'inferno'.

  • figuresize – Base figure size. Default 10.

  • nr – Maximum number of images to plot. Default 1.

  • print_object_number – If True, annotate each object with its label ID. Default True.

Returns:

None

spacr.plot.plot_comparison_results(comparison_results)[source]

Plot Jaccard, Dice, boundary-F1 and average-precision distributions per comparison.

Parameters:

comparison_results – Iterable of dicts with per-file metrics (each key ending in jaccard/dice/boundary_f1/ average_precision).

Returns:

The generated Figure.

spacr.plot.plot_data_from_csv(settings)[source]

Load per-plate CSVs, filter/outlier-clean and render a spacrGraph plot.

Parameters:

settings – Settings dict — see settings.get_plot_data_from_csv_default_settings for keys (src, data_column, grouping_column, keep_groups, remove_outliers, graph_type, graph_name, …).

Returns:

(fig, results_df, df) — the figure, stats DataFrame and plotted DataFrame.

Raises:

ValueError – if src is not a string or list.

spacr.plot.plot_data_from_db(settings)[source]

Read one or more measurement DBs, annotate conditions and render a spacrGraph plot.

Concatenates results across source directories, derives the recruitment column if requested, drops missing rows, then hands the data to spacrGraph for statistics + plotting.

Parameters:

settings – Settings dict. See settings.set_default_plot_data_from_db for accepted keys (notably src, database, table_names, data_column, grouping_column, graph_type, graph_name).

Returns:

The plotted DataFrame, or None when the requested data or grouping column is missing.

Raises:

ValueError – if src is neither a string nor a list.

spacr.plot.plot_feature_importance(feature_importance_df, title='')[source]

Plot a horizontal bar chart of raw feature importances.

Parameters:
  • feature_importance_df – DataFrame with columns feature and importance.

  • title – what the bars MEAN, when it is not the model’s own importances. Four of the classifiers spaCR offers expose no feature_importances_ and are drawn from permutation importance instead – a different quantity, measuring what the fitted model loses when a column is shuffled – and a panel that did not say so would be passing one off as the other.

Returns:

The generated Figure.

spacr.plot.plot_histogram(df, column, dst=None)[source]

Plot a histogram of df[column] and optionally save it as PDF.

Parameters:
  • df – DataFrame containing column.

  • column – Column to plot.

  • dst – If set, save under <dst>/<column>_histogram.pdf.

Returns:

None

spacr.plot.plot_image_grid(image_paths, percentiles)[source]

Render a square grid of percentile-normalised images with a black background.

Each tile carries its source file and the per-channel display range it was stretched to, so a checked export (see save_figure()) can say when tiles are scaled differently and write a provenance sidecar that rebuilds every tile from its file.

Parameters:
  • image_paths – Image files to display; extra tiles are filled black.

  • percentiles – Two-element percentile pair used to normalise each channel.

Returns:

The generated Figure.

spacr.plot.plot_image_mask_overlay(file, channels, cell_channel, nucleus_channel, pathogen_channel, organelle_channel=None, figuresize=10, percentiles=(2, 98), thickness=3, save_pdf=True, mode='outlines', export_tiffs=False, all_on_all=False, all_outlines=False, filter_dict=None, outline_palette='default', organelle_channels=None)[source]

Plot image and mask overlays.

Loads the merged .npy stack, draws one panel per requested channel with the object masks applied as contours or filled labels, and closes with a panel showing every object combined.

Parameters:
  • file – Path to the merged .npy stack for one field of view.

  • channels – Indices of the image channels to draw, one panel each.

  • cell_channel – Intensity channel the cell mask belongs to, or None when there is no cell mask.

  • nucleus_channel – Intensity channel the nucleus mask belongs to, or None.

  • pathogen_channel – Intensity channel the pathogen mask belongs to, or None.

  • organelle_channel – Intensity channel the organelle mask belongs to, or None. Default None.

  • figuresize – Figure height in inches; the figure is drawn four times as wide. Default 10.

  • percentiles – Two-element percentile pair used to normalise each channel. Default (2, 98).

  • thickness – Contour line width in pixels. Default 3.

  • save_pdf – If True, save the figure into results/overlay/ two directories above file, in the configured figure format rather than always as PDF. Default True.

  • mode – 'outlines' draws mask contours; any other value overlays filled, randomly coloured labels. Default 'outlines'.

  • export_tiffs – If True, also write every stack plane as a grayscale TIFF into results/<stem>/tiff/ alongside it. Default False.

  • all_on_all – If True, draw every mask on every channel. Default False.

  • all_outlines – If True, draw every mask on the channels that own no mask themselves. Default False.

  • filter_dict – Optional per-object limits keyed by 'cell', 'nucleus', 'pathogen' or 'organelle', each holding ((min_area, max_area), (min_intensity, max_intensity)); objects outside the limits are dropped before plotting.

  • outline_palette – which outline colours to draw, a key of OUTLINE_PALETTES. 'default' is what spaCR has always drawn; 'colourblind' is legible under red-green and blue-yellow deficiency, where the default’s worst pair scores 27 out of 255 – cell is drawn red and pathogen green, the one pair the commonest deficiency removes. Default 'default', because changing every figure a user has already made would be worse than the defect.

  • organelle_channels – Optional {slot role: channel} for the organelle slots after the first ({'organelleb': 3}), each drawn in its own colour. Default None draws the first slot only, as before. Mask planes are located through the merged folder’s plane layout sidecar when one is present, so a slot left out here no longer shifts the planes of the objects that are drawn.

Returns:

The generated matplotlib Figure.

spacr.plot.plot_image_mask_overlay_magenta_outlines(file, channels, cell_channel, nucleus_channel, pathogen_channel, figuresize=10, percentiles=(2, 98), thickness=3, save_pdf=True, mode='outlines', export_tiffs=False, all_on_all=False, all_outlines=False, filter_dict=None)[source]

Plot image and mask overlays, outlining each channel’s own mask in magenta.

Variant of plot_image_mask_overlay() with no organelle_channel: when mode is 'outlines' and all_on_all is False, the mask belonging to a channel is outlined in magenta rather than in that object’s colour. In every other mode it falls back to filled, randomly coloured labels as that function does, but seeded per call rather than per object, so the colours differ between runs and between panels.

Parameters:
  • file – Path to the merged .npy stack for one field of view.

  • channels – Indices of the image channels to draw, one panel each.

  • cell_channel – Intensity channel the cell mask belongs to, or None when there is no cell mask.

  • nucleus_channel – Intensity channel the nucleus mask belongs to, or None.

  • pathogen_channel – Intensity channel the pathogen mask belongs to, or None.

  • figuresize – Figure height in inches; the figure is drawn four times as wide. Default 10.

  • percentiles – Two-element percentile pair used to normalise each channel. Default (2, 98).

  • thickness – Contour line width in pixels. Default 3.

  • save_pdf – If True, save the figure into results/overlay/ two directories above file, in the configured figure format rather than always as PDF. Default True.

  • mode – 'outlines' draws mask contours; any other value overlays filled, randomly coloured labels. Default 'outlines'.

  • export_tiffs – If True, also write every stack plane as a grayscale TIFF into results/<stem>/tiff/ alongside it. Default False.

  • all_on_all – If True, draw every mask on every channel in its own colour. Default False.

  • all_outlines – If True, draw every mask on the channels that own no mask themselves. Default False.

  • filter_dict – Optional per-object limits with a 'cell', 'nucleus' and 'pathogen' entry, each holding ((min_area, max_area), (min_intensity, max_intensity)); objects outside the limits are dropped before plotting.

Returns:

The generated matplotlib Figure.

spacr.plot.plot_images_and_arrays(folders, lower_percentile=1, upper_percentile=99, threshold=1000, extensions=None, overlay=False, max_nr=None, randomize=True)[source]

Show side-by-side images and arrays found across multiple folders.

Each image is either percentile-normalised (values below threshold) or shown as a label mask. Optionally overlays object outlines from a matching mask file.

Parameters:
  • folders – Folders to scan for image/array files.

  • lower_percentile – Lower percentile clip. Default 1.

  • upper_percentile – Upper percentile clip. Default 99.

  • threshold – Values <= threshold are treated as label data instead of intensity. Default 1000.

  • extensions – File extensions to include. Default ['.npy', '.tif', '.tiff', '.png'].

  • overlay – If True, overlay object outlines. Default False.

  • max_nr – Maximum number of key groups to plot.

  • randomize – If True, shuffle key order before plotting. Default True.

Returns:

None

spacr.plot.plot_lorenz_curves(csv_files, name_column='grna_name', value_column='count', remove_keys=None, x_lim=None, y_lim=None, remove_outliers=False, save=True)[source]

Overlay Lorenz curves from multiple gRNA count CSVs with per-plate Gini coefficients.

Parameters:
  • csv_files – Paths to per-plate CSVs, each with columns name_column and value_column.

  • name_column – Identifier column used for outlier filtering. Default 'grna_name'.

  • value_column – Column whose distribution is analysed. Default 'count'.

  • remove_keys – Names to exclude before analysis. Default [] (exclude nothing).

  • x_lim – X-axis limits [lo, hi]. Default [0.0, 1].

  • y_lim – Y-axis limits [lo, hi]. Default [0, 1].

  • remove_outliers – If True, drop names whose number of observations falls outside a fence extending 1.5 times the 5th-to-95th-percentile spread of group sizes. Count values do not enter this filter. Default False.

  • save – If True, save the figure alongside the first CSV under results/lorenz_curve_with_gini.pdf. Default True.

Returns:

None

spacr.plot.plot_masks(batch, masks, flows, cmap='inferno', figuresize=10, nr=1, file_type='.npz', print_object_number=True)[source]

Display per-channel images, label masks and flow fields for a batch.

Parameters:
  • batch – Image batch — shape (N, H, W, C) or a single image of shape (H, W, C).

  • masks – Label masks, one per image (list or ndarray).

  • flows – Flow arrays, one per image.

  • cmap – Colormap for image channels. Default 'inferno'.

  • figuresize – Base figure size. Default 10.

  • nr – Maximum number of images to plot. Default 1.

  • file_type – Source file type of flows — 'png' selects the first element of each flow entry. Default '.npz'.

  • print_object_number – If True, annotate each object with its label ID. Default True.

Returns:

None

spacr.plot.plot_merged(src, settings)[source]

Show multi-channel image stacks with per-object outlines overlaid.

Parameters:
  • src – Folder containing .npy merged stacks.

  • settings – Plot settings dict — includes channel/mask dims, overlay colours, normalisation, filter and object-count keys.

Returns:

The last generated Figure when settings['nr'] is exceeded; otherwise None.

spacr.plot.plot_object_outlines(src, objects=None, channels=None, max_nr=10)[source]

Overlay mask outlines on the matching channel image for each object type.

Parameters:
  • src – Experiment root; masks/<object>_mask_stack and channel folders live directly under it.

  • objects – Object types to plot. Default ['nucleus', 'cell', 'pathogen'].

  • channels – Channel indices paired with objects (channel folders are named <channel + 1>). Default [0, 1, 2].

  • max_nr – Maximum number of images to plot per object. Default 10.

Returns:

None

spacr.plot.plot_organelle_output(img_batch, masks, settings, cmap='inferno', figuresize=10, nr=1, print_object_number=True)[source]

Plot organelle segmentation results: raw channel, label mask, morphology-specific diagnostic.

Parameters:
  • img_batch – Single-channel image batch of shape (N, H, W).

  • masks – Label masks, one per image.

  • settings – Organelle settings dict; organelle_morphology and organelle_method drive the diagnostic panel.

  • cmap – Colormap for the raw channel. Default 'inferno'.

  • figuresize – Base figure size. Default 10.

  • nr – Maximum number of images to plot. Default 1.

  • print_object_number – If True, annotate each object with its label ID. Default True.

Returns:

None

spacr.plot.plot_permutation(permutation_df)[source]

Plot a horizontal bar chart of permutation feature importances with error bars.

Parameters:

permutation_df – DataFrame with columns feature, importance_mean and importance_std.

Returns:

The generated Figure.

spacr.plot.plot_plates(df, variable, grouping, min_max, cmap, min_count=0, verbose=True, dst=None)[source]

Render every plate of a screen as ONE panel, wells square, on one colour scale.

The layout, the colour scale and the treatment of unmeasured wells live in spacr.figures.plates; this function is the call the pipeline already makes, kept at its own signature.

WHAT CHANGED, AND WHY (the design – “the lpates look super small on the collected figure”): the plates were laid out four-per-row on a 40 x 5 inch figure, an 8:1 strip that uses about an eighth of a square tile in the figure grid, with wells 1.14:1 rather than square. They are now a SMALL MULTIPLE – 2 x 2 for a four-plate screen – which is a 1.3:1 composite, and the figure is sized from the well grid so the wells come out exactly square.

Two things that were wrong with the picture and not only with its size: each plate carried its OWN colour scale, so the same blue meant a different number on the plate beside it; and a well that was never measured was drawn as a measurement of zero, which on a screen with 155 of 384 wells used is more than half the panel — and set the bottom of the scale. One scale is now shared across the plates, and an unmeasured well is drawn as a neutral wash and left out of the scale.

Parameters:
  • df – Long-format DataFrame with a prc column of the form plateID_rowID_columnID and the column named by variable.

  • variable – Column to aggregate (see generate_plate_heatmap()).

  • grouping – Aggregation mode — 'count', 'mean' or 'sum'.

  • min_max – Color-scale spec ('all', 'allq', [vmin, vmax]), applied ONCE over every plate rather than once per plate.

  • cmap – Matplotlib colormap name or object. None — or the legacy 'viridis' literal — uses the house single-hue ramp.

  • min_count – Drop wells with fewer than this many rows before plotting. Default 0.

  • verbose – If True, call plt.show() after building the figure. Default True.

  • dst – If given, save the figure as <dst>/plate_heatmap_<variable>.pdf.

Returns:

The generated matplotlib Figure.

Example

from spacr.plot import plot_plates
fig = plot_plates(
    df, variable='recruitment', grouping='mean',
    min_max='allq', cmap=None, min_count=20,
)

See also

spacr.figures.plates.build_plates() — the panel itself, which also returns the legend sentence for it. spacr.ml.generate_ml_scores() — produces score dataframes typically fed to this plotter.

spacr.plot.plot_proportion_stacked_bars(settings, df, group_column, bin_column, prc_column='prc', level='object', cmap='viridis')[source]

Plot stacked proportion bars per group with chi-squared and pairwise stats.

Parameters:
  • settings – Settings dict — verbose toggles pairwise chi-squared verbosity.

  • df – Long-format DataFrame with categorical group_column and bin_column.

  • group_column – Group axis of the stacked bars.

  • bin_column – Categorical column stacked within each bar.

  • prc_column – Per-well identifier used when aggregating at the well or plate level. Default 'prc'.

  • level – Aggregation level — 'object' for direct counts, or 'well' / 'plateID' for per-well means with SD bars.

  • cmap – Matplotlib colormap. Default 'viridis'.

Returns:

(results_df, pairwise_results, fig) — chi-squared summary, pairwise comparison table and the plot figure.

spacr.plot.plot_region(settings)[source]

Render mask overlay, cropped PNG grid and activation-map grid for one FOV.

Reads the FOV’s merged NPY, resolves its PNG crops and activation maps from the measurements and activation DBs, and writes the three figures under <src>/results/<name>/ when possible — in the configured figure format, so PDF only while that is the preference.

Parameters:

settings – Settings dict with src, name, channels, cell_channel, nucleus_channel, pathogen_channel, percentiles, activation_mode, activation_db, mode, export_tiffs.

Returns:

Tuple (fig_mask_overlay, fig_png_grid, fig_activation_grid) — any element may be None when the corresponding assets were not found.

spacr.plot.plot_resize(images, resized_images, labels, resized_labels)[source]

Show original vs. resized image/label pairs in a 2x2 grid.

Parameters:
  • images – Sequence of original images (first element shown).

  • resized_images – Sequence of resized images.

  • labels – Sequence of original label arrays.

  • resized_labels – Sequence of resized label arrays.

Returns:

None

spacr.plot.print_mask_and_flows(stack, mask, flows, overlay=True, max_size=1000, thickness=2)[source]

Show a single image, its label mask (optionally outlined) and flow image.

Parameters:
  • stack – Original 2D image or (H, W, C) stack.

  • mask – Label mask matching stack spatially.

  • flows – Optional list of flow arrays; skipped when None.

  • overlay – If True, draw mask contours over the image instead of showing the mask alone. Default True.

  • max_size – Downsample any dimension exceeding this size. Default 1000.

  • thickness – Contour line thickness in pixels. Default 2.

Returns:

None

spacr.plot.print_ready(fig, mode=None, announce=True)[source]

Repaint fig’s chrome for paper for the length of the block.

THE CONTRACT IS THAT NOTHING SURVIVES IT. Every artist touched is restored in a finally, so a user watching a plot while it saves must not see it flash and the figure is byte-identical afterwards – the own acceptance, and the reason this is a context manager rather than a function that fixes a figure up.

WHAT MOVES: an illegible ground becomes the page; illegible chrome becomes the ink; an illegible GRID becomes a faint print grey rather than the ink, because a grid repainted in the ink is a cage over the data.

WHAT DOES NOT MOVE: every data colour, and any chrome that was already legible on the page. A light-mode save therefore changes nothing at all, which is the property that makes this safe to switch on by default.

Parameters:
  • fig – matplotlib figure whose page, chrome and grid artists are repainted temporarily and restored when the context exits.

  • mode – one of spacr.figure_style.SAVE_MODES; None asks the preference. 'screen' is a no-op by construction.

  • announce – print the 150 D sentence when a data colour has stopped working on the page. Off for a caller that saves in a loop.

spacr.plot.proportion_mixed_model(df, group_column, bin_column, unit_column)[source]

A binomial GLM on the per-object outcome, standard errors clustered by unit.

Parameters:
  • df – object-level observations carrying group, bin, and unit fields.

  • group_column – column naming the conditions to compare.

  • bin_column – categorical outcome column modelled one bin at a time.

  • unit_column – column naming clusters used for robust standard errors.

The proportions test throws away how many objects each well contributed; this keeps them while still charging the degrees of freedom the DESIGN supports, by clustering on the unit. Reported beside the other two because when it disagrees with them, the disagreement is the finding.

spacr.plot.proportion_test_by_unit(df, group_column, bin_column, unit_column)[source]

Compare conditions on their PER-UNIT proportions, one row per bin.

Parameters:
  • df – object-level observations carrying group, bin, and unit fields.

  • group_column – column naming the conditions to compare.

  • bin_column – categorical outcome column whose bins are tested.

  • unit_column – column naming independent replication units.

The object-level chi-squared asks whether 20,000 objects came from one distribution. Objects in a well share a treatment, a transfection, an imaging session and a monolayer, so that is not the question anyone asked, and its p-value is smaller than the experiment supports by orders of magnitude. This asks the question the design supports: do the WELLS differ, with n = the number of wells.

spacr.plot.proportions_per_unit(df, group_column, bin_column, unit_column)[source]

Each unit’s share of every bin, one row per unit.

Parameters:
  • df – object-level observations carrying group, bin, and unit fields.

  • group_column – column naming the conditions to compare.

  • bin_column – categorical outcome column whose shares are computed.

  • unit_column – column naming independent wells, plates, or other replication units.

Returns:

a frame with group_column, unit_column and one column per bin holding a proportion in [0, 1]. Units contributing no objects do not appear.

spacr.plot.random_cmap(num_objects=100)[source]

Return a random ListedColormap with num_objects + 1 colours.

Parameters:

num_objects – Number of foreground colours to generate. Default 100.

Returns:

Colormap with index 0 = black and remaining indices random opaque RGBA colours.

spacr.plot.read_and_plot__vision_results(base_dir, y_axis='accuracy', name_split='_time', y_lim=None)[source]

Aggregate vision-model test CSVs under base_dir and plot mean score per model.

Parameters:
  • base_dir – Root directory containing *_test_result.csv files nested per epoch.

  • y_axis – Metric column to average. Default 'accuracy'.

  • name_split – Substring that splits filename into model name and epoch info. Default '_time'.

  • y_lim – Y-axis limits [lo, hi]. Default [0.8, 0.9].

Returns:

None

spacr.plot.save_figure(fig, path, *, fmt=None, dpi=None, close=False, save_mode=None, announce_colours=True, integrity=None, **kwargs)[source]

Write fig to path, honouring the figure preferences.

The single place a spaCR figure the user keeps gets written. Before this existed there were sixty-odd savefig calls, each with its own hard-coded format and DPI, and the “Figure format” and “Resolution” preferences reached exactly two of them – both writing to a temp directory. Everything a pipeline saved into its results folder ignored both settings entirely.

Three things are decided here rather than left to the caller or to matplotlib.

The format follows the preference, and the file NAME follows the format: a PNG written to figure.pdf is a file no viewer opens. An explicit fmt= still wins, for the few callers that genuinely need one particular format.

Fonts are embedded as TrueType (pdf.fonttype = 42) for the length of the save. matplotlib’s default is Type 3, which draws every glyph as its own content stream: the file is still vector, but Illustrator and Inkscape open the text as unselectable outlines, and the preference that selects this path is labelled “PDF (vector, editable)”. Scoped with rc_context so a caller that has deliberately chosen otherwise is not changed underneath it.

The DPI is passed, always. A PDF page is resolution-independent, but spaCR figures are full of imshow panels – cell montages, mask overlays, plate heatmaps – and those are rasterised at the figure’s own 100 DPI unless told otherwise. Without it, 100, 300 and 600 produced byte-identical files. What is passed is deliverable_dpi(), which says so out loud when the requested number is not achievable here.

The chrome is repainted for paper, and only for the length of the write. A figure saved from a dark-themed session was white ink on a white page – spacr.qt.preferences.get_figure_colors hands both renderers a white foreground on a dark theme and nothing inverted it at export time, so the file was a blank rectangle with some coloured dots in it, and a transparent PNG even looked right in a dark file manager and disappeared when it was pasted into a manuscript. print_ready() moves the furniture and puts it back; the DATA is never touched, because a white point turned black is, on a volcano, the colour of “not a hit”.

The figure first takes the settings changed in the Preferences figure settings (spacr.figures.style._apply_user_style()), and a second copy is written when “Also save” names another format.

Parameters:
  • fig – a matplotlib Figure.

  • path – destination; its extension is corrected to the format.

  • fmt – force a format, bypassing the preference.

  • dpi – force a DPI, bypassing the preference.

  • close – close the figure once written.

  • save_mode – force one of spacr.figure_style.SAVE_MODES – 'print' (light page, dark chrome), 'screen' (exactly what is on screen, the old behaviour) or 'transparent'. None asks the preference, which defaults to 'print'.

  • announce_colours – say when a data colour has stopped working on the light page. It is NAMED, never substituted.

  • integrity – check the figure’s image panels before writing – display ranges that differ between panels meant for comparison, saturated or clipped pixels, repeated panels, a lossy format, and panels written with fewer pixels than they hold – print any warning, stamp the provenance (source files, display settings, processing steps, spaCR version) into the PNG or PDF metadata and write it to a <file>.provenance.json sidecar beside the figure. None follows SPACR_FIGURE_INTEGRITY and then the Preferences toggle, which is off by default. A figure without image panels is written unchanged either way.

  • kwargs – passed through to savefig (bbox_inches etc.). An explicit facecolor still wins over the print ground.

Returns:

the path actually written, as a str.

spacr.plot.visualize_cellpose_masks(masks, titles=None, filename=None, save=False, src=None)[source]

Display several Cellpose-style label masks side by side for a quick visual QC.

Handy for sanity-checking the masks produced by spacr.core.preprocess_generate_masks() (e.g. compare the cell, nucleus and pathogen masks of the same field, or two runs against each other). Each mask is rendered with a random-color palette so neighbouring objects stay distinguishable.

Parameters:
  • masks – Sequence of 2D label mask arrays.

  • titles – Titles paired positionally with masks. Falls back to 'Mask 1', 'Mask 2', …

  • filename – Used in the figure suptitle and, when save, as the output PDF filename.

  • save – If True, save the figure under <src>/results/<filename>.pdf. Default False.

  • src – Root folder for saving. Defaults to the current working directory.

Returns:

None. Displays (and optionally writes) the figure.

Raises:

AssertionError – if titles and masks have different lengths.

Example

from spacr.plot import visualize_cellpose_masks
visualize_cellpose_masks(
    [cell_mask, nucleus_mask, pathogen_mask],
    titles=['cell','nucleus','pathogen'],
    filename='field_001', save=True, src='/data/plate01',
)

See also

spacr.core.preprocess_generate_masks() — produces the masks visualized here.

spacr.plot.visualize_masks(mask1, mask2, mask3, title='Masks Comparison')[source]

Show three masks side by side with random colormaps.

Parameters:
  • mask1 – First label mask.

  • mask2 – Second label mask.

  • mask3 – Third label mask.

  • title – Figure suptitle. Default "Masks Comparison".

Returns:

None

spacr.plot.volcano_plot(data: str | pandas.DataFrame, *, fold_change_col: str, p_value_col: str, name_col: str | None = None, x_transform: str = 'none', y_transform: str = '-log10', fold_change_threshold: float | None = None, p_value_threshold: float | None = None, annotate: bool = True, annotate_max: int | None = None, point_size: float = 20.0, alpha: float = 0.7, figsize: Tuple[float, float] = (8.0, 6.0), title: str | None = None, xlim: Tuple[float, float] | None = None, ylim: Tuple[float, float] | None = None, threshold_line_kwargs: dict | None = None, scatter_kwargs: dict | None = None, text_kwargs: dict | None = None, save_path: str | None = None, show: bool = True, ax: matplotlib.pyplot.Axes | None = None, sheet_name: int | str = 0) → Tuple[matplotlib.pyplot.Figure, matplotlib.pyplot.Axes, list][source]

Read a table (CSV/TSV/XLS/XLSX or a DataFrame) and render a volcano plot.

Auto-detects file type from extension (.csv, .tsv/.tab, .xls/.xlsx) and applies the requested x/y transforms before drawing.

Parameters:
  • data – Path to table file or a pandas DataFrame.

  • fold_change_col – Column of raw fold change (or logFC when x_transform='none').

  • p_value_col – Column of p-values.

  • name_col – Optional column supplying point labels.

  • x_transform – One of 'none', 'log2', 'log10', 'ln'. Use 'none' when the column already stores logFC (may be negative).

  • y_transform – One of 'none', '-log10', '-ln', 'log10', 'ln'. Default '-log10'.

  • fold_change_threshold – Threshold on x — in plotted units when x_transform='none', otherwise in raw FC units.

  • p_value_threshold – Threshold on raw p; drawn as a dashed horizontal line in plotted units.

  • annotate – Annotate significant points when a name column is supplied.

  • annotate_max – Cap on the number of annotated points (highest y first).

  • point_size – Scatter marker size.

  • alpha – Scatter marker alpha.

  • figsize – Figure size in inches.

  • title – Optional figure title.

  • xlim – Optional x-axis limits.

  • ylim – Optional y-axis limits.

  • threshold_line_kwargs – Extra kwargs for threshold lines.

  • scatter_kwargs – Extra kwargs for the scatter call.

  • text_kwargs – Extra kwargs for label texts.

  • save_path – If given, save the figure to this path.

  • show – Call plt.show() at the end. Default True.

  • ax – Existing axes to draw on; a new figure is created if None.

  • sheet_name – Excel sheet index/name for .xls/.xlsx inputs.

Returns:

(fig, ax, hits) where hits are the labels drawn.

Raises:

ValueError – on unknown transforms, or numeric columns that cannot be coerced.

Nested helpers

_panel_thumbnail._zscored(values)

Standardize a signature, or return None when its spread is unusable.

spacr/plot.py:1038

_plot_cropped_arrays.plot_single_array(array, ax, title, chosen_cmap)

Render one channel from stack onto ax. No colorbar is drawn.

Parameters:
  • array (ndarray) – One 2D plane of stack. Its count of distinct values, not its dtype, is what decides whether it is treated as an intensity image or as a label mask – so a uint8 plane, which can hold at most 256 distinct values, is always taken for a mask under the default threshold of 500.

  • ax (matplotlib.axes.Axes) – Axes drawn on in place; its frame and ticks are switched off, the title is fixed at size 18, and nothing is returned.

  • title (str) – Panel title. When the plane is treated as a mask, the object count is appended as ", N (obj.)". That count is the number of distinct non-zero values, so the background value 0 is never counted, but a negative value is counted as an object.

  • chosen_cmap (Colormap) – Colormap for the intensity case only. It is discarded when the plane has no more than threshold unique values (the enclosing function’s argument, default 500), because a random colormap – black at index 0, one random opaque colour per non-zero label – is generated instead so neighbouring objects stay distinguishable.

spacr/plot.py:5129

_save_scimg_plot._visualize_scimgs(src, channel_indices=None, um_per_pixel=0.1, scale_bar_length_um=10, show_filename=True, standardize=True, nr_imgs=None, fontsize=8, channel_names=None, plot=False)

Visualize single-cell images.

Parameters:
  • src (str) – The source directory path.

  • channel_indices (list, optional) – List of channel indices to visualize. Defaults to None.

  • um_per_pixel (float, optional) – Micrometers per pixel. Defaults to 0.1.

  • scale_bar_length_um (float, optional) – Length of the scale bar in micrometers. Defaults to 10.

  • show_filename (bool, optional) – Whether to show the filename on the image. Defaults to True.

  • standardize (bool, optional) – Whether to standardize the image sizes. Defaults to True.

  • nr_imgs (int, optional) – The number of images to visualize. Defaults to None.

  • fontsize (int, optional) – Font size for the filename. Defaults to 8.

  • channel_names (list, optional) – List of channel names. Defaults to None.

  • plot (bool, optional) – Whether to plot the images. Defaults to False.

Returns:

matplotlib.figure.Figure – The figure object containing the plotted images.

spacr/plot.py:5029

_save_scimg_plot._visualize_scimgs._generate_filelist(src)

Generate a list of image files in the specified directory.

Parameters:

src (str) – The source directory path.

Returns:

list – A list of image file paths.

spacr/plot.py:5049

_save_scimg_plot._visualize_scimgs._random_sample(file_list, nr_imgs=None)

Randomly selects a subset of files from the given file list.

Parameters:
  • file_list (list) – A list of file names.

  • nr_imgs (int, optional) – The number of files to select. If None, all files are selected. Defaults to None.

Returns:

list – A list of randomly selected file names.

spacr/plot.py:5064

_visualize_and_save_timelapse_stack_with_tracks._view_frame_with_tracks(frame=0)

Display the frame with tracks overlaid.

Parameters: frame (int): The frame number to display.

Returns: None

spacr/plot.py:5212

jitterplot_by_annotation._resolve_well_column(frame, *bases)

Return the first _x, bare, or _y well-column spelling.

spacr/plot.py:6711

jitterplot_by_annotation.join_measurments_and_annotation(src, tables)

Join per-object measurement tables with the png_list annotation table.

Parameters:
  • src – spaCR experiment directory; the database is read from <src>/measurements/measurements.db and no other layout is supported — pass the experiment folder, not the .db file.

  • tables – Object tables to merge, joined on prcfo. Every name listed must exist in the database.

Returns:

One row per object, with the png_list crop path attached by a left join.

Raises:

pandas.errors.MergeError – if png_list holds more than one crop per prcfo; the join is validated one_to_one precisely so duplicated crops cannot silently multiply the measurement rows and inflate the jitter plot.

spacr/plot.py:6668

overlay_masks_on_images.normalize_image(image)

Normalize the image to the 1st and 99th percentiles.

Parameters:

image – Image array of any numeric dtype, typically the raw 16-bit TIFF. The percentiles are taken over the whole array, so a multi-channel image is stretched by one shared window rather than per channel, and the brightest and darkest 1% saturate. A flat or near-constant image puts both percentiles on the same value; the rescale then divides by zero and the uint8 cast turns the resulting nan into an undefined value, so guard empty fields upstream.

Returns:

A uint8 array on 0-255, ready to blend with the mask overlay.

spacr/plot.py:8797

plot_data_from_csv.filter_rows_by_column_values(df: pd.DataFrame, column: str, values: list) → pd.DataFrame

Return a filtered DataFrame where only rows with the column value in the list are kept.

Parameters:
  • df – Frame to filter; it is not modified, and the result is a .copy() so later assignment to it raises no SettingWithCopyWarning.

  • column – Column to test. Must exist, or KeyError is raised — here it is the caller’s grouping_column.

  • values – Values to keep, matched with isin so comparison is exact and type-sensitive: the string '1' will not match an integer 1 read from the CSV. An empty list keeps nothing and yields an empty frame rather than passing everything through.

Returns:

A new filtered DataFrame.

spacr/plot.py:8530

plot_image_mask_overlay._filter_object(mask, intensity_image, min_max_area=(0, 10000000), min_max_intensity=(0, 65000), type_='object')

Filter objects in a mask based on their area (size) and mean intensity.

Parameters:
  • mask (ndarray) – The input mask.

  • intensity_image (ndarray) – The corresponding intensity image.

  • min_max_area (tuple) – A tuple (min_area, max_area) specifying the minimum and maximum area thresholds.

  • min_max_intensity (tuple) – A tuple (min_intensity, max_intensity) specifying the minimum and maximum intensity thresholds.

Returns:

ndarray – The filtered mask.

spacr/plot.py:3643

plot_image_mask_overlay._plot_merged_plot(image, outlines, outline_colors, figuresize, thickness, percentiles, mode='outlines', all_on_all=False, all_outlines=False, channels=None, channel_to_outline=None, channel_to_label=None, save_pdf=True)

Plot the merged plot with overlay, image channels, and masks.

spacr/plot.py:3439

plot_image_mask_overlay._plot_merged_plot._apply_contours(image, mask, color, thickness)

Apply contours to the image.

spacr/plot.py:3481

plot_image_mask_overlay._plot_merged_plot._generate_colored_mask(mask, cmap)

Generate a colored mask using the given colormap.

spacr/plot.py:3456

plot_image_mask_overlay._plot_merged_plot._generate_contours(mask)

Generate contours from the mask using OpenCV.

spacr/plot.py:3474

plot_image_mask_overlay._plot_merged_plot._normalize_image(image, percentiles)

Normalize the image based on given percentiles.

spacr/plot.py:3468

plot_image_mask_overlay._plot_merged_plot._overlay_mask(image, mask)

Overlay the colored mask onto the original image.

spacr/plot.py:3463

plot_image_mask_overlay._save_channels_as_tiff(stack, save_dir, filename)

Save each channel in the stack as a grayscale TIFF.

spacr/plot.py:3634

plot_image_mask_overlay.random_color_cmap(n_labels, seed=None)

Generate a random-looking but deterministic colormap with a unique seed.

Parameters:
  • n_labels – How many object colours to draw. Index 0 of the returned colormap is forced to black for background, so the map holds n_labels + 1 entries; callers here pass int(outline.max() + 1) per object type, or int(combined_mask.max() + 1) for the merged panel. A value <= 0 short-circuits to a black-only colormap rather than raising.

  • seed – Seed for a local default_rng; the same seed always produces the same hue assignment, which is why each object type is given its own fixed seed and so keeps its colours across panels. None draws fresh entropy and colours change per call.

Returns:

A ListedColormap of vivid, well-separated hues.

spacr/plot.py:3408

plot_image_mask_overlay_magenta_outlines._filter_object(mask, intensity_image, min_max_area=(0, 10000000), min_max_intensity=(0, 65000), type_='object')

Filter objects in a mask based on their area (size) and mean intensity.

Parameters:
  • mask (ndarray) – The input mask.

  • intensity_image (ndarray) – The corresponding intensity image.

  • min_max_area (tuple) – A tuple (min_area, max_area) specifying the minimum and maximum area thresholds.

  • min_max_intensity (tuple) – A tuple (min_intensity, max_intensity) specifying the minimum and maximum intensity thresholds.

Returns:

ndarray – The filtered mask.

spacr/plot.py:4029

plot_image_mask_overlay_magenta_outlines._plot_merged_plot(image, outlines, outline_colors, figuresize, thickness, percentiles, mode='outlines', all_on_all=False, all_outlines=False, channels=None, cell_channel=None, nucleus_channel=None, pathogen_channel=None, cell_outlines=None, nucleus_outlines=None, pathogen_outlines=None, save_pdf=True)

Plot the merged plot with overlay, image channels, and masks.

spacr/plot.py:3874

plot_image_mask_overlay_magenta_outlines._plot_merged_plot._apply_contours(image, mask, color, thickness)

Apply contours to the image.

spacr/plot.py:3920

plot_image_mask_overlay_magenta_outlines._plot_merged_plot._generate_colored_mask(mask, cmap)

Generate a colored mask using the given colormap.

spacr/plot.py:3895

plot_image_mask_overlay_magenta_outlines._plot_merged_plot._generate_contours(mask)

Generate contours from the mask using OpenCV.

spacr/plot.py:3913

plot_image_mask_overlay_magenta_outlines._plot_merged_plot._normalize_image(image, percentiles)

Normalize the image based on given percentiles.

spacr/plot.py:3907

plot_image_mask_overlay_magenta_outlines._plot_merged_plot._overlay_mask(image, mask)

Overlay the colored mask onto the original image.

spacr/plot.py:3902

plot_image_mask_overlay_magenta_outlines._save_channels_as_tiff(stack, save_dir, filename)

Save each channel in the stack as a grayscale TIFF.

spacr/plot.py:4020

plot_image_mask_overlay_magenta_outlines.random_color_cmap(n_labels, seed)

Generates a random color map for a given number of labels.

Parameters:
  • n_labels – How many object colours to draw. Index 0 is prepended as black for background, so the map holds n_labels + 1 entries; callers here pass int(outline.max() + 1) per object type, or int(combined_mask.max() + 1) for the merged panel. Colours are drawn as uniform RGB, so unlike the HSV variant in plot_image_mask_overlay() some come out dark and low-contrast against the image.

  • seed – Seeds the global numpy.random state, not a local generator, so passing it also shifts every later np.random draw in the process. Callers here pass a fresh random.randint(0, 100) per panel, which is why the same object gets a different colour in each panel and each run. None leaves the global state alone.

Returns:

A ListedColormap.

spacr/plot.py:3849

plot_images_and_arrays.find_files(folders, extensions)

Return a dict keyed by base filename mapping to files with the requested extensions.

Parameters:
  • folders – Folder paths, each walked recursively. Grouping is by basename without extension, and only names found in every folder survive the final filter — one missing file drops that name from the result entirely, and two files with the same basename under one folder keep only the last one walked.

  • extensions – Extensions to accept, matched with str.endswith so they must include the dot and match case.

Returns:

{basename: {folder: path}} for complete groups only.

spacr/plot.py:4480

plot_images_and_arrays.normalize_image(image, lower=1, upper=99)

Percentile-clip and rescale image to [0, 1].

Parameters:
  • image – Any numeric array; normalisation is over the whole array at once, so a multi-channel stack is scaled by a single pair of percentiles rather than per channel.

  • lower – Lower percentile, in 0-100. Default 1.

  • upper – Upper percentile, in 0-100, and must be strictly greater than lower: when the two percentiles evaluate equal (a flat image) the rescale divides by zero and returns nan rather than a blank frame, and swapping the two inverts the image instead of raising. Default 99.

Returns:

A float array clipped to [0, 1].

spacr/plot.py:4463

plot_images_and_arrays.plot_from_file_dict(file_dict, threshold=1000, lower_percentile=1, upper_percentile=99, overlay=False)

Show image/mask pairs collected in file_dict side-by-side.

Parameters:
  • file_dict – {filename: {folder: path}} produced by find_files.

  • threshold – Values above this unique-count are treated as intensity images; otherwise as label masks. Default 1000.

  • lower_percentile – Lower percentile clip. Default 1.

  • upper_percentile – Upper percentile clip. Default 99.

  • overlay – If True, overlay mask outlines on the image. Default False.

Returns:

None

spacr/plot.py:4507

plot_lorenz_curves.gini_coefficient(data)

Calculate Gini coefficient from data.

Parameters:

data – 1D array of non-negative counts, sorted internally. It is normalised by np.sum(data), so an all-zero input yields nan. Unlike lorenz_curve, an empty array does not raise here — it silently returns 1.0, the value for maximum inequality — so filter empty plates out upstream.

Returns:

The Gini coefficient as a float, from 0.0 for a perfectly even distribution up towards 1.0 as the counts concentrate on a few gRNAs. The area is taken with the trapezoid rule, so an even distribution reports exactly 0.0 rather than 1/n.

spacr/plot.py:6391

plot_lorenz_curves.lorenz_curve(data)

Calculate Lorenz curve.

Parameters:

data – 1D array of non-negative counts; it is sorted here, so the caller’s order does not matter. The curve is normalised by the running total’s last element, so an all-zero input divides by zero and an input mixing signs is not a Lorenz curve at all. Must be non-empty — an empty array indexes past the end.

Returns:

len(data) + 1 cumulative shares rising from 0 to 1, one longer than the input because the origin is prepended.

spacr/plot.py:6374

plot_lorenz_curves.remove_outliers_by_wells(data, name_col, wells_col)

Remove outliers based on 95% confidence interval for well counts.

Parameters:
  • data – DataFrame with one row per well-and-name observation. Whole names are kept or dropped together, never individual rows, so the surviving frame still has every well of every name it keeps.

  • name_col – Column identifying the gRNA (or other name). Rows are grouped on it and the group size — the number of wells a name appears in — is what the fence is applied to, so the count column’s values play no part in this filter.

  • wells_col – Accepted so the call reads symmetrically with the enclosing function’s value_column, but never read: the well count is derived from the group sizes above. Passing a wrong or missing column name changes nothing.

Returns:

data restricted to the names inside the fence. The fence is 1.5 * the 5th-to-95th-percentile spread, not the interquartile range, so it is far wider than a textbook IQR rule and its lower edge is usually negative — in practice only unusually widespread names are removed.

spacr/plot.py:6412

plot_region._sort_paths_by_basename(paths)

Return paths sorted by their basename.

spacr/plot.py:8638

plot_region.save_figure_as_pdf(fig, path)

Save fig in the user’s chosen figure format.

Named for the format it used to hard-code; it follows the preference now, like every other figure the user keeps, and save_figure creates the parent directory itself.

Parameters:
  • fig – Figure to write. It is left open, so the caller can still return it to the notebook after saving.

  • path – Destination path. Its extension is rewritten to whichever format the preference selected, so passing a .pdf name does not force PDF; missing parent directories are created. The path actually written is printed, not returned.

spacr/plot.py:8642

plot_resize.prepare_image(img)

Return (display_array, cmap) handling 2D/3D input shapes.

Parameters:

img – A 2D array, or a 3D array in channels-last order. One channel is squeezed to 2D and three or four are passed through as RGB/RGBA with a None colormap; any other channel count (a 5-channel spaCR stack, or a channels-first array read straight off disk) falls back to the mean across the last axis, which is a legal but usually misleading picture.

Returns:

(array, cmap) to hand straight to imshow, where cmap is None for true-colour data.

Raises:

ValueError – if img is neither 2D nor 3D.

spacr/plot.py:6060

print_mask_and_flows.apply_contours_on_image(image, mask, color=(255, 0, 0), thickness=2)

Draw the contours on the original image.

Parameters:
  • image – Base image. A 2D array is normalised to uint8 and promoted to RGB first, which assumes it is already scaled to [0, 1]; anything already 3D is copied and drawn on as-is, so the caller owns its dtype and value range.

  • mask – Label mask the outlines come from, traced with generate_contours. It must line up pixel-for-pixel with image, so resize both with the same max_size.

  • color – Contour colour as a BGR/RGB triple in 0-255, matching however the image channels are ordered. Default (255, 0, 0).

  • thickness – Line width in pixels; a negative value fills each contour solid instead of outlining it. Default 2.

Returns:

A new RGB array; the input image is not modified.

spacr/plot.py:5960

print_mask_and_flows.generate_contours(mask)

Generate contours for each object in the mask using OpenCV.

Parameters:

mask – Label mask, cast to uint8 before tracing — labels above 255 wrap around, and because only external contours are retrieved, touching objects trace as one outline and holes inside an object are not outlined.

Returns:

The OpenCV contour list, ready for cv2.drawContours.

spacr/plot.py:5948

print_mask_and_flows.normalize_to_uint8(image)

Normalize and convert image to uint8.

Parameters:

image – Array whose values are assumed to be scaled to [0, 1] already — the function only clips and multiplies by 255, it does not rescale. Raw 16-bit camera data therefore saturates to solid white apart from its zero pixels, and negative values clip to black.

Returns:

A uint8 array of the same shape.

spacr/plot.py:5984

print_mask_and_flows.resize_if_needed(image, max_size)

Resize image if any dimension exceeds max_size while maintaining aspect ratio.

Parameters:
  • image – 2D or (H, W, C) array. The channel axis is left untouched, and the result is cast back to the input dtype, so a label mask keeps integer labels — but the interpolation is anti-aliased, which can invent label values that belong to no object along object borders.

  • max_size – Cap on the larger of height and width, in pixels. The image is returned unchanged when it already fits, so no upscaling ever happens; a non-positive value drives the scale factor to zero, so pass a real pixel budget.

Returns:

The resized array, or image itself when it fits.

spacr/plot.py:5926

spacrGraph.create_plot._generate_tabels(unique_groups)

Generate row labels and a symbol table for multi-level grouping.

spacr/plot.py:7624

spacrGraph.create_plot._get_positions(self, ax)

Return plotted group centers in left-to-right table order.

spacr/plot.py:7678

spacrGraph.create_plot._place_symbols(row_labels, transposed_table, x_positions, ax)

Places symbols and row labels aligned under the bars or jitter points on the graph.

Parameters: - row_labels: List of row titles to be displayed along the y-axis. - transposed_table: Data to be placed under each bar/jitter as symbols. - x_positions: X-axis positions for each group to align the symbols. - ax: The matplotlib Axes object where the plot is drawn.

spacr/plot.py:7650

volcano_plot._as_numeric(s: pd.Series, colname: str) → np.ndarray

Coerce a column to floats, refusing an entirely nonnumeric result.

spacr/plot.py:9412

volcano_plot._read_table_auto(path: str) → pd.DataFrame

Read Excel or delimited text, sniffing comma versus tab as fallback.

spacr/plot.py:9385

volcano_plot._threshold_x_in_plot_units(thresh: float) → float

Convert a raw fold-change threshold to absolute plotted units.

spacr/plot.py:9454

volcano_plot._threshold_y_in_plot_units(pthresh: float) → float

Validate and transform a raw p-value threshold for the y axis.

spacr/plot.py:9467

volcano_plot._transform_x(x: np.ndarray, mode: str) → np.ndarray

Apply the selected x transform, requiring positive log inputs.

spacr/plot.py:9419

volcano_plot._transform_y(p: np.ndarray, mode: str) → np.ndarray

Apply the selected y transform after clipping logarithm inputs.

spacr/plot.py:9437