spacr.model_compare

Workflow inputs and outputs

Model Compare

Compare model masks on identical fields. Without independent reference labels, agreement does not measure segmentation accuracy.

Open: Make Masks → Model Compare.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.

  • Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.

  • Segmentation checkpoint — Saved Cellpose-compatible checkpoint or a compatible installed backend selected with its own configuration.

Outputs

  • Run/model comparison — Comparison tables and figures from compatible saved runs or masks; agreement is not ground-truth accuracy.

API reference.

Module tutorial.

Run two segmentation models over the same fields and say what changed.

Why this exists

Choosing a segmentation model in spaCR is currently done by running a whole plate twice and squinting at the montages. That is hours per candidate, and the comparison it produces is “these look about the same” — which is exactly the answer you get whether the models differ or not.

This module makes the comparison small and quantitative: three fields, two models, one table. It is deliberately split in two halves so the Model Zoo (TO-DO #29, “test a model on 3 fields”) can reuse the whole thing:

  • the metric layer (compare_masks() and everything under it) takes label arrays and nothing else. No Cellpose, no torch, no GUI, no file paths.

  • the orchestration layer (compare_models()) runs a segmentation callable twice and folds the per-field results into a ComparisonReport. The callable is an argument, so the zoo can hand it a different backend, and a test can hand it a stub.

Neither half imports torch or cellpose at module level; only segment_with_cellpose() does, and only when it is called.

Neither model is ground truth

This is an A/B comparison, not an evaluation. There is no correct mask here, so nothing in this module calls one model right and the other wrong: every count is reported directionally (“B found 12 objects A did not”), never as precision or recall, and the wording of format_comparison() follows the same rule. The only symmetric summary numbers are the ones that genuinely are symmetric — the ARI and the matched fraction.

The metrics, and why each one is computed the way it is

ari — background is excluded, and that is the whole trick.

The Adjusted Rand Index over two raw label images is nearly useless: a 1400x1400 field is ~95 % background, both models agree it is background, and those agreed pairs swamp everything else. Two models that disagree about every single object still score about 0.999 (tests/ test_model_compare.py asserts exactly that number on a field where the foreground agreement is 0.0).

So the index is computed over the union foreground — the pixels at least one model assigned to an object — with pixels the other model left unassigned treated as unclustered singletons rather than as one big background cluster. That second half matters as much as the first: with background as a single cluster, a model that misses an entire object still scores 1.0, because “one object plus a background blob” and “one object plus a background blob” are the same partition. As singletons, missing half the objects scores 0.50, which is what it deserves.

Computed in closed form from the object overlap matrix (see adjusted_rand_index()) rather than by materialising millions of singleton labels; tests/test_model_compare.py checks it against sklearn.metrics.adjusted_rand_score on the expanded arrays.

ARI is a pixel-pair index, so it is sensitive to boundaries and it is degenerate when a field holds one object (a single cluster has no pair structure to agree about). It is reported next to the object-level numbers for that reason, never alone.

iou_matched_fraction and mean_matched_iou — from an optimal

assignment, not a greedy one. Matching objects between two segmentations is a bipartite assignment problem; picking each object’s best partner double-assigns, and picking greedily in descending IoU can leave a pair stranded that an optimal assignment would have kept. Both failures are reproduced in the tests. match_objects() thresholds the IoU matrix and then runs scipy.optimize.linear_sum_assignment() over it, which is also what cellpose.metrics does.

Above an IoU of 0.5 the assignment is provably unique — two objects cannot both overlap a third by more than half of it — so at the default threshold greedy would in fact get the same answer. The optimal assignment is used anyway because iou_threshold is a knob, and every value below 0.5 (which is where “roughly the same object” lives) makes greedy wrong.

split_events / merge_events — fragmentation is not discovery.

“Model B found 20 more objects” means completely different things if they are 20 new cells or 20 fragments of cells A already found, and fragmentation is the common Cellpose failure. So the object-count delta is decomposed:

  • a B object is a fragment of an A object when the majority of it (see containment) lies inside that A object and it was not assigned to some other A object. An A object with two or more such fragments is one split_event, and it explains k - 1 of B’s extra objects.

  • the mirror image gives merge_events at B objects that swallow two or more A objects, explaining k - 1 of A’s objects going missing.

  • whatever is left over — new_objects_b and missing_objects_a — is the genuine difference in what the two models detected.

The “not assigned elsewhere” clause is what keeps the attribution honest: a B object that straddles two A objects but is paired with one of them is that object’s counterpart, not a fragment of its neighbour.

qc_a / qc_b

Each field’s masks are additionally run through spacr.seg_qc, so the table can say which of the two disagreeing masks looks broken on its own terms (fused, shattered, empty, all on the border). Those thresholds are argued in that module and are not duplicated here.

Degenerate fields, and what they are defined to be

  • both masks empty — ari = 1.0, iou_matched_fraction = 1.0, mean_matched_iou = nan. Two models that both say “there is nothing in this field” have made the same statement, and that is agreement; the alternative (nan) would drop the field out of every aggregate, so a channel that is legitimately empty would silently shrink the sample instead of showing up as the unanimous verdict it is. It is counted separately in ComparisonReport.n_both_empty so it can never be mistaken for agreement about objects, and mean_matched_iou stays nan because there is no matched pair to take an IoU of.

  • one mask empty — ari = 0.0 and no matches; falls out of the definitions with no special case, and is the right answer: the models agree about nothing.

  • one object each — matched normally, but the ARI is degenerate (a single cluster carries no pair information) and can even be negative for two masks that overlap well. The object-level numbers carry the field in that case.

  • completely disjoint labels — ARI near zero, nothing matched, every object reported as new/missing.

What Cellpose 4 accepts and then ignores

A comparison that differs only in an argument the model never sees reports “no difference” and wastes the run. On the installed Cellpose 4, these are accepted and then dropped (see IGNORED_ARGUMENTS): model_type, diam_mean and nchan at construction, channels and rescale at eval, plus spaCR’s own restore_type. Every pre-SAM model name resolves to cpsam too, so “cyto3 versus nuclei” is one model against itself.

diameter is not in that list: eval still honours it by rescaling the image by 30 / diameter before inference. It is the one size knob that does anything, which is why format_comparison() prints it first.

compare_models() therefore records both what reached the model and what was dropped on the floor, and raises a loud warning when the two configurations differ only in arguments nothing will read.

Public API

ModelConfig one model’s settings, plus what of it survives. SegComparison one field, two masks: every number above. ComparisonReport the whole run: configs, per-field rows, aggregates. compare_masks the pure metric entry point (label arrays in). compare_models run two models over the same fields. adjusted_rand_index background-excluded ARI on its own. match_objects the optimal object assignment on its own. segment_with_cellpose the default segmentation backend. load_fields pull N fields out of a folder for a comparison. format_comparison the printable report.

Classes

ComparisonReport

Two models over a set of fields: the configs, the rows, the aggregate.

ModelConfig

One side of the comparison: which model, run how.

SegComparison

One field, two masks — every number the comparison produces.

Functions

adjusted_rand_index(→ float)

Adjusted Rand Index over the union foreground, background excluded.

compare_configs(→ Dict[str, Any])

Diff two configurations into what matters and what cannot matter.

compare_masks(→ SegComparison)

Compare two label images of the same field. Pure: arrays in, numbers out.

compare_models(→ ComparisonReport)

Run two models over the same fields and compare what they produced.

format_comparison(→ str)

Render the report a human reads before picking a model.

iou_matrix(→ numpy.ndarray)

Object-by-object IoU from an object_overlap() table.

load_fields(→ Tuple[List[str], List[numpy.ndarray]])

Pull the first n_fields images out of a folder (or a list).

match_objects(→ Dict[str, Any])

Pair the objects of two masks by optimal assignment.

object_overlap(→ Dict[str, Any])

Contingency between the objects of two label images.

segment_with_cellpose(→ List[numpy.ndarray])

Segment images with one Cellpose model. The default backend.

Module Contents

class spacr.model_compare.ComparisonReport[source]

Two models over a set of fields: the configs, the rows, the aggregate.

Parameters:
  • model_a – the A configuration.

  • model_b – the B configuration.

  • comparisons – one SegComparison per field, in field order.

  • config_diff – what compare_configs() found.

  • warnings – the lines a reader must see before the numbers — ignored arguments, remapped model names, an A/B that cannot differ.

  • seconds_a – wall-clock seconds model A spent segmenting.

  • seconds_b – the same for B.

  • masks_a – A’s label images, kept so a GUI can draw them.

  • masks_b – B’s label images.

  • images – the source images, kept for the same reason.

  • object_type – what was segmented, for the seg_qc scorecards.

property count_ratio: float[source]

B’s object count as a multiple of A’s; nan when A found none.

property fields: List[str][source]

The field names, in order.

property identical_masks: bool[source]

True when every field’s masks agree object-for-object and pixel-pair.

property mean_ari: float[source]

Mean ARI over the fields that have one; nan when none do.

property mean_matched_fraction: float[source]

The mean fraction of A’s objects that B also found.

NAN-SAFE: a field where neither model found anything contributes no fraction rather than a zero, which would drag the mean down for a field that says nothing about either model.

Returns:

the mean, or NaN when no field had objects.

property mean_matched_iou: float[source]

Mean of the per-field mean matched IoU.

property n_both_empty: int[source]

Fields where neither model found anything — trivial agreement.

property n_fields: int[source]

How many fields the two models were compared over.

Returns:

the field count.

property object_count_delta: int[source]

total_objects_b - total_objects_a, directional.

property summary: str[source]

One directional sentence about the whole run.

property total_fragments: int[source]

Extra B objects that are pieces of A objects, not new detections.

property total_merged_away: int[source]

How many of A’s objects disappeared into a merge.

Distinct from the merge COUNT: one merge can swallow several objects, and the number of objects lost is what changes a per-object measurement downstream.

Returns:

the object count.

property total_merges: int[source]

How many of A’s objects B joined together.

Returns:

the merge count.

property total_missing_objects_a: int[source]

Objects A found that B did not.

Returns:

the object count.

property total_new_objects_b: int[source]

Objects B found that A did not.

Returns:

the object count.

property total_objects_a: int[source]

Every object model A found, across all fields.

Returns:

the object count.

property total_objects_b: int[source]

Every object model B found, across all fields.

Returns:

the object count.

property total_splits: int[source]

How many of A’s objects B broke into several.

DIRECTIONAL. A split and a merge are the same event seen from the two sides, so the pair only means anything if you know which model is A.

Returns:

the split count.

class spacr.model_compare.ModelConfig[source]

One side of the comparison: which model, run how.

Only the fields below and the honoured keys of extra reach CellposeModel.eval. Anything in extra that Cellpose 4 ignores is kept, reported and not passed on, so it shows up in the report as the no-op it is instead of quietly making two runs look identical.

Parameters:
  • name – what to call this side in the report. Defaults to the model.

  • model – 'cpsam', a legacy Cellpose name (resolved to cpsam), or a path to a custom checkpoint.

  • diameter – expected object diameter in pixels, or None to let Cellpose run at native scale. This is the one size argument Cellpose 4 still acts on — it resizes the image by 30 / diameter.

  • flow_threshold – flow-error cutoff.

  • cellprob_threshold – mask-probability cutoff.

  • normalize – per-image percentile normalisation inside Cellpose.

  • invert – invert the image before inference.

  • resample – run the dynamics at full resolution.

  • min_size – objects smaller than this are dropped by Cellpose.

  • niter – dynamics iterations, or None for Cellpose’s default.

  • augment – 8-way test-time augmentation.

  • batch_size – tiles per forward pass — speed only, not results.

  • extra – any other eval keyword. Honoured keys are forwarded; ignored ones are reported.

__post_init__()[source]

Give the model a name if the caller did not.

The checkpoint’s file name is used, falling back to the path itself – a comparison report that says “Model A” twice is unreadable.

eval_kwargs() → Dict[str, Any][source]

The keyword arguments to hand CellposeModel.eval.

honoured_parameters minus model, which is a constructor argument rather than an eval one.

classmethod from_mapping(source: Any) → ModelConfig[source]

Build a config from a dict (or pass a ModelConfig through).

Keys that are not fields of this class land in extra, which is how an ignored argument such as diam_mean survives long enough to be reported instead of silently doing nothing.

Parameters:

source – a mapping or an existing ModelConfig.

Returns:

a ModelConfig.

honoured_parameters() → Dict[str, Any][source]

Everything that reaches the model, resolved model included.

This is what the report displays. If two configurations produce the same dict here, they will produce the same masks, and the comparison has nothing to show.

ignored_parameters() → Dict[str, Any][source]

What was set and will not be read, {name: value}.

The requested model appears here as model when it was remapped: asking for cyto3 and getting cpsam is the same class of surprise as setting diam_mean.

notes() → List[str][source]

One line per argument of this config that will not be read.

property model_was_remapped: bool[source]

True when the requested model is not the model that will be loaded.

property resolved_model: str[source]

The checkpoint Cellpose will actually load.

Every pre-SAM name maps to cpsam; a path is left alone so a custom checkpoint is loaded as asked.

class spacr.model_compare.SegComparison[source]

One field, two masks — every number the comparison produces.

Directional by construction: *_a describes the first mask, *_b the second, and neither is treated as the truth. unmatched_a is “objects A found that B did not pair with”, not “false negatives”.

Parameters:
  • field – the field’s name.

  • n_objects_a – labels in mask A.

  • n_objects_b – labels in mask B.

  • ari – background-excluded Adjusted Rand Index (see adjusted_rand_index()). 1.0 for identical masks, ~0 for unrelated ones, 1.0 for two empty masks, nan only when the union foreground holds fewer than two pixels.

  • iou_matched_fraction – 2 * matched / (n_a + n_b) — the symmetric share of objects that have a partner. Symmetric on purpose: it says how much of the two segmentations correspond without calling either right.

  • mean_matched_iou – mean IoU over matched pairs; nan with none.

  • unmatched_a – A objects with no partner.

  • unmatched_b – B objects with no partner.

  • split_events – A objects that B broke into two or more pieces.

  • merge_events – B objects that swallowed two or more A objects.

  • n_matched – pairs in the optimal assignment above the threshold.

  • fragments_from_splits – extra B objects explained by split_events.

  • merged_away – A objects that disappeared into merge_events.

  • new_objects_b – B objects that are neither matched nor fragments — the genuinely new detections.

  • missing_objects_a – A objects neither matched nor merged away.

  • iou_threshold – the threshold the matching used.

  • matches – [(label_a, label_b, iou), ...], for drawing.

  • qc_a – spacr.seg_qc.FieldQC for mask A, when computed.

  • qc_b – the same for mask B.

  • note – the field’s verdict in prose, with its numbers in it.

__str__() → str[source]

Return the one-line summary: the field, both counts, the delta and the ARI.

property both_empty: bool[source]

True when neither model found anything in this field.

property object_count_delta: int[source]

n_objects_b - n_objects_a. Positive means B found more.

spacr.model_compare.adjusted_rand_index(mask_a: Any, mask_b: Any) → float[source]

Adjusted Rand Index over the union foreground, background excluded.

The index is taken over the pixels at least one mask assigned to an object. Pixels the other mask left unassigned stay in the sample but belong to no cluster (each is its own singleton), so a model that misses an object is penalised for it — with background as one shared cluster it would not be. Pixels neither mask claimed are dropped entirely: they are the agreement that would otherwise be the whole answer.

Computed in closed form from the object overlap table. Singleton clusters contribute nothing to any C(n, 2) term, so the expansion never has to be materialised; the only thing they change is the total pair count, which is taken over the union foreground.

Parameters:
  • mask_a – a 2-D label image.

  • mask_b – a second label image of the same shape.

Returns:

1.0 for identical partitions, ~0 for unrelated ones, negative when the two masks agree less than chance. 1.0 when both masks are empty (see the module docstring for why), nan when the union foreground holds fewer than two pixels and there is no pair to score.

spacr.model_compare.compare_configs(config_a: ModelConfig, config_b: ModelConfig) → Dict[str, Any][source]

Diff two configurations into what matters and what cannot matter.

Parameters:
  • config_a – the A side.

  • config_b – the B side.

Returns:

{'honoured': {key: (a, b)}, 'ignored': {key: (a, b)}, 'identical': bool, 'warnings': [str]}. identical is True when every argument that reaches the model is the same on both sides — the case where the run cannot show a difference and the report has to say so before anybody reads a number off it.

spacr.model_compare.compare_masks(mask_a: Any, mask_b: Any, field: str = 'field', iou_threshold: float = DEFAULT_IOU_THRESHOLD, containment: float = DEFAULT_CONTAINMENT, qc_a: spacr.seg_qc.FieldQC | None = None, qc_b: spacr.seg_qc.FieldQC | None = None) → SegComparison[source]

Compare two label images of the same field. Pure: arrays in, numbers out.

Nothing about Cellpose, models, files or the GUI reaches this function, so it works just as well on masks from any other source — which is the point: the Model Zoo reuses it unchanged.

The comparison is directional but not judgemental. *_a and *_b describe the two masks; neither is the reference, so there is no precision, no recall, and no “correct” column. See the module docstring for how the ARI excludes background, why the object assignment is optimal rather than greedy, and how splits and merges are told apart from genuine differences.

Parameters:
  • mask_a – a 2-D label image (bool and float masks are coerced).

  • mask_b – a second label image of the same shape.

  • field – the field’s name, carried into the row.

  • iou_threshold – minimum IoU for two objects to be the same object.

  • containment – fraction of an object that must lie inside another for it to count as a fragment of it.

  • qc_a – an optional spacr.seg_qc.FieldQC for mask A, attached so the row can say which mask looks broken on its own terms.

  • qc_b – the same for mask B.

Returns:

a SegComparison.

Raises:

ValueError – when the masks are not the same shape, or are not label images.

spacr.model_compare.compare_models(images: Sequence[numpy.ndarray], model_a: Any, model_b: Any, field_names: Sequence[str] | None = None, segment_fn: Callable[[Sequence[numpy.ndarray], ModelConfig], Sequence[numpy.ndarray]] | None = None, iou_threshold: float = DEFAULT_IOU_THRESHOLD, containment: float = DEFAULT_CONTAINMENT, object_type: str = 'cell', qc: bool = True, keep_images: bool = True, progress: Callable[[str, int, int], None] | None = None) → ComparisonReport[source]

Run two models over the same fields and compare what they produced.

Each model is loaded once and run over every field, A first, then B — one model in memory at a time, and the per-model wall clock in the report is therefore a fair (if small-sample) comparison of their cost.

The segmentation call is an argument. That is what makes this reusable: the Model Zoo can hand in a different backend, and a test can hand in a stub, so nothing about this function needs Cellpose to be exercised.

Parameters:
  • images – one array per field.

  • model_a – a ModelConfig or a mapping (see ModelConfig.from_mapping()).

  • model_b – the other side.

  • field_names – names for the rows; defaults to field_0000….

  • segment_fn – fn(images, config) -> masks; defaults to segment_with_cellpose().

  • iou_threshold – passed to compare_masks().

  • containment – passed to compare_masks().

  • object_type – what is being segmented, for the seg_qc scorecards.

  • qc – score each model’s masks with spacr.seg_qc as well.

  • keep_images – keep the images and masks on the report so a GUI can draw them. Pass False for a headless sweep over many fields.

  • progress – fn(message, done, total), called as the run proceeds.

Returns:

a ComparisonReport.

Raises:

ValueError – when there is no field, or a model returns the wrong number of masks.

spacr.model_compare.format_comparison(report: ComparisonReport) → str[source]

Render the report a human reads before picking a model.

Order is deliberate: the warnings first (a run that cannot show a difference must say so before anybody reads a number off it), then the parameters that reached each model, then the per-field table, then the aggregate — and the aggregate is phrased directionally, because neither model is the truth.

Parameters:

report – what compare_models() returned.

Returns:

a multi-line string, ready to print.

spacr.model_compare.iou_matrix(parts: Mapping[str, Any]) → numpy.ndarray[source]

Object-by-object IoU from an object_overlap() table.

Parameters:

parts – the dict object_overlap() returned.

Returns:

an n_a x n_b float array; empty when either mask has no object.

spacr.model_compare.load_fields(source: Any, n_fields: int = DEFAULT_N_FIELDS, channel: int | None = None) → Tuple[List[str], List[numpy.ndarray]][source]

Pull the first n_fields images out of a folder (or a list).

Handles the shapes spaCR actually leaves on disk: a folder of .tif / .png fields, a folder of .npy arrays, and the .npz batches the Mask module writes (data + filenames). Reading stops as soon as n_fields images are in hand, so pointing this at a 1536-field plate costs three files.

Parameters:
  • source – a folder, or an already-loaded sequence of arrays.

  • n_fields – how many fields to take.

  • channel – index into the last axis for a multi-channel field; None keeps the array as it is.

Returns:

(names, images).

Raises:
spacr.model_compare.match_objects(mask_a: Any, mask_b: Any, iou_threshold: float = DEFAULT_IOU_THRESHOLD) → Dict[str, Any][source]

Pair the objects of two masks by optimal assignment.

Bipartite, not greedy. The IoU matrix is thresholded first (everything below iou_threshold becomes 0) and scipy.optimize.linear_sum_assignment() then maximises the total IoU over what is left, which is the same procedure cellpose.metrics._true_positive uses. Taking each object’s best partner double-assigns; taking pairs greedily in descending IoU can strand a pair that the optimal assignment keeps. Both are exercised in the tests.

Parameters:
  • mask_a – a 2-D label image.

  • mask_b – a second label image of the same shape.

  • iou_threshold – minimum IoU for a pair to count as the same object.

Returns:

{'matches': [(label_a, label_b, iou), ...], 'iou': matrix, 'parts': the overlap table, 'unmatched_a': [labels], 'unmatched_b': [labels]}, matches sorted by descending IoU.

spacr.model_compare.object_overlap(mask_a: Any, mask_b: Any) → Dict[str, Any][source]

Contingency between the objects of two label images.

Both masks are relabelled to 0..n first, so arbitrary label values (and gaps in them) are handled, and the pixel-pair counting below can index straight into the table.

Parameters:
  • mask_a – a 2-D label image.

  • mask_b – a second label image of the same shape.

Returns:

{'labels_a', 'labels_b', 'areas_a', 'areas_b', 'overlap', 'n_pixels', 'union_foreground'}. overlap is the n_a x n_b object-by-object intersection count with background already dropped; labels_a maps row i back to the original label.

Raises:

ValueError – when the two masks are not the same shape.

spacr.model_compare.segment_with_cellpose(images: Sequence[numpy.ndarray], config: ModelConfig) → List[numpy.ndarray][source]

Segment images with one Cellpose model. The default backend.

Only ModelConfig.eval_kwargs() is forwarded, so an argument Cellpose 4 ignores never reaches eval — it is reported by the caller instead of being passed on to be silently dropped, which is the difference between a comparison that explains itself and one that says “no difference”.

Model and device dependencies are loaded on demand for this call. Cellpose versions without a separate invert argument receive that setting in the normalization dictionary. Inversion requires normalization on those versions; other normalization options are retained.

Parameters:
  • images – 2-D or 3-D arrays, one per field.

  • config – the model to run.

Returns:

one integer label image per input image.

Nested helpers

compare_models._tick(message: str, done: int) → None

Forward message and done with the fixed total; return None.

spacr/model_compare.py:1283