spacr.model_compare¶
Workflow inputs and outputs¶
Model Compare¶
Compare model masks on identical fields. Without independent reference labels, agreement does not measure segmentation accuracy.
Open: Make Masks → Model Compare.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.
Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.
Segmentation checkpoint — Saved Cellpose-compatible checkpoint or a compatible installed backend selected with its own configuration.
Outputs
Run/model comparison — Comparison tables and figures from compatible saved runs or masks; agreement is not ground-truth accuracy.
Run two segmentation models over the same fields and say what changed.
Why this exists¶
Choosing a segmentation model in spaCR is currently done by running a whole plate twice and squinting at the montages. That is hours per candidate, and the comparison it produces is “these look about the same” — which is exactly the answer you get whether the models differ or not.
This module makes the comparison small and quantitative: three fields, two models, one table. It is deliberately split in two halves so the Model Zoo (TO-DO #29, “test a model on 3 fields”) can reuse the whole thing:
the metric layer (
compare_masks()and everything under it) takes label arrays and nothing else. No Cellpose, no torch, no GUI, no file paths.the orchestration layer (
compare_models()) runs a segmentation callable twice and folds the per-field results into aComparisonReport. The callable is an argument, so the zoo can hand it a different backend, and a test can hand it a stub.
Neither half imports torch or cellpose at module level; only
segment_with_cellpose() does, and only when it is called.
Neither model is ground truth¶
This is an A/B comparison, not an evaluation. There is no correct mask here, so
nothing in this module calls one model right and the other wrong: every count
is reported directionally (“B found 12 objects A did not”), never as
precision or recall, and the wording of format_comparison() follows the
same rule. The only symmetric summary numbers are the ones that genuinely are
symmetric — the ARI and the matched fraction.
The metrics, and why each one is computed the way it is¶
ari— background is excluded, and that is the whole trick.The Adjusted Rand Index over two raw label images is nearly useless: a 1400x1400 field is ~95 % background, both models agree it is background, and those agreed pairs swamp everything else. Two models that disagree about every single object still score about 0.999 (
tests/ test_model_compare.pyasserts exactly that number on a field where the foreground agreement is 0.0).So the index is computed over the union foreground — the pixels at least one model assigned to an object — with pixels the other model left unassigned treated as unclustered singletons rather than as one big background cluster. That second half matters as much as the first: with background as a single cluster, a model that misses an entire object still scores 1.0, because “one object plus a background blob” and “one object plus a background blob” are the same partition. As singletons, missing half the objects scores 0.50, which is what it deserves.
Computed in closed form from the object overlap matrix (see
adjusted_rand_index()) rather than by materialising millions of singleton labels;tests/test_model_compare.pychecks it againstsklearn.metrics.adjusted_rand_scoreon the expanded arrays.ARI is a pixel-pair index, so it is sensitive to boundaries and it is degenerate when a field holds one object (a single cluster has no pair structure to agree about). It is reported next to the object-level numbers for that reason, never alone.
iou_matched_fractionandmean_matched_iou— from an optimalassignment, not a greedy one. Matching objects between two segmentations is a bipartite assignment problem; picking each object’s best partner double-assigns, and picking greedily in descending IoU can leave a pair stranded that an optimal assignment would have kept. Both failures are reproduced in the tests.
match_objects()thresholds the IoU matrix and then runsscipy.optimize.linear_sum_assignment()over it, which is also whatcellpose.metricsdoes.Above an IoU of 0.5 the assignment is provably unique — two objects cannot both overlap a third by more than half of it — so at the default threshold greedy would in fact get the same answer. The optimal assignment is used anyway because
iou_thresholdis a knob, and every value below 0.5 (which is where “roughly the same object” lives) makes greedy wrong.split_events/merge_events— fragmentation is not discovery.“Model B found 20 more objects” means completely different things if they are 20 new cells or 20 fragments of cells A already found, and fragmentation is the common Cellpose failure. So the object-count delta is decomposed:
a B object is a fragment of an A object when the majority of it (see
containment) lies inside that A object and it was not assigned to some other A object. An A object with two or more such fragments is onesplit_event, and it explainsk - 1of B’s extra objects.the mirror image gives
merge_eventsat B objects that swallow two or more A objects, explainingk - 1of A’s objects going missing.whatever is left over —
new_objects_bandmissing_objects_a— is the genuine difference in what the two models detected.
The “not assigned elsewhere” clause is what keeps the attribution honest: a B object that straddles two A objects but is paired with one of them is that object’s counterpart, not a fragment of its neighbour.
qc_a/qc_bEach field’s masks are additionally run through
spacr.seg_qc, so the table can say which of the two disagreeing masks looks broken on its own terms (fused, shattered, empty, all on the border). Those thresholds are argued in that module and are not duplicated here.
Degenerate fields, and what they are defined to be¶
both masks empty —
ari = 1.0,iou_matched_fraction = 1.0,mean_matched_iou = nan. Two models that both say “there is nothing in this field” have made the same statement, and that is agreement; the alternative (nan) would drop the field out of every aggregate, so a channel that is legitimately empty would silently shrink the sample instead of showing up as the unanimous verdict it is. It is counted separately inComparisonReport.n_both_emptyso it can never be mistaken for agreement about objects, andmean_matched_ioustaysnanbecause there is no matched pair to take an IoU of.one mask empty —
ari = 0.0and no matches; falls out of the definitions with no special case, and is the right answer: the models agree about nothing.one object each — matched normally, but the ARI is degenerate (a single cluster carries no pair information) and can even be negative for two masks that overlap well. The object-level numbers carry the field in that case.
completely disjoint labels — ARI near zero, nothing matched, every object reported as new/missing.
What Cellpose 4 accepts and then ignores¶
A comparison that differs only in an argument the model never sees reports “no
difference” and wastes the run. On the installed Cellpose 4, these are accepted
and then dropped (see IGNORED_ARGUMENTS): model_type, diam_mean
and nchan at construction, channels and rescale at eval, plus
spaCR’s own restore_type. Every pre-SAM model name resolves to cpsam
too, so “cyto3 versus nuclei” is one model against itself.
diameter is not in that list: eval still honours it by rescaling the
image by 30 / diameter before inference. It is the one size knob that does
anything, which is why format_comparison() prints it first.
compare_models() therefore records both what reached the model and what
was dropped on the floor, and raises a loud warning when the two configurations
differ only in arguments nothing will read.
Public API¶
ModelConfig one model’s settings, plus what of it survives.
SegComparison one field, two masks: every number above.
ComparisonReport the whole run: configs, per-field rows, aggregates.
compare_masks the pure metric entry point (label arrays in).
compare_models run two models over the same fields.
adjusted_rand_index background-excluded ARI on its own.
match_objects the optimal object assignment on its own.
segment_with_cellpose the default segmentation backend.
load_fields pull N fields out of a folder for a comparison.
format_comparison the printable report.
Classes¶
Two models over a set of fields: the configs, the rows, the aggregate. |
|
One side of the comparison: which model, run how. |
|
One field, two masks — every number the comparison produces. |
Functions¶
|
Adjusted Rand Index over the union foreground, background excluded. |
|
Diff two configurations into what matters and what cannot matter. |
|
Compare two label images of the same field. Pure: arrays in, numbers out. |
|
Run two models over the same fields and compare what they produced. |
|
Render the report a human reads before picking a model. |
|
Object-by-object IoU from an |
|
Pull the first |
|
Pair the objects of two masks by optimal assignment. |
|
Contingency between the objects of two label images. |
|
Segment |
Module Contents¶
- class spacr.model_compare.ComparisonReport[source]¶
Two models over a set of fields: the configs, the rows, the aggregate.
- Parameters:
model_a – the A configuration.
model_b – the B configuration.
comparisons – one
SegComparisonper field, in field order.config_diff – what
compare_configs()found.warnings – the lines a reader must see before the numbers — ignored arguments, remapped model names, an A/B that cannot differ.
seconds_a – wall-clock seconds model A spent segmenting.
seconds_b – the same for B.
masks_a – A’s label images, kept so a GUI can draw them.
masks_b – B’s label images.
images – the source images, kept for the same reason.
object_type – what was segmented, for the seg_qc scorecards.
- property identical_masks: bool[source]¶
True when every field’s masks agree object-for-object and pixel-pair.
- property mean_matched_fraction: float[source]¶
The mean fraction of A’s objects that B also found.
NAN-SAFE: a field where neither model found anything contributes no fraction rather than a zero, which would drag the mean down for a field that says nothing about either model.
- Returns:
the mean, or NaN when no field had objects.
- property n_fields: int[source]¶
How many fields the two models were compared over.
- Returns:
the field count.
- property total_fragments: int[source]¶
Extra B objects that are pieces of A objects, not new detections.
- property total_merged_away: int[source]¶
How many of A’s objects disappeared into a merge.
Distinct from the merge COUNT: one merge can swallow several objects, and the number of objects lost is what changes a per-object measurement downstream.
- Returns:
the object count.
- property total_merges: int[source]¶
How many of A’s objects B joined together.
- Returns:
the merge count.
- property total_missing_objects_a: int[source]¶
Objects A found that B did not.
- Returns:
the object count.
- property total_new_objects_b: int[source]¶
Objects B found that A did not.
- Returns:
the object count.
- property total_objects_a: int[source]¶
Every object model A found, across all fields.
- Returns:
the object count.
- class spacr.model_compare.ModelConfig[source]¶
One side of the comparison: which model, run how.
Only the fields below and the honoured keys of
extrareachCellposeModel.eval. Anything inextrathat Cellpose 4 ignores is kept, reported and not passed on, so it shows up in the report as the no-op it is instead of quietly making two runs look identical.- Parameters:
name – what to call this side in the report. Defaults to the model.
model –
'cpsam', a legacy Cellpose name (resolved tocpsam), or a path to a custom checkpoint.diameter – expected object diameter in pixels, or None to let Cellpose run at native scale. This is the one size argument Cellpose 4 still acts on — it resizes the image by
30 / diameter.flow_threshold – flow-error cutoff.
cellprob_threshold – mask-probability cutoff.
normalize – per-image percentile normalisation inside Cellpose.
invert – invert the image before inference.
resample – run the dynamics at full resolution.
min_size – objects smaller than this are dropped by Cellpose.
niter – dynamics iterations, or None for Cellpose’s default.
augment – 8-way test-time augmentation.
batch_size – tiles per forward pass — speed only, not results.
extra – any other
evalkeyword. Honoured keys are forwarded; ignored ones are reported.
- __post_init__()[source]¶
Give the model a name if the caller did not.
The checkpoint’s file name is used, falling back to the path itself – a comparison report that says “Model A” twice is unreadable.
- eval_kwargs() Dict[str, Any][source]¶
The keyword arguments to hand
CellposeModel.eval.honoured_parametersminusmodel, which is a constructor argument rather than anevalone.
- classmethod from_mapping(source: Any) ModelConfig[source]¶
Build a config from a dict (or pass a
ModelConfigthrough).Keys that are not fields of this class land in
extra, which is how an ignored argument such asdiam_meansurvives long enough to be reported instead of silently doing nothing.- Parameters:
source – a mapping or an existing
ModelConfig.- Returns:
a
ModelConfig.
- honoured_parameters() Dict[str, Any][source]¶
Everything that reaches the model, resolved model included.
This is what the report displays. If two configurations produce the same dict here, they will produce the same masks, and the comparison has nothing to show.
- ignored_parameters() Dict[str, Any][source]¶
What was set and will not be read,
{name: value}.The requested model appears here as
modelwhen it was remapped: asking forcyto3and gettingcpsamis the same class of surprise as settingdiam_mean.
- class spacr.model_compare.SegComparison[source]¶
One field, two masks — every number the comparison produces.
Directional by construction:
*_adescribes the first mask,*_bthe second, and neither is treated as the truth.unmatched_ais “objects A found that B did not pair with”, not “false negatives”.- Parameters:
field – the field’s name.
n_objects_a – labels in mask A.
n_objects_b – labels in mask B.
ari – background-excluded Adjusted Rand Index (see
adjusted_rand_index()). 1.0 for identical masks, ~0 for unrelated ones, 1.0 for two empty masks,nanonly when the union foreground holds fewer than two pixels.iou_matched_fraction –
2 * matched / (n_a + n_b)— the symmetric share of objects that have a partner. Symmetric on purpose: it says how much of the two segmentations correspond without calling either right.mean_matched_iou – mean IoU over matched pairs;
nanwith none.unmatched_a – A objects with no partner.
unmatched_b – B objects with no partner.
split_events – A objects that B broke into two or more pieces.
merge_events – B objects that swallowed two or more A objects.
n_matched – pairs in the optimal assignment above the threshold.
fragments_from_splits – extra B objects explained by
split_events.merged_away – A objects that disappeared into
merge_events.new_objects_b – B objects that are neither matched nor fragments — the genuinely new detections.
missing_objects_a – A objects neither matched nor merged away.
iou_threshold – the threshold the matching used.
matches –
[(label_a, label_b, iou), ...], for drawing.qc_a –
spacr.seg_qc.FieldQCfor mask A, when computed.qc_b – the same for mask B.
note – the field’s verdict in prose, with its numbers in it.
- spacr.model_compare.adjusted_rand_index(mask_a: Any, mask_b: Any) float[source]¶
Adjusted Rand Index over the union foreground, background excluded.
The index is taken over the pixels at least one mask assigned to an object. Pixels the other mask left unassigned stay in the sample but belong to no cluster (each is its own singleton), so a model that misses an object is penalised for it — with background as one shared cluster it would not be. Pixels neither mask claimed are dropped entirely: they are the agreement that would otherwise be the whole answer.
Computed in closed form from the object overlap table. Singleton clusters contribute nothing to any
C(n, 2)term, so the expansion never has to be materialised; the only thing they change is the total pair count, which is taken over the union foreground.- Parameters:
mask_a – a 2-D label image.
mask_b – a second label image of the same shape.
- Returns:
1.0 for identical partitions, ~0 for unrelated ones, negative when the two masks agree less than chance. 1.0 when both masks are empty (see the module docstring for why),
nanwhen the union foreground holds fewer than two pixels and there is no pair to score.
- spacr.model_compare.compare_configs(config_a: ModelConfig, config_b: ModelConfig) Dict[str, Any][source]¶
Diff two configurations into what matters and what cannot matter.
- Parameters:
config_a – the A side.
config_b – the B side.
- Returns:
{'honoured': {key: (a, b)}, 'ignored': {key: (a, b)}, 'identical': bool, 'warnings': [str]}.identicalis True when every argument that reaches the model is the same on both sides — the case where the run cannot show a difference and the report has to say so before anybody reads a number off it.
- spacr.model_compare.compare_masks(mask_a: Any, mask_b: Any, field: str = 'field', iou_threshold: float = DEFAULT_IOU_THRESHOLD, containment: float = DEFAULT_CONTAINMENT, qc_a: spacr.seg_qc.FieldQC | None = None, qc_b: spacr.seg_qc.FieldQC | None = None) SegComparison[source]¶
Compare two label images of the same field. Pure: arrays in, numbers out.
Nothing about Cellpose, models, files or the GUI reaches this function, so it works just as well on masks from any other source — which is the point: the Model Zoo reuses it unchanged.
The comparison is directional but not judgemental.
*_aand*_bdescribe the two masks; neither is the reference, so there is no precision, no recall, and no “correct” column. See the module docstring for how the ARI excludes background, why the object assignment is optimal rather than greedy, and how splits and merges are told apart from genuine differences.- Parameters:
mask_a – a 2-D label image (bool and float masks are coerced).
mask_b – a second label image of the same shape.
field – the field’s name, carried into the row.
iou_threshold – minimum IoU for two objects to be the same object.
containment – fraction of an object that must lie inside another for it to count as a fragment of it.
qc_a – an optional
spacr.seg_qc.FieldQCfor mask A, attached so the row can say which mask looks broken on its own terms.qc_b – the same for mask B.
- Returns:
- Raises:
ValueError – when the masks are not the same shape, or are not label images.
- spacr.model_compare.compare_models(images: Sequence[numpy.ndarray], model_a: Any, model_b: Any, field_names: Sequence[str] | None = None, segment_fn: Callable[[Sequence[numpy.ndarray], ModelConfig], Sequence[numpy.ndarray]] | None = None, iou_threshold: float = DEFAULT_IOU_THRESHOLD, containment: float = DEFAULT_CONTAINMENT, object_type: str = 'cell', qc: bool = True, keep_images: bool = True, progress: Callable[[str, int, int], None] | None = None) ComparisonReport[source]¶
Run two models over the same fields and compare what they produced.
Each model is loaded once and run over every field, A first, then B — one model in memory at a time, and the per-model wall clock in the report is therefore a fair (if small-sample) comparison of their cost.
The segmentation call is an argument. That is what makes this reusable: the Model Zoo can hand in a different backend, and a test can hand in a stub, so nothing about this function needs Cellpose to be exercised.
- Parameters:
images – one array per field.
model_a – a
ModelConfigor a mapping (seeModelConfig.from_mapping()).model_b – the other side.
field_names – names for the rows; defaults to
field_0000….segment_fn –
fn(images, config) -> masks; defaults tosegment_with_cellpose().iou_threshold – passed to
compare_masks().containment – passed to
compare_masks().object_type – what is being segmented, for the seg_qc scorecards.
qc – score each model’s masks with
spacr.seg_qcas well.keep_images – keep the images and masks on the report so a GUI can draw them. Pass False for a headless sweep over many fields.
progress –
fn(message, done, total), called as the run proceeds.
- Returns:
- Raises:
ValueError – when there is no field, or a model returns the wrong number of masks.
- spacr.model_compare.format_comparison(report: ComparisonReport) str[source]¶
Render the report a human reads before picking a model.
Order is deliberate: the warnings first (a run that cannot show a difference must say so before anybody reads a number off it), then the parameters that reached each model, then the per-field table, then the aggregate — and the aggregate is phrased directionally, because neither model is the truth.
- Parameters:
report – what
compare_models()returned.- Returns:
a multi-line string, ready to print.
- spacr.model_compare.iou_matrix(parts: Mapping[str, Any]) numpy.ndarray[source]¶
Object-by-object IoU from an
object_overlap()table.- Parameters:
parts – the dict
object_overlap()returned.- Returns:
an
n_a x n_bfloat array; empty when either mask has no object.
- spacr.model_compare.load_fields(source: Any, n_fields: int = DEFAULT_N_FIELDS, channel: int | None = None) Tuple[List[str], List[numpy.ndarray]][source]¶
Pull the first
n_fieldsimages out of a folder (or a list).Handles the shapes spaCR actually leaves on disk: a folder of
.tif/.pngfields, a folder of.npyarrays, and the.npzbatches the Mask module writes (data+filenames). Reading stops as soon asn_fieldsimages are in hand, so pointing this at a 1536-field plate costs three files.- Parameters:
source – a folder, or an already-loaded sequence of arrays.
n_fields – how many fields to take.
channel – index into the last axis for a multi-channel field; None keeps the array as it is.
- Returns:
(names, images).- Raises:
FileNotFoundError – when the folder does not exist.
ValueError – when it holds no readable field.
- spacr.model_compare.match_objects(mask_a: Any, mask_b: Any, iou_threshold: float = DEFAULT_IOU_THRESHOLD) Dict[str, Any][source]¶
Pair the objects of two masks by optimal assignment.
Bipartite, not greedy. The IoU matrix is thresholded first (everything below
iou_thresholdbecomes 0) andscipy.optimize.linear_sum_assignment()then maximises the total IoU over what is left, which is the same procedurecellpose.metrics._true_positiveuses. Taking each object’s best partner double-assigns; taking pairs greedily in descending IoU can strand a pair that the optimal assignment keeps. Both are exercised in the tests.- Parameters:
mask_a – a 2-D label image.
mask_b – a second label image of the same shape.
iou_threshold – minimum IoU for a pair to count as the same object.
- Returns:
{'matches': [(label_a, label_b, iou), ...], 'iou': matrix, 'parts': the overlap table, 'unmatched_a': [labels], 'unmatched_b': [labels]}, matches sorted by descending IoU.
- spacr.model_compare.object_overlap(mask_a: Any, mask_b: Any) Dict[str, Any][source]¶
Contingency between the objects of two label images.
Both masks are relabelled to
0..nfirst, so arbitrary label values (and gaps in them) are handled, and the pixel-pair counting below can index straight into the table.- Parameters:
mask_a – a 2-D label image.
mask_b – a second label image of the same shape.
- Returns:
{'labels_a', 'labels_b', 'areas_a', 'areas_b', 'overlap', 'n_pixels', 'union_foreground'}.overlapis then_a x n_bobject-by-object intersection count with background already dropped;labels_amaps rowiback to the original label.- Raises:
ValueError – when the two masks are not the same shape.
- spacr.model_compare.segment_with_cellpose(images: Sequence[numpy.ndarray], config: ModelConfig) List[numpy.ndarray][source]¶
Segment
imageswith one Cellpose model. The default backend.Only
ModelConfig.eval_kwargs()is forwarded, so an argument Cellpose 4 ignores never reacheseval— it is reported by the caller instead of being passed on to be silently dropped, which is the difference between a comparison that explains itself and one that says “no difference”.Model and device dependencies are loaded on demand for this call. Cellpose versions without a separate
invertargument receive that setting in the normalization dictionary. Inversion requires normalization on those versions; other normalization options are retained.- Parameters:
images – 2-D or 3-D arrays, one per field.
config – the model to run.
- Returns:
one integer label image per input image.