spacr.scorecard

Score a segmentation against GROUND TRUTH, which nothing in spaCR did.

Every published model carries a table comparing the finetuned model with the vanilla one on a held-out set, with all the common metrics for that model type. Auditing what existed found the gap this module fills:

  • spacr.model_compare compares two models to EACH OTHER and says so in its own docstring – “neither model is ground truth”. It answers “do these two disagree”, which is a different and also useful question.

  • spacr.seg_qc scores a field with NO labels at all: it flags splits, merges and implausible diameters from the mask alone.

Both are useful and neither is Dice. Nothing measured a mask against a hand-drawn truth, so no published number could say a model was BETTER – only that it was different.

WHAT AN OBJECT METRIC HAS TO DECIDE FIRST

Every number here rests on one choice: when is a predicted object THE SAME object as a labelled one? The literature’s answer is an IoU threshold, and the threshold is not a detail – a model that finds every cell and outlines them loosely scores well at 0.5 and badly at 0.9, and a model that finds half of them perfectly does the opposite. That is why average_precision() reports the whole sweep from 0.5 to 0.9 rather than one number, and why every precision/recall figure states the threshold it was matched at.

MATCHING IS ONE-TO-ONE AND OPTIMAL, not greedy. A greedy pass down the IoU matrix is cheaper and gives a different answer when one prediction overlaps two truths: it takes the first pair it sees rather than the assignment that maximises total overlap, so the score depends on label ORDER, which is an implementation detail of whoever wrote the mask. match_objects() uses scipy.optimize.linear_sum_assignment, so relabelling either mask cannot change the result.

WHAT IS DELIBERATELY NOT HERE

No torch, and no import of anything that pulls it. The model zoo imports without torch on purpose and a test asserts it, so reading a scorecard must not be the thing that drags in a GPU stack – the reader is browsing a list. Everything here is numpy plus two scipy/skimage helpers.

Nothing writes files, uploads, or reads a catalogue. This module turns two label arrays into numbers; publishing them is 370’s other half.

Classes

HoldoutField

One labelled field: the image a model reads, and the truth it is

HoldoutScore

One model's result on one named, versioned held-out set.

HoldoutSet

A named, versioned set of labelled fields, as declared by a manifest.

Match

One IoU threshold's worth of matching, and what it implies.

Functions

average_precision(→ Dict[str, float])

AP at each IoU in the sweep, and the mean across it.

boundary_f1(→ Dict[str, float])

Boundary precision, recall and F1 at each pixel tolerance.

compare_against_baseline(→ Dict[str, Dict[str, float]])

Score two models on the same truth and report the difference.

compare_classifier_against_baseline(→ Dict[str, ...)

Two classifiers on the same labels, and the difference.

counts_and_areas(→ Dict[str, float])

Object-count error (signed) and the area error a measurement inherits.

dice(→ Dict[str, float])

Per-object Dice, averaged over matched pairs, plus the pixel-wise one.

headline(→ List[str])

The few lines a tooltip should lead with, and where the rest is.

iou_matrix(→ Tuple[numpy.ndarray, numpy.ndarray, ...)

(iou, truth ids, pred ids). IoU of every overlapping pair.

load_holdout(→ HoldoutSet)

Read a hold-out manifest.

match_objects(→ Match)

Match predicted objects to labelled ones, one to one and optimally.

read_scorecard_csv(→ Dict[str, Any])

Parse a published scorecard CSV back into an entry's metrics.

score_classifier(→ Dict[str, float])

Every classifier metric the published table reports, at a STATED threshold.

score_holdout(→ HoldoutScore)

Score a model over a whole held-out set.

score_model_on_holdout(→ HoldoutScore)

Run predict over a hold-out set and score it against the truth.

score_segmentation(→ Dict[str, float])

Every segmentation metric the published table reports, on one field.

scorecard_csv(→ str)

The rows as CSV text. Written by the caller, wherever it belongs.

scorecard_figure(metrics, path, *[, title, dpi])

Draw finetuned against stock as paired bars, and write it to path.

scorecard_is_present(→ bool)

Whether metrics actually carries a scorecard.

scorecard_rows(→ List[Dict[str, object]])

The published table, one row per metric. THE SOURCE THE REST DERIVE FROM.

splits_and_merges(→ Dict[str, int])

How many labelled objects were split, and how many were merged.

verify_holdout(→ List[str])

Check every truth mask is present and matches its digest.

Module Contents

class spacr.scorecard.HoldoutField[source]

One labelled field: the image a model reads, and the truth it is scored against.

Parameters:
  • image – path to the field, relative to the manifest.

  • truth – path to the label mask, relative to the manifest.

  • sha256 – digest of the TRUTH mask, or "".

class spacr.scorecard.HoldoutScore[source]

One model’s result on one named, versioned held-out set.

Parameters:
  • name – the set’s name, e.g. toxo_pv_holdout.

  • version – the set’s version. Two numbers are comparable only when this matches, and it is carried into every published row for that reason rather than being assumed from context.

  • metrics – the pooled scorecard.

  • per_field – one scorecard per field, in the order they were given.

property n_fields: int[source]

How many labelled fields this score was computed over.

Returns:

the field count.

class spacr.scorecard.HoldoutSet[source]

A named, versioned set of labelled fields, as declared by a manifest.

Parameters:
  • name – stable name, e.g. toxo_pv_holdout.

  • version – the version two numbers must share to be comparable.

  • fields – the labelled fields, in manifest order.

  • root – directory the relative paths resolve against.

path_to(relative: str) → pathlib.Path[source]

Resolve a manifest-relative path against this set’s root.

MANIFEST PATHS ARE RELATIVE ON PURPOSE, so a held-out set can be moved or shared without rewriting it – which is what makes the versioned name mean the same thing on two machines.

Parameters:

relative – the path as the manifest spells it.

Returns:

the absolute path.

class spacr.scorecard.Match[source]

One IoU threshold’s worth of matching, and what it implies.

Parameters:
  • threshold – the IoU at which a pair counted as the same object.

  • pairs – (truth index, pred index) for each matched pair, into the id arrays iou_matrix() returns.

  • ious – the IoU of each matched pair, in the same order.

  • n_truth – labelled objects present.

  • n_pred – predicted objects present.

property average_precision: float[source]

TP / (TP + FP + FN) – the segmentation-benchmark convention.

NOT the area under a precision-recall curve, despite the name. The segmentation literature uses this quantity and calls it AP, and reporting the other thing under the same label is how two papers’ numbers stop being comparable. Named explicitly here for that reason.

property f1: float[source]

Harmonic mean of precision and recall, 0.0 when both are 0.

property false_negatives: int[source]

Truth objects the model did not find.

Returns:

the count.

property false_positives: int[source]

Predicted objects with no truth object to match.

Returns:

the count.

property precision: float[source]

Of what the model predicted, how much was real.

Returns:

the ratio, or 0 when the model predicted nothing.

property recall: float[source]

Of what was really there, how much the model found.

Returns:

the ratio, or 0 when there was nothing to find.

property true_positives: int[source]

Predicted objects that matched a truth object at this threshold.

Returns:

the count.

spacr.scorecard.average_precision(truth: numpy.ndarray, pred: numpy.ndarray, thresholds: Sequence[float] = DEFAULT_IOU_THRESHOLDS) → Dict[str, float][source]

AP at each IoU in the sweep, and the mean across it.

Parameters:
  • truth – the ground-truth label image, 0 for background and one integer id per object.

  • pred – the predicted label image, the same shape as truth.

spacr.scorecard.boundary_f1(truth: numpy.ndarray, pred: numpy.ndarray, tolerances: Sequence[int] = DEFAULT_BOUNDARY_TOLERANCES) → Dict[str, float][source]

Boundary precision, recall and F1 at each pixel tolerance.

A predicted boundary pixel counts as correct when a true boundary pixel lies within tolerance pixels of it, and vice versa. The tolerance is not slack for the model’s benefit: two people labelling the same cell disagree by a pixel or two, so a zero-tolerance boundary score measures the annotator as much as the model.

Parameters:
  • truth – the ground-truth label image, 0 for background and one integer id per object.

  • pred – the predicted label image, the same shape as truth.

spacr.scorecard.compare_against_baseline(truth: numpy.ndarray, finetuned: numpy.ndarray, vanilla: numpy.ndarray, **kwargs) → Dict[str, Dict[str, float]][source]

Score two models on the same truth and report the difference.

Parameters:
  • truth – the ground-truth label image, 0 for background and one integer id per object.

  • finetuned – the finetuned model’s label image for the same field.

  • vanilla – the vanilla model’s label image for the same field.

Returns:

{"finetuned": ..., "vanilla": ..., "delta": ...}.

THE DELTA IS THE ANSWER THE TABLE EXISTS FOR. The table compares the finetuned model against the vanilla one, and a reader given two columns of eleven numbers will do this subtraction by eye and get it wrong somewhere. Reporting it is not a convenience.

Both models are scored on the SAME array with the SAME settings, which is the only thing that makes the subtraction meaningful – and is why this takes three masks rather than two scorecards.

spacr.scorecard.compare_classifier_against_baseline(labels: Sequence[int], finetuned: Sequence[float], vanilla: Sequence[float], **kwargs) → Dict[str, Dict[str, float]][source]

Two classifiers on the same labels, and the difference.

The same shape as compare_against_baseline(), and for the same reason: the delta is what the published table is for. threshold and the counts are dropped from it – a difference in n is not a result, it is a sign the two were scored on different data, which this signature makes impossible.

Parameters:
  • labels – ground truth per object, 0 or 1.

  • finetuned – the finetuned classifier’s predicted probability of the positive class, one per label.

  • vanilla – the vanilla classifier’s predicted probability of the positive class, one per label.

spacr.scorecard.counts_and_areas(truth: numpy.ndarray, pred: numpy.ndarray, threshold: float = 0.5) → Dict[str, float][source]

Object-count error (signed) and the area error a measurement inherits.

SIGNED, because the direction is the diagnosis: a model that finds too many objects is over-segmenting and one that finds too few is merging or missing, and an absolute count error says neither.

Parameters:
  • truth – the ground-truth label image, 0 for background and one integer id per object.

  • pred – the predicted label image, the same shape as truth.

spacr.scorecard.dice(truth: numpy.ndarray, pred: numpy.ndarray, threshold: float = 0.5) → Dict[str, float][source]

Per-object Dice, averaged over matched pairs, plus the pixel-wise one.

TWO NUMBERS BECAUSE THEY ANSWER DIFFERENT QUESTIONS and are routinely confused. The per-object mean says how well a typical object is outlined; the pixel-wise figure ignores objects entirely and is dominated by the largest ones, so a model that misses ten small cells and nails one big one scores well on it and badly on the other.

Parameters:
  • truth – the ground-truth label image, 0 for background and one integer id per object.

  • pred – the predicted label image, the same shape as truth.

spacr.scorecard.headline(metrics: Mapping[str, object], *, baseline: Mapping[str, object] | None = None, limit: int = 3) → List[str][source]

The few lines a tooltip should lead with, and where the rest is.

Parameters:
  • metrics – a scorecard, or any mapping; unknown keys are ignored.

  • baseline – the vanilla model’s scorecard, to show the difference.

  • limit – how many metric lines to return before the pointer.

Returns:

lines, or [] when the mapping holds no scorecard at all – an entry with free-form metrics is left exactly as it was, because this must not turn somebody’s two-line note into a truncated table.

spacr.scorecard.iou_matrix(truth: numpy.ndarray, pred: numpy.ndarray) → Tuple[numpy.ndarray, numpy.ndarray, numpy.ndarray][source]

(iou, truth ids, pred ids). IoU of every overlapping pair.

Parameters:
  • truth – the ground-truth label image, 0 for background and one integer id per object.

  • pred – the predicted label image, the same shape as truth.

spacr.scorecard.load_holdout(manifest_path) → HoldoutSet[source]

Read a hold-out manifest.

Parameters:

manifest_path – path to the hold-out manifest JSON, which must declare name, version and fields.

Raises:

ValueError – when the manifest lacks a name, a version or any field. All three are refusals rather than defaults, and the version most of all: a set that does not say which version it is cannot be compared against anything, and a default would let one be published as though it could.

spacr.scorecard.match_objects(truth: numpy.ndarray, pred: numpy.ndarray, threshold: float = 0.5) → Match[source]

Match predicted objects to labelled ones, one to one and optimally.

Parameters:
  • truth – the ground-truth label image, 0 for background and one integer id per object.

  • pred – the predicted label image, the same shape as truth.

  • threshold – minimum IoU for a pair to count as the same object.

Uses linear_sum_assignment rather than a greedy pass so the result cannot depend on label order – see this module’s docstring.

spacr.scorecard.read_scorecard_csv(text: str) → Dict[str, Any][source]

Parse a published scorecard CSV back into an entry’s metrics.

THE CSV IS THE SOURCE AND THE OTHER SURFACES RENDER IT: the Hugging Face artifact, the tooltip, the API page and the Zoo screen must not be able to disagree, and they cannot if only one of them holds numbers. This is the reader that makes the other three derived.

DEPENDENCY-FREE ON PURPOSE. The Model Zoo imports without torch and a test asserts it, so browsing models and reading a scorecard has to work on a machine with neither torch nor cellpose. Only RE-RUNNING an evaluation needs them. That rules out pandas here too – csv is in the standard library.

Parameters:

text – the CSV as written by scorecard_csv().

Returns:

{metric: {"finetuned": ..., "vanilla": ..., "delta": ...}} plus holdout and holdout_version when the rows carry them. Empty when the text has no rows.

spacr.scorecard.score_classifier(labels: Sequence[int], scores: Sequence[float], *, threshold: float = 0.5, calibration_bins: int = 10) → Dict[str, float][source]

Every classifier metric the published table reports, at a STATED threshold.

Parameters:
  • labels – ground truth, 0 or 1.

  • scores – predicted probability of the positive class.

  • threshold – where a score becomes a positive call. Reported back in the result, because “precision 0.94” without it is not a claim anybody can check or reproduce.

AUROC, AUPRC, Brier and ECE are threshold-FREE and are the numbers that survive a reader disagreeing with the cutoff; everything else moves when the threshold does. Both kinds are here, and the result says which is which by carrying threshold beside them.

spacr.scorecard.score_holdout(pairs: Sequence[Tuple[numpy.ndarray, numpy.ndarray]], *, name: str, version: str, **kwargs) → HoldoutScore[source]

Score a model over a whole held-out set.

Parameters:
  • pairs – (truth, prediction) per field.

  • name – the held-out set’s name, carried into the result.

  • version – the held-out set’s version, carried into the result.

POOLED FROM THE COUNTS, NOT AVERAGED FROM THE RATIOS. A field with three objects and a field with three hundred are not equal evidence, and a mean of per-field precisions treats them as though they were – so a model that fails on one sparse field is punished as hard as one that fails on a confluent one. Precision, recall and F1 are recomputed from the summed true and false positives; only the genuinely per-object means (IoU, Dice, area error) are averaged, and those are weighted by the objects behind them.

spacr.scorecard.score_model_on_holdout(holdout: HoldoutSet, predict, *, read_mask=None, **kwargs) → HoldoutScore[source]

Run predict over a hold-out set and score it against the truth.

Parameters:
  • holdout – the hold-out set whose fields’ images are predicted and whose truth masks are read.

  • predict – image path -> label array. Injected rather than imported: this module must keep importing without torch, and a segmentation model is the one thing that cannot.

  • read_mask – path -> label array. Defaults to tifffile, which the package already depends on.

THE SET’S NAME AND VERSION TRAVEL WITH THE SCORE, so a published number can never be read without knowing what it was measured on.

spacr.scorecard.score_segmentation(truth: numpy.ndarray, pred: numpy.ndarray, *, threshold: float = 0.5, iou_thresholds: Sequence[float] = DEFAULT_IOU_THRESHOLDS, boundary_tolerances: Sequence[int] = DEFAULT_BOUNDARY_TOLERANCES) → Dict[str, float][source]

Every segmentation metric the published table reports, on one field.

Parameters:
  • truth – the ground-truth label image, 0 for background and one integer id per object.

  • pred – the predicted label image, the same shape as truth.

  • threshold – the IoU at which precision, recall, F1, Dice and the area error are matched. Reported in the result as match_iou so a published number can never be read without it.

EVERY NUMBER CARRIES ITS N. n_truth and n_pred are in the result and are not optional: a Dice of 0.91 on eleven objects is not a result, and a table that omits the count invites exactly that reading.

spacr.scorecard.scorecard_csv(rows: Sequence[Dict[str, object]]) → str[source]

The rows as CSV text. Written by the caller, wherever it belongs.

Parameters:

rows – the table rows, as from scorecard_rows(); the first row’s keys are the CSV columns, and no rows gives "".

spacr.scorecard.scorecard_figure(metrics: Mapping[str, object], path, *, title: str = '', dpi: int | None = None)[source]

Draw finetuned against stock as paired bars, and write it to path.

The scorecard in graph form, for Hugging Face and for the API page, beside the CSV and the table. It reads the SAME parsed metrics the tooltip and the zoo screen read, so the picture cannot disagree with the numbers printed next to it.

MATPLOTLIB IS IMPORTED INSIDE, deliberately. This module’s contract is that the Model Zoo can import it with neither torch nor cellpose present – a test asserts it – and a drawing dependency at module scope would break that for every caller who only wanted to READ a scorecard.

ONLY THE METRICS THAT ANSWER THE QUESTION. A scorecard holds forty numbers; a chart of forty bars is a wall, not an answer. The ones in CHART_METRICS that the scorecard actually carries are drawn, in that order, and anything absent is skipped rather than drawn as zero – a missing metric and a metric that scored nothing look identical at a glance and mean opposite things.

Parameters:
  • metrics – as read_scorecard_csv() returns.

  • path – where to write the PNG.

  • title – heading; the model’s name is the useful thing to pass.

Returns:

the path written.

Raises:

ValueError – when none of the chart metrics are present, since an empty chart published beside a model would imply it scored zero.

spacr.scorecard.scorecard_is_present(metrics: Mapping[str, Any]) → bool[source]

Whether metrics actually carries a scorecard.

A MISSING SCORECARD MUST SAY SO RATHER THAN SHOW BLANKS, which 370 asks for by name and which ModelEntry.provenance_known already does for training provenance. A model with no numbers is not a model that scored zero, and a table of empty cells reads as the second.

Parameters:

metrics – an entry’s metrics mapping.

spacr.scorecard.scorecard_rows(finetuned: HoldoutScore, vanilla: HoldoutScore) → List[Dict[str, object]][source]

The published table, one row per metric. THE SOURCE THE REST DERIVE FROM.

The table has four renderings – the CSV on Hugging Face, the tooltip, the API section and the zoo screen – and if the tooltip and the API page can disagree, they eventually will. This is the one place a number is computed; everything else formats these rows.

Parameters:
  • finetuned – the finetuned model’s score.

  • vanilla – the vanilla model’s score, on the same set name and version.

Raises:

ValueError – when the two were scored on different sets. That is not a defensive check for its own sake: a table headed “finetuned against vanilla” whose two columns came from different data is exactly the mistake that cannot be seen by reading it.

spacr.scorecard.splits_and_merges(truth: numpy.ndarray, pred: numpy.ndarray, minimum_overlap: float = 0.1) → Dict[str, int][source]

How many labelled objects were split, and how many were merged.

Parameters:
  • truth – the ground-truth label image, 0 for background and one integer id per object.

  • pred – the predicted label image, the same shape as truth.

  • minimum_overlap – fraction of the truth object a prediction must cover to count as overlapping it, so a one-pixel graze is not a split.

THE TWO FAILURES ARE NOT SYMMETRIC IN WHAT THEY COST. A split inflates the object count and halves the areas; a merge deletes an object and doubles one. Both are invisible to Dice at the field level, which is why they are counted separately rather than folded into it.

spacr.scorecard.verify_holdout(holdout: HoldoutSet) → List[str][source]

Check every truth mask is present and matches its digest.

Parameters:

holdout – the hold-out set whose truth masks are checked on disk.

Returns:

one line per problem; empty when the set is intact.

A DIGEST IS OPTIONAL AND ITS ABSENCE IS REPORTED. A set published without them can still be scored, and nobody can then tell whether two people scored the same masks – which is the entire point of naming and versioning it. So “no digest” is a finding, not a pass.

Nested helpers

read_scorecard_csv.number(value)

A cell as a number, or the raw string when it is not one.

A WHOLE NUMBER COMES BACK AS AN int. The CSV cannot distinguish a count from a measurement, and reading everything as float renders “on 12517.0 objects” in a tooltip – a count with a decimal point reads as a rounding, which it is not.

spacr/scorecard.py:1013

score_holdout.weighted(key: str, weight_key: str = 'n_matched') → float

Average key across fields, weighted by how much each holds.

A plain mean would let a field with four objects count as much as one with four hundred, so a sparse corner of the plate could move the published number more than the plate does.

spacr/scorecard.py:673

score_model_on_holdout.read_mask(path)

Read a label image off disk.

The default, used when the caller passes none. Imported here rather than at module scope so scoring an already-loaded pair of arrays needs no tifffile.

spacr/scorecard.py:908