spacr.classifier_evaluation

Workflow inputs and outputs

Classifier Evaluation

Inspect held-out predictions, calibration and split/leakage evidence for the matching classifier.

Open: Classify → Classifier Evaluation.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Classifier evaluation bundle — Held-out predictions, labels, split metadata and calibration/leakage metrics for a saved classifier run.

Outputs

  • Quality-control results — Stored project checks and QC reports; a missing check is not a passing result.

  • Figures and table exports — The output location chosen by the tool; exports describe the selected data and filters.

Before this module

  • Classify: Use the matching held-out predictions and split metadata.

API reference.

Module tutorial.

Classifier evaluation, calibration, and split-leakage diagnostics.

The module is model-agnostic: it consumes labels, probabilities, fold ids, and sample paths. Deep-learning CV and future classical-ML pipelines can therefore write the same evaluation bundle and the Qt workbench can display one stable artifact format.

Attributes

EVALUATION_FILES

Stable file names produced by write_evaluation_bundle().

Exceptions

LeakageError

Raised when related samples cross a protected split boundary.

Classes

FoldLeakageAudit

Whole-CV proof that each related sample family belongs to one fold.

LeakageReport

Overlap counts and examples for one train/validation boundary.

SplitReport

Provenance and realised sizes for one train/test split.

Functions

audit_cv_folds(→ FoldLeakageAudit)

Verify partition coverage and identity isolation across every CV fold.

audit_dataset_splits(→ LeakageReport)

Audit the permanent train/ versus test/ dataset boundary.

audit_split_leakage(→ LeakageReport)

Detect related images crossing a train/validation boundary.

augmentation_family(→ str)

Return the original-sample family after removing augmentation suffixes.

calibration_table(→ pandas.DataFrame)

Return per-class reliability bins for calibration plots.

cross_calibrate_probabilities(→ Tuple[numpy.ndarray, ...)

Calibrate each held-out fold using every other out-of-fold prediction.

dataset_split_paths(→ List[str])

Return sorted image paths under root/<split>/<class>/.

evaluate_predictions(→ Dict[str, Any])

Build overall, confusion, calibration, and per-plate evaluation tables.

expected_calibration_error(→ float)

Return top-label expected calibration error (ECE).

find_evaluation_bundles(→ List[pathlib.Path])

Return evaluation manifests below root, newest first.

fit_temperature(→ float)

Fit one scalar temperature by minimizing multiclass log loss.

grouped_split(→ Tuple[numpy.ndarray, numpy.ndarray, ...)

Return a stratified holdout while keeping named groups intact.

load_evaluation_bundle(→ Dict[str, Any])

Load one evaluation bundle for the Qt workbench.

nested_group_folds(→ List[Dict[str, Any]])

Build nested stratified/grouped outer and inner index partitions.

normalize_probabilities(→ numpy.ndarray)

Return a finite, row-normalized (n_samples, n_classes) matrix.

normalize_split_level(→ str)

Return a canonical split level; legacy none means cell.

sample_identity(→ Dict[str, str])

Parse plate/well/field/object identities from a crop filename.

split_columns_for(→ Tuple[str, List[str]])

Resolve a split level to complete metadata columns or refuse it.

split_group_values(→ Tuple[str, numpy.ndarray])

Build one strict group id per row from metadata, prcfo, or paths.

write_evaluation_bundle(→ pathlib.Path)

Atomically write a complete evaluation bundle and diagnostic figures.

write_leakage_audit(→ pathlib.Path)

Atomically write a leakage report/audit as JSON and return its path.

Module Contents

exception spacr.classifier_evaluation.LeakageError[source]

Bases: ValueError

Raised when related samples cross a protected split boundary.

Initialize self. See help(type(self)) for accurate signature.

class spacr.classifier_evaluation.FoldLeakageAudit[source]

Whole-CV proof that each related sample family belongs to one fold.

Parameters:
  • group_by – canonical identity level required to remain within one validation fold.

  • n_samples – number of source paths whose fold membership was audited.

  • n_folds – number of train/validation fold pairs inspected.

  • validation_membership_missing – up to twenty sample indexes that were never held out for validation.

  • validation_membership_duplicate – up to twenty sample indexes held out in more than one fold.

  • overlap_counts – identities assigned to multiple validation folds, counted at each exact, content, family, object, and acquisition level.

  • examples – up to ten conflicting identities per level, annotated with the folds that contain them.

  • critical_levels – completeness, overlap, identity, or label failures that make passed false.

  • warnings – non-fatal caveats and explanations accompanying failures.

  • unverifiable_counts – samples whose requested identity or optional byte-content hash could not be checked, counted by level.

  • hash_errors – up to twenty file-specific content-hashing failures.

  • split_name – stable label for this whole-CV audit record.

to_dict() → Dict[str, Any][source]

Return a JSON-serializable audit record.

property passed: bool[source]

Return True only when the fold partition is complete and isolated.

class spacr.classifier_evaluation.LeakageReport[source]

Overlap counts and examples for one train/validation boundary.

Parameters:
  • group_by – protected split level (cell, field, well, or plate); the legacy none spelling is normalized to cell.

  • train_samples – number of training paths.

  • validation_samples – number of validation paths.

  • overlap_counts – overlap count at each identity level.

  • examples – up to ten shared identities per level.

  • split_name – caller-supplied label for the audited boundary.

  • critical_levels – levels that invalidate the requested split.

  • warnings – non-fatal caveats.

  • unverifiable_counts – samples lacking a requested identity or content hash, counted by the level that could not be verified.

  • hash_errors – up to twenty file-specific failures from optional byte-content hashing.

to_dict() → Dict[str, Any][source]

Return a JSON-serializable report.

property passed: bool[source]

Return True when no protected identity crosses the boundary.

class spacr.classifier_evaluation.SplitReport[source]

Provenance and realised sizes for one train/test split.

Parameters:
  • group_by – canonical isolation unit used for the split, from objects through fields, wells, and plates.

  • requested_fraction – held-out object fraction requested by the caller.

  • cell_fraction – realised share of objects assigned to the test side.

  • group_fraction – realised share of distinct groups assigned to test.

  • train_cells – number of object rows used for fitting.

  • test_cells – number of object rows held out for evaluation.

  • train_groups – number of distinct groups used for fitting.

  • test_groups – number of distinct groups held out for evaluation.

  • total_groups – distinct groups across both sides of the split.

  • rule – human-readable algorithm and isolation guarantee that produced the realised split.

summary() → str[source]

Describe both whole-group and object-level holdout costs.

to_dict() → Dict[str, Any][source]

Return JSON-safe split provenance for model cards/settings.

spacr.classifier_evaluation.audit_cv_folds(paths: Sequence[Any], folds: Sequence[Tuple[Sequence[int], Sequence[int]]], *, labels: Sequence[Any] | None = None, group_by: str = 'well', hash_content: bool = False, require_identity: bool = True, raise_on_leakage: bool = False) → FoldLeakageAudit[source]

Verify partition coverage and identity isolation across every CV fold.

This checks the fold assignment as one object, rather than trusting a sample of pairwise boundaries. Each index must be validation exactly once; exact paths, byte-identical content, source/augmentation families and the requested plate/well/field group must map to one held-out fold only.

Parameters:
  • paths – one source path per sample, indexed by the fold indices.

  • folds – (train_indices, validation_indices) per fold. Indices outside range(len(paths)) are reported rather than ignored.

  • labels – optional class per path, one value per path. Only used to describe the folds; it does not affect leakage detection.

  • group_by – identity level that may not cross a fold boundary – none, field, well or plate. A well-grouped split permits the same plate on both sides but never the same well.

  • hash_content – also compare file CONTENT, so a byte-identical copy under a different name is caught. Costs one read per file.

  • require_identity – treat UNVERIFIABLE as critical. A filename that does not encode the requested identity, or a file that cannot be hashed, is otherwise only a warning – so leaving this False means a clean report can still hide leakage nobody could check for.

  • raise_on_leakage – raise LeakageError instead of returning a report whose passed is False.

Returns:

a FoldLeakageAudit carrying the per-level overlap counts, examples, and which levels were critical.

Raises:

ValueError – for an unsupported group_by, or labels whose length does not match paths.

spacr.classifier_evaluation.audit_dataset_splits(root: Any, *, group_by: str = 'well', hash_content: bool = True, require_identity: bool = True, raise_on_leakage: bool = False) → LeakageReport[source]

Audit the permanent train/ versus test/ dataset boundary.

Parameters:
  • root – dataset directory holding train/ and test/. Both are searched recursively, and only .png, .jpg, .jpeg, .tif, .tiff, .bmp and .npy files are collected, so a folder of any other format audits as empty.

  • group_by – identity level that may not appear on both sides – none, field, well or plate. The default well still permits the same plate in train and test. It is validated only after the images are found, so a bad value over an empty tree reports the missing images instead.

  • hash_content – defaults to True here, unlike audit_split_leakage(), so a byte-identical copy saved under a different name fails the audit – at the cost of one read per file.

  • require_identity – defaults to True here: filenames that do not encode the group_by level, and files that cannot be hashed, become critical instead of a warning, so an unverifiable split cannot report as clean.

  • raise_on_leakage – raise LeakageError instead of returning a report whose passed is False.

Raises:

FileNotFoundError – when either side collects no image – a missing or unreadable folder is never treated as a passing split.

Returns:

a LeakageReport whose split_name is always train_vs_test; the test/ side is counted as validation_samples.

spacr.classifier_evaluation.audit_split_leakage(train_paths: Sequence[Any], validation_paths: Sequence[Any], *, group_by: str = 'well', raise_on_leakage: bool = False, split_name: str = '', hash_content: bool = False, require_identity: bool = False) → LeakageReport[source]

Detect related images crossing a train/validation boundary.

Exact/object/augmentation-family overlap is always critical. The requested group_by level is also critical: a well-grouped split permits the same plate on both sides but never the same well.

Parameters:
  • train_paths – source paths used to fit the model.

  • validation_paths – paths used only for evaluation.

  • group_by – cell, field, well, or plate. Legacy none aliases cell.

  • raise_on_leakage – raise LeakageError on a critical overlap.

  • split_name – optional fold/split label stored in the report.

  • hash_content – also compare file CONTENT, so a byte-identical copy under a different name is caught. Costs one read per file.

  • require_identity – treat UNVERIFIABLE as critical. A filename that does not encode the requested identity, or a file that cannot be hashed, is otherwise only a warning – so leaving this False means a clean report can still hide leakage nobody could check for.

Returns:

LeakageReport.

spacr.classifier_evaluation.augmentation_family(path: Any) → str[source]

Return the original-sample family after removing augmentation suffixes.

In-memory spaCR augmentations retain the exact source filename, while older exported datasets use suffixes such as _aug3, _rot90 or _flip_h. Both forms collapse to one family.

Parameters:

path – crop path or file name; only its basename without the final extension is used, with trailing augmentation suffixes stripped repeatedly.

spacr.classifier_evaluation.calibration_table(y_true: Sequence[int], probabilities: Any, *, classes: Sequence[str] | None = None, n_bins: int = 10) → pandas.DataFrame[source]

Return per-class reliability bins for calibration plots.

Parameters:
  • y_true – class INDEX per sample, compared column by column. A label outside the probability columns is REFUSED: it matches no class, so it used to read observed_frequency 0.0 everywhere and render as a catastrophically miscalibrated curve rather than an error.

  • probabilities – predicted probabilities; a 1-D positive-class vector is expanded to two columns and every row is renormalized.

  • classes – display names in column order. An EMPTY sequence falls back to class_0 ... class_n; a non-empty one whose length disagrees with the columns is refused.

  • n_bins – equal-width confidence bins over [0, 1], truncated to an int and floored at 2, so 0, 1 and any negative value all give two bins. The top bin is closed on the right, so confidence 1.0 lands in it rather than falling out of the table.

Raises:

ValueError – when y_true and probabilities differ in length, classes has the wrong length, or a class index falls outside the probability columns.

Returns:

one row per class and NON-EMPTY bin, so the frame is shorter than n_classes * n_bins rows. An empty result still carries CALIBRATION_COLUMNS, so a column can be indexed on it.

spacr.classifier_evaluation.cross_calibrate_probabilities(y_true: Sequence[int], probabilities: Any, fold_ids: Sequence[Any], *, method: str = 'temperature', warnings_out: List[str] | None = None) → Tuple[numpy.ndarray, Dict[str, float]][source]

Calibrate each held-out fold using every other out-of-fold prediction.

This cross-fitting prevents a sample from fitting the calibrator that is evaluated on that same sample.

Parameters:
  • y_true – class INDEX per sample, read only to fit the temperatures. A value outside the probability columns is REFUSED – it used to make every fold’s fit fail, fall back to temperature 1.0, and return uncalibrated probabilities while reporting that calibration ran.

  • probabilities – predicted probabilities; a 1-D positive-class vector is expanded to two columns and every row is renormalized, so the result is always a new normalized matrix – never the object passed in, even when no calibration is applied.

  • fold_ids – held-out block per sample; any hashable value works, and each distinct value is calibrated from all the others, so at least two distinct values are required. The returned map is keyed by str(fold_id).

  • method – temperature, or none/off/false (also None and the empty string) to return the normalized probabilities unchanged with an empty temperature map. Case and surrounding whitespace are ignored; anything else is refused rather than skipped.

  • warnings_out – list appended in place when a fold cannot be calibrated from the others. The message is printed either way, so leaving this None only discards the machine-readable copy.

Raises:

ValueError – on mismatched lengths, an unrecognized method, or fewer than two distinct fold_ids when calibrating.

Returns:

(calibrated_probabilities, temperature_by_held_out_fold).

spacr.classifier_evaluation.dataset_split_paths(root: Any, split: str) → List[str][source]

Return sorted image paths under root/<split>/<class>/.

Parameters:
  • root – dataset folder that contains the split folders; ~ is expanded.

  • split – name of the split subfolder, such as train or test. A missing folder returns an empty list; files with an image or .npy suffix are collected recursively below it.

spacr.classifier_evaluation.evaluate_predictions(y_true: Sequence[int], probabilities: Any, sample_paths: Sequence[Any], *, classes: Sequence[str] | None = None, fold_ids: Sequence[Any] | None = None, calibration_method: str = 'none', calibration_bins: int = 10) → Dict[str, Any][source]

Build overall, confusion, calibration, and per-plate evaluation tables.

Parameters:
  • y_true – true class INDEX per sample, in range(n_classes). A value outside that range is refused rather than clipped.

  • probabilities – (n_samples, n_classes) predicted probabilities. Its column count defines n_classes.

  • sample_paths – one source path per sample, used to derive the per-plate tables. Must be the same length as y_true.

  • classes – display names for the columns, in column order. Defaults to class_0 ... class_n; a length that disagrees with the probability columns is refused.

  • fold_ids – which held-out fold each sample came from. REQUIRED for temperature calibration and unused otherwise – the temperature is fitted per held-out fold so no sample is calibrated on itself, which needs at least two distinct folds.

  • calibration_method – 'none' (default) or 'temperature'. Anything else is refused rather than silently ignored.

  • calibration_bins – bin count for the expected-calibration-error estimate. More bins resolve the reliability curve better and make each bin noisier.

Raises:

ValueError – on mismatched lengths, a class index outside the probability columns, an unknown calibration_method, or temperature calibration with fewer than two distinct folds.

Returns:

the evaluation bundle, a dict with six keys:

  • summary — scalar metrics (n, accuracy, balanced accuracy, macro/weighted F1, macro precision/recall, log loss, multiclass Brier, expected calibration error, mean confidence) plus classes, n_classes, calibration_method, raw_expected_calibration_error, temperatures_by_held_out_fold, calibration_warnings and probability_column_names;

  • predictions — one row per sample with fold, the sample_identity() columns, true_label / true_class, predicted_label / predicted_class, correct, confidence (the calibrated probability of the chosen class) and a raw_prob_<name> / prob_<name> pair per class, where <name> is the sanitized, de-duplicated class name listed in probability_column_names rather than the class name itself;

  • confusion_counts — counts indexed and columned by class name;

  • confusion_normalized — the same matrix divided by its true-class row totals, with all-zero rows left at zero;

  • per_plate — the same scalar metrics per plate group, with plate as the first column;

  • calibration — the calibration_table() reliability bins for the calibrated probabilities.

spacr.classifier_evaluation.expected_calibration_error(y_true: Sequence[int], probabilities: Any, *, n_bins: int = 10) → float[source]

Return top-label expected calibration error (ECE).

Parameters:
  • y_true – true class index per sample, aligned with the rows of probabilities.

  • probabilities – predicted probabilities, either a positive-class vector or a samples-by-classes matrix; validated and row-normalised by normalize_probabilities(). Its length must equal that of y_true.

spacr.classifier_evaluation.find_evaluation_bundles(root: Any) → List[pathlib.Path][source]

Return evaluation manifests below root, newest first.

Parameters:

root – folder searched recursively for evaluation manifests, or a manifest file (returned alone) or any other file (its folder is searched). A missing path raises FileNotFoundError.

spacr.classifier_evaluation.fit_temperature(y_true: Sequence[int], probabilities: Any) → float[source]

Fit one scalar temperature by minimizing multiclass log loss.

Parameters:
  • y_true – true class index per sample; at least two samples and two distinct classes are required.

  • probabilities – uncalibrated predicted probabilities, a positive-class vector or a samples-by-classes matrix, normalised by normalize_probabilities() before the fit.

spacr.classifier_evaluation.grouped_split(groups: Sequence[Any], labels: Sequence[Any], holdout: float, seed: int = 0, *, group_by: Any = 'well', hold_out_groups: Sequence[Any] | None = None) → Tuple[numpy.ndarray, numpy.ndarray, SplitReport][source]

Return a stratified holdout while keeping named groups intact.

A grouped design is refused when either side cannot contain every class. This is intentionally stricter than silently scoring a model on siblings of its training rows or on a holdout that contains only one class.

Parameters:
  • groups – group identifier for every labelled object. All members of one group remain on the same side of the split.

  • labels – class label for every object, aligned to groups.

  • holdout – requested test fraction, strictly between zero and one.

  • hold_out_groups –

    groups that go to the TEST side whatever the fraction says. This is what holdout_plate is: cross-validation splits within the data it is given, so a model can learn the plate rather than the phenotype and every number it reports will look fine. Naming a plate here trains without it and scores on it, which is the one number that says whether a classifier generalises.

    The class check still applies: a named holdout that leaves either side without every class is refused, for the same reason a random one is.

spacr.classifier_evaluation.load_evaluation_bundle(path: Any) → Dict[str, Any][source]

Load one evaluation bundle for the Qt workbench.

Parameters:

path – an evaluation manifest file, or the bundle folder that contains it; ~ is expanded. Tables named by the manifest that are absent load as empty frames.

spacr.classifier_evaluation.nested_group_folds(labels: Sequence[int], *, outer_splits: int, inner_splits: int, groups: Sequence[Any] | None = None, seed: int = 0) → List[Dict[str, Any]][source]

Build nested stratified/grouped outer and inner index partitions.

Inner indexes are returned in the original/global coordinate system.

Parameters:
  • labels – class label per sample, coerced to int. Only its length and class balance matter – the folds carry indices, not data.

  • outer_splits – outer fold count, coerced with int (2.9 gives 2); below 2 is refused.

  • inner_splits – inner fold count built inside each outer TRAINING set, so it is bounded by that subset, not by the dataset: three inner folds over four samples fails on the outer training half even though the outer split itself succeeded.

  • groups – optional group key per sample (well or plate id) kept whole within a fold at BOTH levels. None gives a plain stratified split that will scatter crops of the same well across folds.

  • seed – outer folds use seed; the inner folds of outer fold k use seed + k, so inner partitions differ between outer folds instead of repeating one layout. Must be an int – a float reaches numpy.random.default_rng and raises TypeError.

Raises:

ValueError – when either split count is below 2, or when the sample count (or the number of distinct groups) cannot supply that many folds.

Returns:

one dict per outer fold, with outer_fold (1-based), train, validation and inner – a list of (train, validation) index pairs expressed in GLOBAL indices, not as positions inside train.

spacr.classifier_evaluation.normalize_probabilities(probabilities: Any, *, n_classes: int | None = None) → numpy.ndarray[source]

Return a finite, row-normalized (n_samples, n_classes) matrix.

Parameters:

probabilities – a one-dimensional positive-class vector, a one-column matrix (both expanded to two classes), or a samples-by-classes matrix. Values must be finite and within 0 to 1, and no row may sum to zero.

spacr.classifier_evaluation.normalize_split_level(group_by: Any) → str[source]

Return a canonical split level; legacy none means cell.

Parameters:

group_by – requested split level: cell, field, well or plate (exact, lower-case). None, False, 'none' and 'off' map to cell; anything else raises ValueError.

spacr.classifier_evaluation.sample_identity(path: Any) → Dict[str, str][source]

Parse plate/well/field/object identities from a crop filename.

Unknown levels are returned as empty strings rather than guessed. The object identity is the augmentation-normalized full stem.

Parameters:

path – crop path or file name. Augmentation suffixes are removed and the stem is split on underscores to find the plate, well and field parts.

spacr.classifier_evaluation.split_columns_for(group_by: Any, columns: Sequence[str], table: str = 'data') → Tuple[str, List[str]][source]

Resolve a split level to complete metadata columns or refuse it.

prcfo and crop filenames are handled by split_group_values(); this helper describes the direct-column route used by measurement tables. Partial keys are never accepted because, for example, columnID='c1' is not a well identity across rows and plates.

Parameters:
  • group_by – split level, normalised by normalize_split_level().

  • columns – column names available in the table; every identity column the level needs must be present or ValueError is raised.

spacr.classifier_evaluation.split_group_values(*, group_by: Any = 'well', frame: pandas.DataFrame | None = None, paths: Sequence[Any] | None = None, table: str = 'data') → Tuple[str, numpy.ndarray][source]

Build one strict group id per row from metadata, prcfo, or paths.

For grouped levels every identity must be verifiable. Inventing singleton ids for unparseable rows would make a leaking random split look grouped.

spacr.classifier_evaluation.write_evaluation_bundle(output_dir: Any, evaluation: Mapping[str, Any], *, leakage_reports: Sequence[LeakageReport] | None = None) → pathlib.Path[source]

Atomically write a complete evaluation bundle and diagnostic figures.

Parameters:
  • output_dir – folder that receives the bundle; created with its parents if needed.

  • evaluation – mapping as returned by evaluate_predictions(), with summary, predictions, confusion_counts, confusion_normalized, per_plate and calibration entries.

spacr.classifier_evaluation.write_leakage_audit(path: Any, audit: Any) → pathlib.Path[source]

Atomically write a leakage report/audit as JSON and return its path.

Parameters:
  • path – destination JSON file; ~ is expanded and missing parent folders are created. It is replaced atomically.

  • audit – leakage report or audit to write: any object with a to_dict() method, or a mapping.

spacr.classifier_evaluation.EVALUATION_FILES[source]

Stable file names produced by write_evaluation_bundle().

Nested helpers

fit_temperature.objective(log_temperature: float) → float

Return multiclass log loss after applying an exponentiated scale.

spacr/classifier_evaluation.py:1199

load_evaluation_bundle.read_csv(key: str, **kwargs) → pd.DataFrame

Load a manifest-named CSV, or an empty frame when it is absent.

spacr/classifier_evaluation.py:1866

write_evaluation_bundle.write_csv(name: str, frame: pd.DataFrame, *, index: bool = False) → None

Atomically replace name with frame and optional index.

spacr/classifier_evaluation.py:1709

write_evaluation_bundle.write_json(name: str, payload: Any) → None

Atomically replace name with stable indented JSON.

spacr/classifier_evaluation.py:1699