spacr.classifier_evaluation¶
Workflow inputs and outputs¶
Classifier Evaluation¶
Inspect held-out predictions, calibration and split/leakage evidence for the matching classifier.
Open: Classify → Classifier Evaluation.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Classifier evaluation bundle — Held-out predictions, labels, split metadata and calibration/leakage metrics for a saved classifier run.
Outputs
Quality-control results — Stored project checks and QC reports; a missing check is not a passing result.
Figures and table exports — The output location chosen by the tool; exports describe the selected data and filters.
Before this module
Classify: Use the matching held-out predictions and split metadata.
Classifier evaluation, calibration, and split-leakage diagnostics.
The module is model-agnostic: it consumes labels, probabilities, fold ids, and sample paths. Deep-learning CV and future classical-ML pipelines can therefore write the same evaluation bundle and the Qt workbench can display one stable artifact format.
Attributes¶
Stable file names produced by |
Exceptions¶
Raised when related samples cross a protected split boundary. |
Classes¶
Whole-CV proof that each related sample family belongs to one fold. |
|
Overlap counts and examples for one train/validation boundary. |
|
Provenance and realised sizes for one train/test split. |
Functions¶
|
Verify partition coverage and identity isolation across every CV fold. |
|
Audit the permanent |
|
Detect related images crossing a train/validation boundary. |
|
Return the original-sample family after removing augmentation suffixes. |
|
Return per-class reliability bins for calibration plots. |
|
Calibrate each held-out fold using every other out-of-fold prediction. |
|
Return sorted image paths under |
|
Build overall, confusion, calibration, and per-plate evaluation tables. |
|
Return top-label expected calibration error (ECE). |
|
Return evaluation manifests below |
|
Fit one scalar temperature by minimizing multiclass log loss. |
|
Return a stratified holdout while keeping named groups intact. |
|
Load one evaluation bundle for the Qt workbench. |
|
Build nested stratified/grouped outer and inner index partitions. |
|
Return a finite, row-normalized |
|
Return a canonical split level; legacy |
|
Parse plate/well/field/object identities from a crop filename. |
|
Resolve a split level to complete metadata columns or refuse it. |
|
Build one strict group id per row from metadata, |
|
Atomically write a complete evaluation bundle and diagnostic figures. |
|
Atomically write a leakage report/audit as JSON and return its path. |
Module Contents¶
- exception spacr.classifier_evaluation.LeakageError[source]¶
Bases:
ValueErrorRaised when related samples cross a protected split boundary.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.classifier_evaluation.FoldLeakageAudit[source]¶
Whole-CV proof that each related sample family belongs to one fold.
- Parameters:
group_by – canonical identity level required to remain within one validation fold.
n_samples – number of source paths whose fold membership was audited.
n_folds – number of train/validation fold pairs inspected.
validation_membership_missing – up to twenty sample indexes that were never held out for validation.
validation_membership_duplicate – up to twenty sample indexes held out in more than one fold.
overlap_counts – identities assigned to multiple validation folds, counted at each exact, content, family, object, and acquisition level.
examples – up to ten conflicting identities per level, annotated with the folds that contain them.
critical_levels – completeness, overlap, identity, or label failures that make
passedfalse.warnings – non-fatal caveats and explanations accompanying failures.
unverifiable_counts – samples whose requested identity or optional byte-content hash could not be checked, counted by level.
hash_errors – up to twenty file-specific content-hashing failures.
split_name – stable label for this whole-CV audit record.
- class spacr.classifier_evaluation.LeakageReport[source]¶
Overlap counts and examples for one train/validation boundary.
- Parameters:
group_by – protected split level (
cell,field,well, orplate); the legacynonespelling is normalized tocell.train_samples – number of training paths.
validation_samples – number of validation paths.
overlap_counts – overlap count at each identity level.
examples – up to ten shared identities per level.
split_name – caller-supplied label for the audited boundary.
critical_levels – levels that invalidate the requested split.
warnings – non-fatal caveats.
unverifiable_counts – samples lacking a requested identity or content hash, counted by the level that could not be verified.
hash_errors – up to twenty file-specific failures from optional byte-content hashing.
- class spacr.classifier_evaluation.SplitReport[source]¶
Provenance and realised sizes for one train/test split.
- Parameters:
group_by – canonical isolation unit used for the split, from objects through fields, wells, and plates.
requested_fraction – held-out object fraction requested by the caller.
cell_fraction – realised share of objects assigned to the test side.
group_fraction – realised share of distinct groups assigned to test.
train_cells – number of object rows used for fitting.
test_cells – number of object rows held out for evaluation.
train_groups – number of distinct groups used for fitting.
test_groups – number of distinct groups held out for evaluation.
total_groups – distinct groups across both sides of the split.
rule – human-readable algorithm and isolation guarantee that produced the realised split.
- spacr.classifier_evaluation.audit_cv_folds(paths: Sequence[Any], folds: Sequence[Tuple[Sequence[int], Sequence[int]]], *, labels: Sequence[Any] | None = None, group_by: str = 'well', hash_content: bool = False, require_identity: bool = True, raise_on_leakage: bool = False) FoldLeakageAudit[source]¶
Verify partition coverage and identity isolation across every CV fold.
This checks the fold assignment as one object, rather than trusting a sample of pairwise boundaries. Each index must be validation exactly once; exact paths, byte-identical content, source/augmentation families and the requested plate/well/field group must map to one held-out fold only.
- Parameters:
paths – one source path per sample, indexed by the fold indices.
folds –
(train_indices, validation_indices)per fold. Indices outsiderange(len(paths))are reported rather than ignored.labels – optional class per path, one value per path. Only used to describe the folds; it does not affect leakage detection.
group_by – identity level that may not cross a fold boundary –
none,field,wellorplate. A well-grouped split permits the same plate on both sides but never the same well.hash_content – also compare file CONTENT, so a byte-identical copy under a different name is caught. Costs one read per file.
require_identity – treat UNVERIFIABLE as critical. A filename that does not encode the requested identity, or a file that cannot be hashed, is otherwise only a warning – so leaving this False means a clean report can still hide leakage nobody could check for.
raise_on_leakage – raise
LeakageErrorinstead of returning a report whosepassedis False.
- Returns:
a
FoldLeakageAuditcarrying the per-level overlap counts, examples, and which levels were critical.- Raises:
ValueError – for an unsupported
group_by, orlabelswhose length does not matchpaths.
- spacr.classifier_evaluation.audit_dataset_splits(root: Any, *, group_by: str = 'well', hash_content: bool = True, require_identity: bool = True, raise_on_leakage: bool = False) LeakageReport[source]¶
Audit the permanent
train/versustest/dataset boundary.- Parameters:
root – dataset directory holding
train/andtest/. Both are searched recursively, and only.png,.jpg,.jpeg,.tif,.tiff,.bmpand.npyfiles are collected, so a folder of any other format audits as empty.group_by – identity level that may not appear on both sides –
none,field,wellorplate. The defaultwellstill permits the same plate in train and test. It is validated only after the images are found, so a bad value over an empty tree reports the missing images instead.hash_content – defaults to True here, unlike
audit_split_leakage(), so a byte-identical copy saved under a different name fails the audit – at the cost of one read per file.require_identity – defaults to True here: filenames that do not encode the
group_bylevel, and files that cannot be hashed, become critical instead of a warning, so an unverifiable split cannot report as clean.raise_on_leakage – raise
LeakageErrorinstead of returning a report whosepassedis False.
- Raises:
FileNotFoundError – when either side collects no image – a missing or unreadable folder is never treated as a passing split.
- Returns:
a
LeakageReportwhosesplit_nameis alwaystrain_vs_test; thetest/side is counted asvalidation_samples.
- spacr.classifier_evaluation.audit_split_leakage(train_paths: Sequence[Any], validation_paths: Sequence[Any], *, group_by: str = 'well', raise_on_leakage: bool = False, split_name: str = '', hash_content: bool = False, require_identity: bool = False) LeakageReport[source]¶
Detect related images crossing a train/validation boundary.
Exact/object/augmentation-family overlap is always critical. The requested
group_bylevel is also critical: a well-grouped split permits the same plate on both sides but never the same well.- Parameters:
train_paths – source paths used to fit the model.
validation_paths – paths used only for evaluation.
group_by –
cell,field,well, orplate. Legacynonealiasescell.raise_on_leakage – raise
LeakageErroron a critical overlap.split_name – optional fold/split label stored in the report.
hash_content – also compare file CONTENT, so a byte-identical copy under a different name is caught. Costs one read per file.
require_identity – treat UNVERIFIABLE as critical. A filename that does not encode the requested identity, or a file that cannot be hashed, is otherwise only a warning – so leaving this False means a clean report can still hide leakage nobody could check for.
- Returns:
- spacr.classifier_evaluation.augmentation_family(path: Any) str[source]¶
Return the original-sample family after removing augmentation suffixes.
In-memory spaCR augmentations retain the exact source filename, while older exported datasets use suffixes such as
_aug3,_rot90or_flip_h. Both forms collapse to one family.- Parameters:
path – crop path or file name; only its basename without the final extension is used, with trailing augmentation suffixes stripped repeatedly.
- spacr.classifier_evaluation.calibration_table(y_true: Sequence[int], probabilities: Any, *, classes: Sequence[str] | None = None, n_bins: int = 10) pandas.DataFrame[source]¶
Return per-class reliability bins for calibration plots.
- Parameters:
y_true – class INDEX per sample, compared column by column. A label outside the probability columns is REFUSED: it matches no class, so it used to read
observed_frequency0.0 everywhere and render as a catastrophically miscalibrated curve rather than an error.probabilities – predicted probabilities; a 1-D positive-class vector is expanded to two columns and every row is renormalized.
classes – display names in column order. An EMPTY sequence falls back to
class_0 ... class_n; a non-empty one whose length disagrees with the columns is refused.n_bins – equal-width confidence bins over
[0, 1], truncated to an int and floored at 2, so 0, 1 and any negative value all give two bins. The top bin is closed on the right, so confidence 1.0 lands in it rather than falling out of the table.
- Raises:
ValueError – when
y_trueandprobabilitiesdiffer in length,classeshas the wrong length, or a class index falls outside the probability columns.- Returns:
one row per class and NON-EMPTY bin, so the frame is shorter than
n_classes * n_binsrows. An empty result still carriesCALIBRATION_COLUMNS, so a column can be indexed on it.
- spacr.classifier_evaluation.cross_calibrate_probabilities(y_true: Sequence[int], probabilities: Any, fold_ids: Sequence[Any], *, method: str = 'temperature', warnings_out: List[str] | None = None) Tuple[numpy.ndarray, Dict[str, float]][source]¶
Calibrate each held-out fold using every other out-of-fold prediction.
This cross-fitting prevents a sample from fitting the calibrator that is evaluated on that same sample.
- Parameters:
y_true – class INDEX per sample, read only to fit the temperatures. A value outside the probability columns is REFUSED – it used to make every fold’s fit fail, fall back to temperature 1.0, and return uncalibrated probabilities while reporting that calibration ran.
probabilities – predicted probabilities; a 1-D positive-class vector is expanded to two columns and every row is renormalized, so the result is always a new normalized matrix – never the object passed in, even when no calibration is applied.
fold_ids – held-out block per sample; any hashable value works, and each distinct value is calibrated from all the others, so at least two distinct values are required. The returned map is keyed by
str(fold_id).method –
temperature, ornone/off/false(alsoNoneand the empty string) to return the normalized probabilities unchanged with an empty temperature map. Case and surrounding whitespace are ignored; anything else is refused rather than skipped.warnings_out – list appended in place when a fold cannot be calibrated from the others. The message is printed either way, so leaving this None only discards the machine-readable copy.
- Raises:
ValueError – on mismatched lengths, an unrecognized
method, or fewer than two distinctfold_idswhen calibrating.- Returns:
(calibrated_probabilities, temperature_by_held_out_fold).
- spacr.classifier_evaluation.dataset_split_paths(root: Any, split: str) List[str][source]¶
Return sorted image paths under
root/<split>/<class>/.- Parameters:
root – dataset folder that contains the split folders;
~is expanded.split – name of the split subfolder, such as
trainortest. A missing folder returns an empty list; files with an image or.npysuffix are collected recursively below it.
- spacr.classifier_evaluation.evaluate_predictions(y_true: Sequence[int], probabilities: Any, sample_paths: Sequence[Any], *, classes: Sequence[str] | None = None, fold_ids: Sequence[Any] | None = None, calibration_method: str = 'none', calibration_bins: int = 10) Dict[str, Any][source]¶
Build overall, confusion, calibration, and per-plate evaluation tables.
- Parameters:
y_true – true class INDEX per sample, in
range(n_classes). A value outside that range is refused rather than clipped.probabilities –
(n_samples, n_classes)predicted probabilities. Its column count definesn_classes.sample_paths – one source path per sample, used to derive the per-plate tables. Must be the same length as
y_true.classes – display names for the columns, in column order. Defaults to
class_0 ... class_n; a length that disagrees with the probability columns is refused.fold_ids – which held-out fold each sample came from. REQUIRED for temperature calibration and unused otherwise – the temperature is fitted per held-out fold so no sample is calibrated on itself, which needs at least two distinct folds.
calibration_method –
'none'(default) or'temperature'. Anything else is refused rather than silently ignored.calibration_bins – bin count for the expected-calibration-error estimate. More bins resolve the reliability curve better and make each bin noisier.
- Raises:
ValueError – on mismatched lengths, a class index outside the probability columns, an unknown
calibration_method, or temperature calibration with fewer than two distinct folds.- Returns:
the evaluation bundle, a dict with six keys:
summary— scalar metrics (n, accuracy, balanced accuracy, macro/weighted F1, macro precision/recall, log loss, multiclass Brier, expected calibration error, mean confidence) plusclasses,n_classes,calibration_method,raw_expected_calibration_error,temperatures_by_held_out_fold,calibration_warningsandprobability_column_names;predictions— one row per sample withfold, thesample_identity()columns,true_label/true_class,predicted_label/predicted_class,correct,confidence(the calibrated probability of the chosen class) and araw_prob_<name>/prob_<name>pair per class, where<name>is the sanitized, de-duplicated class name listed inprobability_column_namesrather than the class name itself;confusion_counts— counts indexed and columned by class name;confusion_normalized— the same matrix divided by its true-class row totals, with all-zero rows left at zero;per_plate— the same scalar metrics perplategroup, withplateas the first column;calibration— thecalibration_table()reliability bins for the calibrated probabilities.
- spacr.classifier_evaluation.expected_calibration_error(y_true: Sequence[int], probabilities: Any, *, n_bins: int = 10) float[source]¶
Return top-label expected calibration error (ECE).
- Parameters:
y_true – true class index per sample, aligned with the rows of
probabilities.probabilities – predicted probabilities, either a positive-class vector or a samples-by-classes matrix; validated and row-normalised by
normalize_probabilities(). Its length must equal that ofy_true.
- spacr.classifier_evaluation.find_evaluation_bundles(root: Any) List[pathlib.Path][source]¶
Return evaluation manifests below
root, newest first.- Parameters:
root – folder searched recursively for evaluation manifests, or a manifest file (returned alone) or any other file (its folder is searched). A missing path raises
FileNotFoundError.
- spacr.classifier_evaluation.fit_temperature(y_true: Sequence[int], probabilities: Any) float[source]¶
Fit one scalar temperature by minimizing multiclass log loss.
- Parameters:
y_true – true class index per sample; at least two samples and two distinct classes are required.
probabilities – uncalibrated predicted probabilities, a positive-class vector or a samples-by-classes matrix, normalised by
normalize_probabilities()before the fit.
- spacr.classifier_evaluation.grouped_split(groups: Sequence[Any], labels: Sequence[Any], holdout: float, seed: int = 0, *, group_by: Any = 'well', hold_out_groups: Sequence[Any] | None = None) Tuple[numpy.ndarray, numpy.ndarray, SplitReport][source]¶
Return a stratified holdout while keeping named groups intact.
A grouped design is refused when either side cannot contain every class. This is intentionally stricter than silently scoring a model on siblings of its training rows or on a holdout that contains only one class.
- Parameters:
groups – group identifier for every labelled object. All members of one group remain on the same side of the split.
labels – class label for every object, aligned to
groups.holdout – requested test fraction, strictly between zero and one.
hold_out_groups –
groups that go to the TEST side whatever the fraction says. This is what
holdout_plateis: cross-validation splits within the data it is given, so a model can learn the plate rather than the phenotype and every number it reports will look fine. Naming a plate here trains without it and scores on it, which is the one number that says whether a classifier generalises.The class check still applies: a named holdout that leaves either side without every class is refused, for the same reason a random one is.
- spacr.classifier_evaluation.load_evaluation_bundle(path: Any) Dict[str, Any][source]¶
Load one evaluation bundle for the Qt workbench.
- Parameters:
path – an evaluation manifest file, or the bundle folder that contains it;
~is expanded. Tables named by the manifest that are absent load as empty frames.
- spacr.classifier_evaluation.nested_group_folds(labels: Sequence[int], *, outer_splits: int, inner_splits: int, groups: Sequence[Any] | None = None, seed: int = 0) List[Dict[str, Any]][source]¶
Build nested stratified/grouped outer and inner index partitions.
Inner indexes are returned in the original/global coordinate system.
- Parameters:
labels – class label per sample, coerced to int. Only its length and class balance matter – the folds carry indices, not data.
outer_splits – outer fold count, coerced with
int(2.9gives 2); below 2 is refused.inner_splits – inner fold count built inside each outer TRAINING set, so it is bounded by that subset, not by the dataset: three inner folds over four samples fails on the outer training half even though the outer split itself succeeded.
groups – optional group key per sample (well or plate id) kept whole within a fold at BOTH levels. None gives a plain stratified split that will scatter crops of the same well across folds.
seed – outer folds use
seed; the inner folds of outer foldkuseseed + k, so inner partitions differ between outer folds instead of repeating one layout. Must be an int – a float reachesnumpy.random.default_rngand raisesTypeError.
- Raises:
ValueError – when either split count is below 2, or when the sample count (or the number of distinct groups) cannot supply that many folds.
- Returns:
one dict per outer fold, with
outer_fold(1-based),train,validationandinner– a list of(train, validation)index pairs expressed in GLOBAL indices, not as positions insidetrain.
- spacr.classifier_evaluation.normalize_probabilities(probabilities: Any, *, n_classes: int | None = None) numpy.ndarray[source]¶
Return a finite, row-normalized
(n_samples, n_classes)matrix.- Parameters:
probabilities – a one-dimensional positive-class vector, a one-column matrix (both expanded to two classes), or a samples-by-classes matrix. Values must be finite and within 0 to 1, and no row may sum to zero.
- spacr.classifier_evaluation.normalize_split_level(group_by: Any) str[source]¶
Return a canonical split level; legacy
nonemeanscell.- Parameters:
group_by – requested split level:
cell,field,wellorplate(exact, lower-case).None,False,'none'and'off'map tocell; anything else raisesValueError.
- spacr.classifier_evaluation.sample_identity(path: Any) Dict[str, str][source]¶
Parse plate/well/field/object identities from a crop filename.
Unknown levels are returned as empty strings rather than guessed. The object identity is the augmentation-normalized full stem.
- Parameters:
path – crop path or file name. Augmentation suffixes are removed and the stem is split on underscores to find the plate, well and field parts.
- spacr.classifier_evaluation.split_columns_for(group_by: Any, columns: Sequence[str], table: str = 'data') Tuple[str, List[str]][source]¶
Resolve a split level to complete metadata columns or refuse it.
prcfoand crop filenames are handled bysplit_group_values(); this helper describes the direct-column route used by measurement tables. Partial keys are never accepted because, for example,columnID='c1'is not a well identity across rows and plates.- Parameters:
group_by – split level, normalised by
normalize_split_level().columns – column names available in the table; every identity column the level needs must be present or
ValueErroris raised.
- spacr.classifier_evaluation.split_group_values(*, group_by: Any = 'well', frame: pandas.DataFrame | None = None, paths: Sequence[Any] | None = None, table: str = 'data') Tuple[str, numpy.ndarray][source]¶
Build one strict group id per row from metadata,
prcfo, or paths.For grouped levels every identity must be verifiable. Inventing singleton ids for unparseable rows would make a leaking random split look grouped.
- spacr.classifier_evaluation.write_evaluation_bundle(output_dir: Any, evaluation: Mapping[str, Any], *, leakage_reports: Sequence[LeakageReport] | None = None) pathlib.Path[source]¶
Atomically write a complete evaluation bundle and diagnostic figures.
- Parameters:
output_dir – folder that receives the bundle; created with its parents if needed.
evaluation – mapping as returned by
evaluate_predictions(), withsummary,predictions,confusion_counts,confusion_normalized,per_plateandcalibrationentries.
- spacr.classifier_evaluation.write_leakage_audit(path: Any, audit: Any) pathlib.Path[source]¶
Atomically write a leakage report/audit as JSON and return its path.
- Parameters:
path – destination JSON file;
~is expanded and missing parent folders are created. It is replaced atomically.audit – leakage report or audit to write: any object with a
to_dict()method, or a mapping.
- spacr.classifier_evaluation.EVALUATION_FILES[source]¶
Stable file names produced by
write_evaluation_bundle().
Nested helpers¶
- fit_temperature.objective(log_temperature: float) float¶
Return multiclass log loss after applying an exponentiated scale.
spacr/classifier_evaluation.py:1199
- load_evaluation_bundle.read_csv(key: str, **kwargs) pd.DataFrame¶
Load a manifest-named CSV, or an empty frame when it is absent.
spacr/classifier_evaluation.py:1866