spacr.surrogate

Workflow inputs and outputs

Explain CV Model

Fit a measured-feature surrogate to existing computer-vision scores, preserving identity joins. A surrogate explanation is not causality.

Open: Classify → Explain CV Model.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route: cell, nucleus, pathogen, cytoplasm. Relevant columns, depending on the route: plateID, rowID, columnID, fieldID.

  • Object classification scores — Saved score CSVs and, when merged, measurements/measurements.db, table png_list. Relevant tables, depending on the route: png_list. Relevant columns, depending on the route: pred, cv_predictions, ml_pred, predictions.

Outputs

  • Fitted feature classifier — The fitted tabular classifier and its recorded feature list, training settings and validation results.

  • Figures and table exports — The output location chosen by the tool; exports describe the selected data and filters.

Before this module

  • Classify: Select CV predictions and the matching measured objects.

API reference.

Module tutorial.

Explain a computer-vision classifier with a model of measured features.

Activation maps say where a model looked. This says what it responded to, in units a biologist already interprets — area, intensity, texture — by fitting a gradient-boosted model to reproduce the CV model’s predictions from the object features spaCR already measured, and then asking that model what it used.

The two answer different questions and neither replaces the other. A map that lights up the parasite tells you the model attends to the parasite; it does not tell you whether it is responding to the parasite’s size, its brightness, or how far it sits from the nucleus.

Fidelity is reported first, and it is not a formality. A surrogate that cannot reproduce the CV model’s decisions has explained nothing, and its feature ranking is a ranking of how the features predict each other. That number is the one people skip, so SurrogateResult puts it before the importances and SurrogateResult.summary() refuses to lead with anything else.

Three importances, because they disagree.

  • gain — cheap, comes free with the fit, and is biased towards high-cardinality features. Read it last.

  • permutation — measured on held-out data by breaking one feature at a time. Costs a re-scoring per feature and is worth it.

  • SHAP — additive per-object attributions. The only one that says which direction a feature pushed a particular object, and the one that connects to spacr.attribution’s pixel-level story if that gains the SHAP family too.

Where they disagree, the disagreement is the finding: a feature ranked high by gain and low by permutation is usually one the model could have used and did not.

The join goes through png_list. A CV model keys on the crop it scored; the features key on the object. png_list is the only table that holds both, so it is the bridge — and joining on png_path alone is not enough, because a path is not stable across machines.

Exceptions

SurrogateError

The surrogate cannot be built, or cannot be believed.

Classes

SurrogateResult

What the surrogate learned, with the caveat attached.

Functions

available_backends(→ Dict[str, Dict[str, Any]])

Report the supported surrogate families and whether they can run.

available_model_families(→ Dict[str, Dict[str, Any]])

Return every fixed surrogate choice and whether it can run here.

build_surrogate_frame(→ pandas.DataFrame)

Join CV predictions to the measured feature table.

explain_classifier(→ SurrogateResult)

Join predictions to features, fit a surrogate, and rank the features.

explain_cv_default_settings(→ Dict[str, Any])

Defaults for the first-class Explain CV Model module.

fit_surrogate(→ SurrogateResult)

Fit a surrogate to frame['cv_prediction'] and rank the features.

importance_method_availability(→ Dict[str, Dict[str, Any]])

Which importance measures apply to a model, and why the others do not.

rank_feature_importance(, shap_explainer, n_repeats, ...)

Rank the features of any fitted tabular classifier or regressor.

register_explain_cv_settings(→ bool)

Register Explain CV Model's settings with both desktop front ends.

run_explain_cv(→ Dict[str, Any])

Explain an existing CV prediction file without rerunning its model.

write_surrogate_result(→ Dict[str, str])

Write the complete, provenance-bearing surrogate result bundle.

Module Contents

exception spacr.surrogate.SurrogateError[source]

Bases: RuntimeError

The surrogate cannot be built, or cannot be believed.

Initialize self. See help(type(self)) for accurate signature.

class spacr.surrogate.SurrogateResult[source]

What the surrogate learned, with the caveat attached.

Parameters:
  • fidelity – held-out accuracy at reproducing the CV model.

  • baseline – the accuracy of always guessing the commonest class. Fidelity is only meaningful against this.

  • importance – one row per feature, with gain, permutation and shap columns where each was computed.

  • n_objects – rows the surrogate was fitted on.

  • class_counts – how many objects the CV model put in each class.

  • split_report – JSON-safe grouped-split provenance and realised training and held-out group and object counts.

  • warnings – anything that should be read before the ranking is.

  • model_family – canonical requested estimator family: "random_forest", "hist_gradient_boosting", or "xgboost".

  • backend – fully qualified class name of the estimator actually fitted.

  • backend_version – installed distribution version for that estimator.

  • random_seed – seed shared by grouped splitting, estimator fitting, and permutation importance so the run can be reproduced.

  • model_params – effective estimator constructor options after defaults and caller overrides are combined.

  • feature_columns – ordered numeric features actually fitted after identifier, leakage, and explicit exclusions.

  • excluded_columns – sorted union of automatically detected answer-leaking features and columns named by the caller’s exclude.

  • balanced_accuracy – held-out balanced accuracy of surrogate predictions against the CV model’s decisions.

  • f1_macro – macro-averaged F1 on the held-out objects.

  • class_metrics – held-out precision, recall, F1, and support indexed by CV class.

  • confusion – held-out confusion matrix with CV classes on rows and surrogate predictions on columns.

  • held_out – held-out identifiers, CV and surrogate predictions, and per-class probabilities when the backend supplies them.

  • shap_values – signed SHAP contributions for sampled held-out objects; empty when SHAP is unavailable or fails.

  • shap_feature_values – measured values for the same sampled rows and feature columns as shap_values.

  • correlated_features – strong training-split feature pairs and their Spearman coefficients at the configured absolute threshold.

  • feature_distributions – long held-out table of count, mean, median, and standard deviation for each feature within each CV class.

  • family_importance – gain, permutation, and SHAP importance totals grouped by the feature-family heuristic.

  • model – fitted estimator retained for downstream inspection and omitted from the dataclass representation.

  • minimum_fidelity_improvement – minimum accuracy gain over the majority baseline required by is_faithful before rankings are presented.

summary(n: int = 15) → str[source]

A report that leads with the caveat, because the caveat decides whether the rest is worth reading.

top(n: int = 15, by: str = 'permutation') → pandas.DataFrame[source]

The n most important features by one measure.

Parameters:
  • n – how many rows.

  • by – 'permutation', 'shap' or 'gain'. Falls back to whichever was computed, rather than raising, so a run without SHAP still reports something.

property fidelity_improvement: float[source]

Accuracy improvement over the majority-class baseline.

property is_faithful: bool[source]

Whether the surrogate beats the majority-class baseline at all.

Deliberately a low bar. It is not “this explanation is good”; it is “this explanation is about the CV model rather than about noise”.

spacr.surrogate.available_backends() → Dict[str, Dict[str, Any]][source]

Report the supported surrogate families and whether they can run.

XGBoost is optional and is never silently replaced by another estimator.

spacr.surrogate.available_model_families() → Dict[str, Dict[str, Any]][source]

Return every fixed surrogate choice and whether it can run here.

XGBoost is optional and remains visible when absent so the GUI can disable it with an explanation. It must never silently become Random Forest under an XGBoost label.

spacr.surrogate.build_surrogate_frame(db_path: str, predictions: pandas.DataFrame, path_column: str = 'path', prediction_column: str = 'pred') → pandas.DataFrame[source]

Join CV predictions to the measured feature table.

The CV model keys on the crop it scored and the features key on the object, so png_list is the bridge: it is the only table holding both. The join is validated one-to-one — a repeated key would multiply rows and silently reweight the surrogate towards whichever objects happened to duplicate.

Parameters:
  • db_path – the measurements database.

  • predictions – a frame with a crop path and a predicted class, as apply_model_to_tar returns.

  • path_column – the column holding the crop path.

  • prediction_column – the column holding the predicted class.

Returns:

features joined to a cv_prediction column.

Raises:

SurrogateError – a missing column, or a join that matches nothing.

spacr.surrogate.explain_classifier(db_path: str, predictions: pandas.DataFrame, *, path_column: str = 'path', prediction_column: str = 'pred', **fit_kwargs: Any) → SurrogateResult[source]

Join predictions to features, fit a surrogate, and rank the features.

Parameters:
  • db_path – the measurements database.

  • predictions – a frame with a crop path and a predicted class.

  • path_column – column holding the crop path.

  • prediction_column – column holding the predicted class.

  • fit_kwargs – forwarded to fit_surrogate().

Returns:

a SurrogateResult.

spacr.surrogate.explain_cv_default_settings(settings=None) → Dict[str, Any][source]

Defaults for the first-class Explain CV Model module.

spacr.surrogate.fit_surrogate(frame: pandas.DataFrame, *, test_size: float = 0.3, n_estimators: int = 300, random_seed: int = 0, n_repeats: int = 5, shap_max_samples: int = 500, exclude: Sequence[str] | None = None, split_by: str = 'well', model_family: str = 'random_forest', model_options: Mapping[str, Any] | None = None, correlation_threshold: float = 0.9, minimum_fidelity_improvement: float = 0.05, importance_methods: Sequence[str] = IMPORTANCE_METHODS, shap_explainer: str = 'auto', verbose: bool = True) → SurrogateResult[source]

Fit a surrogate to frame['cv_prediction'] and rank the features.

Parameters:
  • frame – as build_surrogate_frame() returns.

  • test_size – held-out fraction. Fidelity and permutation importance are both measured on it, never on the training rows.

  • n_estimators – trees in the surrogate.

  • random_seed – fixed, so a reported ranking can be reproduced.

  • n_repeats – permutation repeats per feature.

  • shap_max_samples – SHAP is O(rows); this caps the explained sample and the cap is REPORTED rather than applied silently.

  • exclude – extra feature columns to drop.

  • split_by – acquisition unit held intact between surrogate fitting and fidelity measurement. Default 'well'.

  • importance_methods – which of IMPORTANCE_METHODS to compute.

  • shap_explainer – one of SHAP_EXPLAINERS.

  • verbose – print the summary when done.

Returns:

a SurrogateResult.

Raises:

SurrogateError – too few objects or classes to fit anything.

spacr.surrogate.importance_method_availability(model: Any = None, model_family: str | None = None) → Dict[str, Dict[str, Any]][source]

Which importance measures apply to a model, and why the others do not.

Parameters:
  • model – a fitted estimator, when there is one.

  • model_family – a key of MODEL_FAMILIES, when there is no fitted model yet (the Explain CV Model form).

Returns:

{name: {'available': bool, 'reason': str}} for gain, permutation, tree_shap and kernel_shap.

spacr.surrogate.rank_feature_importance(model: Any, x: pandas.DataFrame, y: Any, *, methods: Sequence[str] = ('gain', 'permutation', 'shap'), shap_explainer: str = 'auto', n_repeats: int = 5, random_state: int = 0, shap_max_samples: int = 200, scoring: str | None = None, destination: str | None = None, title: str = 'Feature importance') → Tuple[pandas.DataFrame, Dict[str, str]][source]

Rank the features of any fitted tabular classifier or regressor.

The same three measures the surrogate reports, for a model that did not come from fit_surrogate(): native (gain / impurity) importance read from the model, sklearn permutation importance on (x, y), and mean absolute SHAP (TreeSHAP for tree ensembles, KernelSHAP otherwise). Pass held-out rows: permutation importance on the training rows measures what the model memorised.

A measure that does not apply to this model (gain for a model without feature_importances_, SHAP without the shap package) is left out and the reason is recorded in the returned table’s attrs['warnings'].

Parameters:
  • model – a fitted estimator with predict.

  • x – feature rows, one column per feature.

  • y – the target for those rows.

  • methods – which of IMPORTANCE_METHODS to compute.

  • shap_explainer – one of SHAP_EXPLAINERS.

  • n_repeats – permutation repeats per feature.

  • random_state – seed for permutation and SHAP subsampling.

  • shap_max_samples – rows SHAP explains at most.

  • scoring – sklearn scorer name for permutation importance; the model’s own score when None.

  • destination – folder for feature_importance.csv and the ranked bar plot; nothing is written when None.

  • title – bar-plot title.

Returns:

(table, paths) – the table is sorted by the first computed of permutation, SHAP, gain; paths maps artifact role to path.

Raises:

SurrogateError – for an unknown method or explainer.

spacr.surrogate.register_explain_cv_settings(replace: bool = False) → bool[source]

Register Explain CV Model’s settings with both desktop front ends.

spacr.surrogate.run_explain_cv(settings: Mapping[str, Any]) → Dict[str, Any][source]

Explain an existing CV prediction file without rerunning its model.

Parameters:

settings – mapping returned by explain_cv_default_settings(). db_path is the exact measurements database and predictions_file is an existing per-object prediction CSV.

Returns:

{'result': SurrogateResult, 'paths': artifact mapping}.

Raises:

SurrogateError – when either source is missing or the requested optional backend cannot run.

spacr.surrogate.write_surrogate_result(result: SurrogateResult, destination: str) → Dict[str, str][source]

Write the complete, provenance-bearing surrogate result bundle.

Importance tables are always retained as source data, but plots are only presented when the held-out fidelity clears the declared improvement gate.

Parameters:
  • result – fitted SurrogateResult.

  • destination – new or existing output directory.

Returns:

mapping of artifact role to absolute path.

Nested helpers

_kernel_shap_values._predict(values)

Call the model on a frame with the training column names.

spacr/surrogate.py:827

write_surrogate_result._csv(role: str, name: str, frame: pd.DataFrame) → None

Write and register one non-empty result table.

Parameters:
  • role – artifact key to add to the captured path mapping.

  • name – CSV filename beneath the captured output directory.

  • frame – result table to write with its index; None and empty tables are deliberately omitted.

Returns:

None. A written table’s absolute path is stored in the captured artifact mapping.

spacr/surrogate.py:1109