spacr.classify¶
Workflow inputs and outputs¶
Classify¶
Choose Computer Vision for crops or streamed object images, or Tabular Machine Learning for measured feature columns. The two families consume different inputs.
Open: Home → Classify.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route:
cell,nucleus,pathogen,cytoplasm. Relevant columns, depending on the route:plateID,rowID,columnID,fieldID.Object crops — data/**/*_png when save_png is enabled; png_list indexes saved crops. Supported workflows can instead stream crops from merged arrays and masks. Relevant tables, depending on the route:
png_list. Relevant columns, depending on the route:png_path,prcfo.Training annotations — A chosen annotation column in measurements/measurements.db, table png_list; labels belong to object identities. Relevant tables, depending on the route:
png_list. Relevant columns, depending on the route:prcfo.
Outputs
Object classification scores — Saved score CSVs and, when merged, measurements/measurements.db, table png_list. Relevant tables, depending on the route:
png_list. Relevant columns, depending on the route:pred,cv_predictions,ml_pred,predictions.Trained image classifier — Saved classifier checkpoint, model settings and matching channel/preprocessing configuration.
Fitted feature classifier — The fitted tabular classifier and its recorded feature list, training settings and validation results.
Classifier evaluation bundle — Held-out predictions, labels, split metadata and calibration/leakage metrics for a saved classifier run.
Before this module
Measure: Choose the image or tabular family to match your input.
Annotate: Keep labelled training and evaluation groups separate.
Gate Editor: Write reviewed gate selections to an annotation column, then select that same column and object population in Classify. Keep validation objects separate from training labels.
Image UMAP: Write reviewed lasso selections to an annotation column in the matching object database, then select that column in Classify. Inspect crops and validate labels; embedding clusters are not ground truth.
After this module
Regression: Select the intended CV or ML score column and preserve plate/well identity.
Classifier Evaluation: Use the matching held-out predictions and split metadata.
Activation: Also provide the matching crops and channel preprocessing.
Explain CV Model: Select CV predictions and the matching measured objects.
One Classify entry point over both classifier families.
What it is for. Classify is the fourth step of the pipeline, after Mask,
Measure and Annotate. It learns the classes a screen is scored on and
predicts them for every object. classifier_family chooses how: cv
trains an image model on object crops through
spacr.deep_spacr.deep_spacr(), and ml fits a classical model such as
XGBoost, LightGBM or a random forest on measured features through
spacr.ml.generate_ml_scores().
What it needs. A project that has been through Mask and Measure, so that
measurements/measurements.db and its object crops exist, and a definition
of the classes. dataset_mode sets that definition for both families:
metadata takes classes from plate metadata such as the positive and
negative control wells, and annotation takes them from an annotation
column of png_list, as written by Annotate. The CV family reads crops from
PNG files or cuts them from the merged arrays, as crop_source selects.
What it produces. A CV run writes model checkpoints,
DL_model_settings.csv, a dataset tar, a top_examples/ folder with the
most confident crops of each class and an evaluation bundle with the held-out
performance and leakage.json, and merges each object’s predicted class and
probability into png_list. An ML run writes per-object predictions,
feature-importance and permutation tables and a plate heatmap to
results/ beside the measurements database.
What to do next. Check the held-out performance, and the leakage verdict on the QC screen, before trusting the predictions. Then take per-well scores into Regression, which pairs them with the guide counts from Map Barcodes to estimate which guides and genes explain the phenotype.
Classify (CV) trains a Torch model on object crops. Classify (ML) fits a gradient-boosted model on measured features. They answer the same question – which class is this object – and until now they were two modules that shared six setting names out of 78 and 37, two of which disagreed on their default.
spacr.training_basis already unified what defines a CLASS. This module
unifies what runs: one settings dict, one classifier_family switch, one
call. The two original modules stay exactly as they are – a merged screen
that removed them would strand every saved settings CSV and every notebook
that imports their entry points.
Nothing here reimplements either pipeline. deep_spacr and
generate_ml_scores are called unchanged, which is what makes the merged
module honest: a run through it and a run through the module it dispatches to
produce the same result, because they are the same code.
Exceptions¶
A classifier family spaCR does not have. |
Functions¶
|
Run whichever classifier family |
|
Settings belonging to the OTHER family -- what the panel greys out. |
|
Return the classifier family a settings dict asks for. |
|
Return the classical-ML estimator selected by |
Module Contents¶
- exception spacr.classify.ClassifierFamilyError[source]¶
Bases:
ValueErrorA classifier family spaCR does not have.
Initialize self. See help(type(self)) for accurate signature.
- spacr.classify.classify(settings: Mapping[str, Any]) Any[source]¶
Run whichever classifier family
settingsasks for.The merged module’s pipeline entry point. It normalises the shared vocabulary, resolves the family, and calls the existing entry point unchanged – so a run here and a run through Classify (CV) or Classify (ML) are the same run, not two implementations that have to be kept in step.
- Parameters:
settings – the run settings.
- Returns:
whatever the dispatched pipeline returns.
- Raises:
ClassifierFamilyError – an unrecognised family.
ValueError – when pre-dispatch validation rejects the selected ML estimator or CV crop source.
- spacr.classify.inapplicable_settings(family: str) Tuple[str, ...][source]¶
Settings belonging to the OTHER family – what the panel greys out.
Greyed, never removed: INVARIANTS §6. A key absent from the dict makes the pipeline fall back to its own default, which can differ from the value the module needs and says nothing when it does.
- Parameters:
family – the chosen family.
- Returns:
setting keys the other family owns.
- Raises:
ClassifierFamilyError – unknown family.
- spacr.classify.resolve_family(settings: Mapping[str, Any]) str[source]¶
Return the classifier family a settings dict asks for.
Defaults to
'cv', because the merged module’s own default settings are the CV ones and a dict with no family is most likely a Classify (CV) CSV opened in the merged screen.- Parameters:
settings – the run settings.
- Returns:
'cv'or'ml'.- Raises:
ClassifierFamilyError – an unrecognised family. Guessing would train a different kind of model than the user asked for and report success.
- spacr.classify.resolve_ml_model_type(settings: Mapping[str, Any]) str[source]¶
Return the classical-ML estimator selected by
settings.model_type_mlis authoritative in a merged payload.model_typeis accepted only when the ML-specific key is absent, which migrates the short-lived shared-vocabulary settings files written before the two model controls were separated. The default matchesspacr.settings.set_default_analyze_screen().- Parameters:
settings – settings for an ML-family run.
- Returns:
a member of
ML_MODEL_TYPES.- Raises:
ValueError – when the selected value is not an ML estimator.