spacr.training_basis

One vocabulary for “what defines a training class”, shared by Classify (CV) and Classify (ML).

The two modules did the same job in different words. Of 78 CV settings and 37 ML ones, six shared a name — and two of those six disagreed on their default (annotation_column was 'test' in one and None in the other; n_jobs was 28 and -1). Three more pairs were the same setting under different names. A settings CSV was therefore not portable between them, and neither was a user’s understanding.

The training basis. Classify (CV) already had all three: dataset_mode is 'metadata', 'annotation' or 'measurement', and spacr.io builds the dataset from class_metadata/metadata_rules, annotation_columns/annotation_values accordingly. Classify (ML) had two, and chose between them implicitly: ml.py asked whether annotation_column was None. Nothing said so in the settings panel, so a user who filled in an annotation column silently stopped training on their plate controls.

So there is no new concept to invent here. dataset_mode becomes the shared name, ML gains the basis it lacked, and the choice becomes something the user makes rather than something they trigger.

Backward compatibility is the whole difficulty. A settings CSV written before this exists in every user’s project folder, and INVARIANTS §6 is the trap: a key absent from the dict means the pipeline falls back to its own default, which can differ from the GUI’s, and nothing says so. Every rename here is therefore an alias, not a replacement — normalize_settings() translates the old name and the old implicit basis into the new explicit one, and a run from an old CSV does exactly what it did before.

Exceptions

TrainingBasisError

A basis that spaCR does not have, or one that cannot run as configured.

Functions

describe_basis(→ str)

One line for the settings panel, naming what the user must fill in.

inapplicable_settings(→ Tuple[str, ...])

Settings belonging to the OTHER bases -- what the GUI greys out.

normalize_settings(→ Dict[str, Any])

Return settings in the shared vocabulary. Never modifies the input.

resolve_basis(→ str)

Return the training basis a settings dict asks for.

settings_for_basis(→ Tuple[str, ...])

The settings that apply to basis.

Module Contents

exception spacr.training_basis.TrainingBasisError[source]

Bases: ValueError

A basis that spaCR does not have, or one that cannot run as configured.

Initialize self. See help(type(self)) for accurate signature.

spacr.training_basis.describe_basis(basis: str) → str[source]

One line for the settings panel, naming what the user must fill in.

Parameters:

basis – normalized training-basis name to describe.

spacr.training_basis.inapplicable_settings(basis: str) → Tuple[str, ...][source]

Settings belonging to the OTHER bases – what the GUI greys out.

Greyed, not removed. INVARIANTS §6: a key absent from the dict makes the pipeline fall back to its own default, which can differ from the value the module needs. A greyed control keeps its value and stops being editable; a deleted one changes the run.

Parameters:

basis – the chosen basis.

Returns:

setting keys that do not apply to it.

spacr.training_basis.normalize_settings(settings: Mapping[str, Any]) → Dict[str, Any][source]

Return settings in the shared vocabulary. Never modifies the input.

Applies SETTING_ALIASES and pins dataset_mode to whatever resolve_basis() worked out, so every consumer downstream reads one name and one explicit basis.

The new name wins when both are present. Someone who has set the current key has said what they mean; a stale alias left in the same CSV must not override it.

Parameters:

settings – the run settings.

Returns:

a new dict.

spacr.training_basis.resolve_basis(settings: Mapping[str, Any]) → str[source]

Return the training basis a settings dict asks for.

Precedence, and the reason:

  1. dataset_mode, when set. It is the explicit answer.

  2. Otherwise, the ML module’s historical implicit rule: an annotation_column that is set meant “train on annotations”. This is what makes an old settings CSV behave exactly as it used to.

  3. Otherwise 'metadata', which is what both modules defaulted to.

Parameters:

settings – the run settings.

Returns:

one of TRAINING_BASES.

Raises:

TrainingBasisError – an unrecognised dataset_mode. Silently falling back would train on the wrong labels and report success.

spacr.training_basis.settings_for_basis(basis: str) → Tuple[str, ...][source]

The settings that apply to basis.

Parameters:

basis – one of TRAINING_BASES.

Returns:

the setting keys that basis reads.

Raises:

TrainingBasisError – unknown basis.