spacr.train_compare

Workflow inputs and outputs

Training Runs

Compare saved training histories, metrics and settings from compatible model runs.

Open: Classify → Training Runs.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Run history and artifacts — Project run records, settings, output paths, artifact provenance, status and logs.

Outputs

  • Run/model comparison — Comparison tables and figures from compatible saved runs or masks; agreement is not ground-truth accuracy.

API reference.

Module tutorial.

Overlay several training runs’ curves and say exactly what differed between them.

Why this exists

Picking a classifier in spaCR means training it, writing down the number, changing a knob, training it again, and then trying to remember which of the six things you touched is the one that moved the number. The provenance diff (spacr.run_journal.diff_runs()) answers “what changed”; this module answers “and what did it do to the curves”, on one axis, next to the diff.

What is on disk, and therefore what can be shown

spacr.deep_spacr.train_model() evaluates train (and validation, when a validation loader exists) at the end of every epoch and hands the metric dicts to spacr.io._save_progress(), which appends them to two CSVs in the run’s dst folder:

<src>/model/<model_type>/<channels>/epochs_<N>/
    train.csv          # one row per epoch
    validation.csv     # one row per epoch, only when val_loaders exist
    <model_type>_epoch_<e>_channels_<ch>.pth

Columns come from spacr.deep_spacr.evaluate_model_performance(): epoch, loss, accuracy (and its duplicate Accuracy), neg_accuracy, pos_accuracy, prauc, optimal_threshold, plus train_time on the train side and num_classes on the multiclass side.

A k-fold run (cross_validation_folds >= 2) trains a fresh model per fold in <dst>/fold_<i>/, so it produces k independent pairs of curves, plus the per-fold summary CSVs spacr.deep_spacr._cross_validate_model() writes at the dst level.

The settings that produced a run are not in the run folder: save_settings writes them to <src>/settings/train_test_<model_type>_<epochs>.csv as a Key,Value CSV. load_run() walks up from the run folder to find it — but only as far as the run’s own project root, which the layout above makes identifiable: <src> is the folder the model/ tree hangs off. An ancestor above that belongs to a different project, and reading it would report somebody else’s settings as this run’s provenance, so the settings diff would show differences nobody configured and the answer would depend on where in the filesystem the project happened to sit. See _owns_this_run().

Judgement calls baked in

Runs of different length are never aligned. Each series is plotted over its own epoch axis. Truncating a 60-epoch run to match a 25-epoch neighbour invents a result (“B never got better”) and padding invents a different one; both are worse than a legend that says one run is longer.

Train and validation are never silently mixed. Every series label carries run, split and fold, and plot_curves() annotates the axes when more than one split is on it. A train curve above a validation curve is the definition of overfitting, not evidence that one run beat another.

k-fold runs keep their folds. folds='per_fold' (the default) draws every fold; folds='mean' draws the fold mean with a ±1 sd band and labels it mean of k folds ±sd; folds='both' draws both. A mean curve rendered as if it were a single run hides precisely the fold-to-fold variance k-fold was added to expose — and because folds can stop at different epochs (early stopping), the mean carries an n_folds column so a tail computed from two of five folds is visible rather than implied.

Best-epoch and last-epoch are both reported, always. The best epoch of a validation curve is chosen using that curve, so it is an optimistically biased estimate of held-out performance; the last epoch is unbiased but may be well past the optimum. Reporting either one alone is a way to be wrong.

The settings diff reuses the provenance bucketing. A flat key-by-key diff of two real runs reports ~200 differences, of which none are decisions anybody made. So this module reuses spacr.run_journal.values_equal() (structural comparison, so "[0, 1, 2]" from a CSV round-trip equals [0, 1, 2]) and the same three buckets:

changed

keys present in every run whose values differ — the knobs someone turned.

env

environment / non-substantive drift: paths, timestamps, host, versions, worker counts, verbosity. Shown, never mixed into changed.

drift

keys absent from at least one run — schema drift between spaCR releases, summarised rather than enumerated.

Two runs whose settings match say “no differences” in so many words; an empty table reads like a failure.

Broken run folders are reported, not fatal. A folder with no curves, no settings, or a header-only (zero-epoch) log is loaded anyway, carries a note saying what is wrong, and does not stop the other runs being compared.

This module does not import torch. It reads CSVs.

Example:

from spacr.train_compare import find_runs, compare_runs, format_comparison

runs = find_runs('/data/screen1/model')
cmp = compare_runs(runs, folds='both')
print(format_comparison(cmp, metric='accuracy'))
fig = plot_curves(cmp, 'accuracy')

See also

spacr.run_journal.diff_runs() — the two-run provenance diff whose bucketing this module reuses and generalises to N runs.

Classes

Comparison

The result of compare_runs() — series to plot plus the diff.

Series

One line on the plot: a run, a split and (for k-fold) a fold.

TrainingRun

One training run's curves, settings and complaints.

Functions

available_metrics(→ List[str])

Metric columns these runs logged, common ones first.

compare_runs() → Comparison)

Line up several runs' curves and diff their settings.

diff_settings() → Dict[str, Any])

Bucket the settings differences across N runs.

find_runs(→ List[TrainingRun])

Discover every training run under root, newest folder first.

format_comparison(→ str)

Render a Comparison as a console report.

is_env_key() → bool)

True when a settings key records the machine, not a modelling decision.

load_run(→ TrainingRun)

Load one training run folder into a TrainingRun.

metric_direction(→ Optional[str])

Return 'max', 'min' or None for a metric column name.

plot_curves(, mark_best, band)

Overlay every series' metric on one axes and return the figure.

render_setting_value(→ str)

One-line, length-capped rendering of a settings value.

Module Contents

class spacr.train_compare.Comparison[source]

The result of compare_runs() — series to plot plus the diff.

Parameters:
  • runs – source training runs in comparison order.

  • series – plot-ready series in run order.

  • settings_diff – bucketed settings comparison produced by diff_settings().

  • metrics – metric names available on at least one series, with shared metrics ordered first.

  • problems – flattened {'run_id', 'note'} records collected from all runs.

  • fold_mode – "per_fold", "mean", or "both", identifying how series was assembled.

epoch_ranges() → Dict[str, Tuple[int, int]][source]

Return each series label mapped to its finite epoch span.

labels() → List[str][source]

Return the plot-series labels in comparison order.

lengths_differ() → bool[source]

True when the series do not all cover the same epoch span.

series_for(label: str) → Series | None[source]

The series with this label, or None.

Parameters:

label – exact legend label to look up.

series_with(metric: str) → List[Series][source]

Series that actually have finite values for metric.

Parameters:

metric – metric each returned series must contain.

splits() → List[str][source]

Return the distinct split names in sorted order.

class spacr.train_compare.Series[source]

One line on the plot: a run, a split and (for k-fold) a fold.

Parameters:
  • run_id – identifier of the training run that produced this series.

  • split – logged data split, normally "train" or "val".

  • fold – fold name, an empty string for a single split, or "mean" for a fold aggregate.

  • kind – "single", "fold", or "mean", describing how the series was assembled.

  • label – complete legend label for the plotted line.

  • frame – epoch-ordered, unresampled frame containing epoch and this series’ numeric metrics.

  • n_folds – total folds represented by the series; one for an individual split or fold. For a mean, per-epoch support is retained separately in frame["n_folds"].

best(metric: str) → Dict[str, Any] | None[source]

{'epoch', 'value', 'direction'} for the best epoch, or None.

Parameters:

metric – metric whose direction-aware optimum is requested.

None when the metric is absent, entirely NaN, or has no meaningful direction (see metric_direction()).

epoch_range() → Tuple[int, int][source]

Return the minimum and maximum finite epochs, or (0, 0).

has(metric: str) → bool[source]

True when this series has at least one finite value for metric.

Parameters:

metric – metric column to check for finite observations.

last(metric: str) → Dict[str, Any] | None[source]

{'epoch', 'value'} for the last epoch with a finite value.

Parameters:

metric – metric whose last finite observation is requested.

sd(metric: str) → numpy.ndarray | None[source]

Fold-to-fold sd for a 'mean' series, else None.

Parameters:

metric – base metric whose __sd column is requested.

support() → Tuple[int, int] | None[source]

(min, max) folds contributing to a 'mean' series, else None.

Folds stop at different epochs when early stopping fires, so the tail of a fold mean can be an average over two folds where the head was an average over five. min < max is exactly that situation.

support_drops_at() → int | None[source]

First epoch where fewer than all folds contribute, or None.

values(metric: str) → numpy.ndarray[source]

This series’ values for metric (empty array when absent).

Parameters:

metric – metric column whose numeric values are requested.

property epochs: numpy.ndarray[source]

Return the epoch column coerced to a NumPy numeric array.

property n_epochs: int[source]

Return the number of rows in this series.

class spacr.train_compare.TrainingRun[source]

One training run’s curves, settings and complaints.

Parameters:
  • run_id – human-readable run identifier; find_runs() makes it unique within the returned scan.

  • path – training-output directory containing progress CSVs or numbered fold subdirectories.

  • settings – settings recovered from run-local settings.json or settings.csv, or from a matching ancestor settings/*.csv; empty when none is usable.

  • curves – long-form per-epoch metrics with run_id, split, fold, and epoch identity columns; empty with those columns present when no usable curves were logged.

  • folds – numerically ordered fold-directory names, empty for a non-cross-validated run.

  • final_metrics – per-series summaries containing identity, epoch count and range, last finite observations, and direction-aware best observations when defined.

  • notes – non-fatal discovery or data-quality messages; notes do not prevent the run from being compared.

  • settings_path – path of the settings snapshot used, or "" when none was found.

  • manifest – parsed run-journal manifest mapping, or an empty mapping when absent, unreadable, or not an object.

metrics() → List[str][source]

Numeric metric columns this run actually logged, sorted.

summary_line() → str[source]

One line for a list widget: id, shape of the log, key settings.

property has_curves: bool[source]

Return whether at least one usable curve row was loaded.

property is_cv: bool[source]

Return whether one or more fold directories were discovered.

property n_epochs: int[source]

Highest epoch number logged by any split/fold (0 when nothing was).

spacr.train_compare.available_metrics(runs: Sequence[TrainingRun], folds: str = 'per_fold') → List[str][source]

Metric columns these runs logged, common ones first.

Parameters:

runs – training runs whose plottable metrics are requested.

Lets a caller populate a metric picker before anything is compared.

spacr.train_compare.compare_runs(runs: Sequence[TrainingRun], folds: str = 'per_fold', env_keys: Sequence[str] = ()) → Comparison[source]

Line up several runs’ curves and diff their settings.

Parameters:
  • runs – runs from load_run() / find_runs(). Runs with no curves are kept — their notes travel into Comparison.problems — but contribute no series.

  • folds – how to render a cross-validated run: 'per_fold' (default; every fold gets a line), 'mean' (fold mean with a ±1 sd band, labelled as a mean) or 'both'.

  • env_keys – extra settings keys to treat as environment drift.

Returns:

a Comparison.

Raises:

ValueError – on an unknown folds mode.

spacr.train_compare.diff_settings(runs: Sequence[TrainingRun], env_keys: Sequence[str] = ()) → Dict[str, Any][source]

Bucket the settings differences across N runs.

Parameters:

runs – training runs whose settings are compared.

Generalises spacr.run_journal.diff_runs() from two runs to many and reuses its comparison (spacr.run_journal.values_equal()) and its reason for bucketing: a flat diff of two real runs reports ~200 differences of which none are decisions anybody made.

Returns:

a dict with

run_ids

the runs that contributed settings, in order.

changed

[{'key', 'values': {run_id: value}}] for keys present in every contributing run whose values differ — the signal.

env

same shape, for keys is_env_key() classifies as environment drift. Never merged into changed.

drift

[{'key', 'present': [...], 'missing': [...]}] — schema drift.

same / shared

counts of unchanged and shared keys.

identical

True when two or more runs contributed settings and nothing at all differs. Callers must say “no differences” rather than render an empty table.

no_settings

run ids that had no settings to contribute.

env_manifest

the manifest env diff, when the folders are run-journal runs.

spacr.train_compare.find_runs(root: Any, max_depth: int = DEFAULT_SCAN_DEPTH, limit: int = 200) → List[TrainingRun][source]

Discover every training run under root, newest folder first.

Walks at most max_depth levels, skips hidden folders, and never returns a fold_<i> folder found below root in its own right — folds belong to the run above them and are loaded as part of it. Pointing root straight at a fold folder still loads that one fold, because then it is what the caller asked for.

Runs that are broken (no curves, no settings, zero-epoch log) are returned with their notes rather than skipped: a scan that silently drops the folder you were looking for is worse than one that tells you what is wrong with it.

Parameters:
  • root – folder to scan (may itself be a run folder).

  • max_depth – how deep below root to look.

  • limit – stop after this many runs.

Returns:

loaded runs with ids made unique within the result.

spacr.train_compare.format_comparison(comparison: Comparison, metric: str = 'accuracy', max_drift_names: int = 6) → str[source]

Render a Comparison as a console report.

Parameters:

comparison – completed run comparison to render.

Ordering mirrors spacr.run_journal.format_run_diff(): the runs, then the curves, then the settings that changed (the signal), then environment drift, then schema drift reduced to one line. Problems come first, because a run with no curves in the list is the thing you most need to know.

Both the last-epoch and the best-epoch value of metric are printed for every series, with the caveat that makes the second one interpretable.

spacr.train_compare.is_env_key(key: Any, env_keys: Sequence[str] = ()) → bool[source]

True when a settings key records the machine, not a modelling decision.

Parameters:

key – settings key to classify.

Token-wise, not substring: start_time matches, update_freq does not. src deliberately does not match — a run on a different dataset is a real difference and belongs in changed, even though it is a path.

spacr.train_compare.load_run(path: Any, run_id: str | None = None) → TrainingRun[source]

Load one training run folder into a TrainingRun.

path is the folder spacr.deep_spacr.train_model() was given as dst — the one holding train.csv / validation.csv, or holding the fold_<i>/ subfolders of a cross-validated run.

A folder that is missing its curves, missing its settings or holding a zero-epoch log does not raise: the run comes back with whatever could be read and a TrainingRun.notes entry per problem, so one bad folder in a scan cannot stop the rest being compared.

Parameters:
  • path – run folder.

  • run_id – override the generated id.

Returns:

the loaded run.

Raises:

FileNotFoundError – only when path is not a directory.

spacr.train_compare.metric_direction(name: Any) → str | None[source]

Return 'max', 'min' or None for a metric column name.

Parameters:

name – metric column name whose optimisation direction is requested.

None means “no meaningful best” — optimal_threshold and train_time are recorded per epoch but neither has a direction, and calling the largest one “best” would be a fabrication.

spacr.train_compare.plot_curves(comparison: Comparison, metric: str = 'accuracy', ax: Any = None, labels: Sequence[str] | None = None, figsize: Tuple[float, float] = (9.0, 5.0), mark_best: bool = False, band: bool = True) → Any[source]

Overlay every series’ metric on one axes and return the figure.

Each series is drawn over its own epochs. Runs of different lengths are not truncated to the shortest or padded to the longest: the lines end where the runs ended, and the axes annotation says so.

Colour identifies the run, line style the split (train dashed, validation solid), and every legend entry spells out run · split · fold. When more than one split is on the axes a note says so — a train curve sitting above a validation curve is overfitting, not a better run.

'mean' series (from compare_runs(folds='mean')) are drawn with a ±1 sd band; where fewer folds reach an epoch than the run has, the band is thinner because fewer folds contributed, which the legend states.

Parameters:
  • comparison – from compare_runs().

  • metric – metric column to draw.

  • ax – draw into this axes instead of making a figure (the Qt screen reuses one canvas).

  • labels – only draw these series labels.

  • figsize – figure size when ax is None.

  • mark_best – put a marker on each series’ best epoch.

  • band – draw the ±sd band for mean series.

Returns:

the matplotlib.figure.Figure. It carries spacr_series_by_label — {line label: Series} — so a click on a line can name its run.

spacr.train_compare.render_setting_value(value: Any, width: int = 40) → str[source]

One-line, length-capped rendering of a settings value.

Parameters:

value – settings value to render.

Thin public wrapper over spacr.run_journal._render_value() so the GUI renders settings exactly the way the console report does.