spacr.train_compare¶
Workflow inputs and outputs¶
Training Runs¶
Compare saved training histories, metrics and settings from compatible model runs.
Open: Classify → Training Runs.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Run history and artifacts — Project run records, settings, output paths, artifact provenance, status and logs.
Outputs
Run/model comparison — Comparison tables and figures from compatible saved runs or masks; agreement is not ground-truth accuracy.
Overlay several training runs’ curves and say exactly what differed between them.
Why this exists¶
Picking a classifier in spaCR means training it, writing down the number,
changing a knob, training it again, and then trying to remember which of the
six things you touched is the one that moved the number. The provenance diff
(spacr.run_journal.diff_runs()) answers “what changed”; this module
answers “and what did it do to the curves”, on one axis, next to the diff.
What is on disk, and therefore what can be shown¶
spacr.deep_spacr.train_model() evaluates train (and validation, when a
validation loader exists) at the end of every epoch and hands the metric dicts
to spacr.io._save_progress(), which appends them to two CSVs in the run’s
dst folder:
<src>/model/<model_type>/<channels>/epochs_<N>/
train.csv # one row per epoch
validation.csv # one row per epoch, only when val_loaders exist
<model_type>_epoch_<e>_channels_<ch>.pth
Columns come from spacr.deep_spacr.evaluate_model_performance():
epoch, loss, accuracy (and its duplicate Accuracy),
neg_accuracy, pos_accuracy, prauc, optimal_threshold, plus
train_time on the train side and num_classes on the multiclass side.
A k-fold run (cross_validation_folds >= 2) trains a fresh model per fold in
<dst>/fold_<i>/, so it produces k independent pairs of curves, plus the
per-fold summary CSVs spacr.deep_spacr._cross_validate_model() writes at
the dst level.
The settings that produced a run are not in the run folder: save_settings
writes them to <src>/settings/train_test_<model_type>_<epochs>.csv as a
Key,Value CSV. load_run() walks up from the run folder to find it —
but only as far as the run’s own project root, which the layout above makes
identifiable: <src> is the folder the model/ tree hangs off. An
ancestor above that belongs to a different project, and reading it would report
somebody else’s settings as this run’s provenance, so the settings diff would
show differences nobody configured and the answer would depend on where in the
filesystem the project happened to sit. See _owns_this_run().
Judgement calls baked in¶
Runs of different length are never aligned. Each series is plotted over its own epoch axis. Truncating a 60-epoch run to match a 25-epoch neighbour invents a result (“B never got better”) and padding invents a different one; both are worse than a legend that says one run is longer.
Train and validation are never silently mixed. Every series label carries
run, split and fold, and plot_curves() annotates the axes when more than
one split is on it. A train curve above a validation curve is the definition of
overfitting, not evidence that one run beat another.
k-fold runs keep their folds. folds='per_fold' (the default) draws every
fold; folds='mean' draws the fold mean with a ±1 sd band and labels it
mean of k folds ±sd; folds='both' draws both. A mean curve rendered as
if it were a single run hides precisely the fold-to-fold variance k-fold was
added to expose — and because folds can stop at different epochs (early
stopping), the mean carries an n_folds column so a tail computed from two of
five folds is visible rather than implied.
Best-epoch and last-epoch are both reported, always. The best epoch of a validation curve is chosen using that curve, so it is an optimistically biased estimate of held-out performance; the last epoch is unbiased but may be well past the optimum. Reporting either one alone is a way to be wrong.
The settings diff reuses the provenance bucketing. A flat key-by-key diff of
two real runs reports ~200 differences, of which none are decisions anybody
made. So this module reuses spacr.run_journal.values_equal() (structural
comparison, so "[0, 1, 2]" from a CSV round-trip equals [0, 1, 2]) and
the same three buckets:
changedkeys present in every run whose values differ — the knobs someone turned.
envenvironment / non-substantive drift: paths, timestamps, host, versions, worker counts, verbosity. Shown, never mixed into
changed.driftkeys absent from at least one run — schema drift between spaCR releases, summarised rather than enumerated.
Two runs whose settings match say “no differences” in so many words; an empty table reads like a failure.
Broken run folders are reported, not fatal. A folder with no curves, no settings, or a header-only (zero-epoch) log is loaded anyway, carries a note saying what is wrong, and does not stop the other runs being compared.
This module does not import torch. It reads CSVs.
Example:
from spacr.train_compare import find_runs, compare_runs, format_comparison
runs = find_runs('/data/screen1/model')
cmp = compare_runs(runs, folds='both')
print(format_comparison(cmp, metric='accuracy'))
fig = plot_curves(cmp, 'accuracy')
See also
spacr.run_journal.diff_runs() — the two-run provenance diff whose
bucketing this module reuses and generalises to N runs.
Classes¶
The result of |
|
One line on the plot: a run, a split and (for k-fold) a fold. |
|
One training run's curves, settings and complaints. |
Functions¶
|
Metric columns these runs logged, common ones first. |
|
Line up several runs' curves and diff their settings. |
|
Bucket the settings differences across N runs. |
|
Discover every training run under |
|
Render a |
|
True when a settings key records the machine, not a modelling decision. |
|
Load one training run folder into a |
|
Return |
|
Overlay every series' |
|
One-line, length-capped rendering of a settings value. |
Module Contents¶
- class spacr.train_compare.Comparison[source]¶
The result of
compare_runs()— series to plot plus the diff.- Parameters:
runs – source training runs in comparison order.
series – plot-ready series in run order.
settings_diff – bucketed settings comparison produced by
diff_settings().metrics – metric names available on at least one series, with shared metrics ordered first.
problems – flattened
{'run_id', 'note'}records collected from all runs.fold_mode –
"per_fold","mean", or"both", identifying howserieswas assembled.
- epoch_ranges() Dict[str, Tuple[int, int]][source]¶
Return each series label mapped to its finite epoch span.
- series_for(label: str) Series | None[source]¶
The series with this label, or
None.- Parameters:
label – exact legend label to look up.
- class spacr.train_compare.Series[source]¶
One line on the plot: a run, a split and (for k-fold) a fold.
- Parameters:
run_id – identifier of the training run that produced this series.
split – logged data split, normally
"train"or"val".fold – fold name, an empty string for a single split, or
"mean"for a fold aggregate.kind –
"single","fold", or"mean", describing how the series was assembled.label – complete legend label for the plotted line.
frame – epoch-ordered, unresampled frame containing
epochand this series’ numeric metrics.n_folds – total folds represented by the series; one for an individual split or fold. For a mean, per-epoch support is retained separately in
frame["n_folds"].
- best(metric: str) Dict[str, Any] | None[source]¶
{'epoch', 'value', 'direction'}for the best epoch, orNone.- Parameters:
metric – metric whose direction-aware optimum is requested.
Nonewhen the metric is absent, entirely NaN, or has no meaningful direction (seemetric_direction()).
- has(metric: str) bool[source]¶
True when this series has at least one finite value for
metric.- Parameters:
metric – metric column to check for finite observations.
- last(metric: str) Dict[str, Any] | None[source]¶
{'epoch', 'value'}for the last epoch with a finite value.- Parameters:
metric – metric whose last finite observation is requested.
- sd(metric: str) numpy.ndarray | None[source]¶
Fold-to-fold sd for a
'mean'series, elseNone.- Parameters:
metric – base metric whose
__sdcolumn is requested.
- support() Tuple[int, int] | None[source]¶
(min, max)folds contributing to a'mean'series, else None.Folds stop at different epochs when early stopping fires, so the tail of a fold mean can be an average over two folds where the head was an average over five.
min < maxis exactly that situation.
- values(metric: str) numpy.ndarray[source]¶
This series’ values for
metric(empty array when absent).- Parameters:
metric – metric column whose numeric values are requested.
- property epochs: numpy.ndarray[source]¶
Return the epoch column coerced to a NumPy numeric array.
- class spacr.train_compare.TrainingRun[source]¶
One training run’s curves, settings and complaints.
- Parameters:
run_id – human-readable run identifier;
find_runs()makes it unique within the returned scan.path – training-output directory containing progress CSVs or numbered fold subdirectories.
settings – settings recovered from run-local
settings.jsonorsettings.csv, or from a matching ancestorsettings/*.csv; empty when none is usable.curves – long-form per-epoch metrics with
run_id,split,fold, andepochidentity columns; empty with those columns present when no usable curves were logged.folds – numerically ordered fold-directory names, empty for a non-cross-validated run.
final_metrics – per-series summaries containing identity, epoch count and range, last finite observations, and direction-aware best observations when defined.
notes – non-fatal discovery or data-quality messages; notes do not prevent the run from being compared.
settings_path – path of the settings snapshot used, or
""when none was found.manifest – parsed run-journal manifest mapping, or an empty mapping when absent, unreadable, or not an object.
- spacr.train_compare.available_metrics(runs: Sequence[TrainingRun], folds: str = 'per_fold') List[str][source]¶
Metric columns these runs logged, common ones first.
- Parameters:
runs – training runs whose plottable metrics are requested.
Lets a caller populate a metric picker before anything is compared.
- spacr.train_compare.compare_runs(runs: Sequence[TrainingRun], folds: str = 'per_fold', env_keys: Sequence[str] = ()) Comparison[source]¶
Line up several runs’ curves and diff their settings.
- Parameters:
runs – runs from
load_run()/find_runs(). Runs with no curves are kept — their notes travel intoComparison.problems— but contribute no series.folds – how to render a cross-validated run:
'per_fold'(default; every fold gets a line),'mean'(fold mean with a ±1 sd band, labelled as a mean) or'both'.env_keys – extra settings keys to treat as environment drift.
- Returns:
a
Comparison.- Raises:
ValueError – on an unknown
foldsmode.
- spacr.train_compare.diff_settings(runs: Sequence[TrainingRun], env_keys: Sequence[str] = ()) Dict[str, Any][source]¶
Bucket the settings differences across N runs.
- Parameters:
runs – training runs whose settings are compared.
Generalises
spacr.run_journal.diff_runs()from two runs to many and reuses its comparison (spacr.run_journal.values_equal()) and its reason for bucketing: a flat diff of two real runs reports ~200 differences of which none are decisions anybody made.- Returns:
a dict with
run_idsthe runs that contributed settings, in order.
changed[{'key', 'values': {run_id: value}}]for keys present in every contributing run whose values differ — the signal.envsame shape, for keys
is_env_key()classifies as environment drift. Never merged intochanged.drift[{'key', 'present': [...], 'missing': [...]}]— schema drift.same/sharedcounts of unchanged and shared keys.
identicalTrue when two or more runs contributed settings and nothing at all differs. Callers must say “no differences” rather than render an empty table.
no_settingsrun ids that had no settings to contribute.
env_manifestthe manifest
envdiff, when the folders are run-journal runs.
- spacr.train_compare.find_runs(root: Any, max_depth: int = DEFAULT_SCAN_DEPTH, limit: int = 200) List[TrainingRun][source]¶
Discover every training run under
root, newest folder first.Walks at most
max_depthlevels, skips hidden folders, and never returns afold_<i>folder found belowrootin its own right — folds belong to the run above them and are loaded as part of it. Pointingrootstraight at a fold folder still loads that one fold, because then it is what the caller asked for.Runs that are broken (no curves, no settings, zero-epoch log) are returned with their notes rather than skipped: a scan that silently drops the folder you were looking for is worse than one that tells you what is wrong with it.
- Parameters:
root – folder to scan (may itself be a run folder).
max_depth – how deep below
rootto look.limit – stop after this many runs.
- Returns:
loaded runs with ids made unique within the result.
- spacr.train_compare.format_comparison(comparison: Comparison, metric: str = 'accuracy', max_drift_names: int = 6) str[source]¶
Render a
Comparisonas a console report.- Parameters:
comparison – completed run comparison to render.
Ordering mirrors
spacr.run_journal.format_run_diff(): the runs, then the curves, then the settings that changed (the signal), then environment drift, then schema drift reduced to one line. Problems come first, because a run with no curves in the list is the thing you most need to know.Both the last-epoch and the best-epoch value of
metricare printed for every series, with the caveat that makes the second one interpretable.
- spacr.train_compare.is_env_key(key: Any, env_keys: Sequence[str] = ()) bool[source]¶
True when a settings key records the machine, not a modelling decision.
- Parameters:
key – settings key to classify.
Token-wise, not substring:
start_timematches,update_freqdoes not.srcdeliberately does not match — a run on a different dataset is a real difference and belongs inchanged, even though it is a path.
- spacr.train_compare.load_run(path: Any, run_id: str | None = None) TrainingRun[source]¶
Load one training run folder into a
TrainingRun.pathis the folderspacr.deep_spacr.train_model()was given asdst— the one holdingtrain.csv/validation.csv, or holding thefold_<i>/subfolders of a cross-validated run.A folder that is missing its curves, missing its settings or holding a zero-epoch log does not raise: the run comes back with whatever could be read and a
TrainingRun.notesentry per problem, so one bad folder in a scan cannot stop the rest being compared.- Parameters:
path – run folder.
run_id – override the generated id.
- Returns:
the loaded run.
- Raises:
FileNotFoundError – only when
pathis not a directory.
- spacr.train_compare.metric_direction(name: Any) str | None[source]¶
Return
'max','min'orNonefor a metric column name.- Parameters:
name – metric column name whose optimisation direction is requested.
Nonemeans “no meaningful best” —optimal_thresholdandtrain_timeare recorded per epoch but neither has a direction, and calling the largest one “best” would be a fabrication.
- spacr.train_compare.plot_curves(comparison: Comparison, metric: str = 'accuracy', ax: Any = None, labels: Sequence[str] | None = None, figsize: Tuple[float, float] = (9.0, 5.0), mark_best: bool = False, band: bool = True) Any[source]¶
Overlay every series’
metricon one axes and return the figure.Each series is drawn over its own epochs. Runs of different lengths are not truncated to the shortest or padded to the longest: the lines end where the runs ended, and the axes annotation says so.
Colour identifies the run, line style the split (train dashed, validation solid), and every legend entry spells out run · split · fold. When more than one split is on the axes a note says so — a train curve sitting above a validation curve is overfitting, not a better run.
'mean'series (fromcompare_runs(folds='mean')) are drawn with a ±1 sd band; where fewer folds reach an epoch than the run has, the band is thinner because fewer folds contributed, which the legend states.- Parameters:
comparison – from
compare_runs().metric – metric column to draw.
ax – draw into this axes instead of making a figure (the Qt screen reuses one canvas).
labels – only draw these series labels.
figsize – figure size when
axis None.mark_best – put a marker on each series’ best epoch.
band – draw the ±sd band for mean series.
- Returns:
the
matplotlib.figure.Figure. It carriesspacr_series_by_label—{line label: Series}— so a click on a line can name its run.
- spacr.train_compare.render_setting_value(value: Any, width: int = 40) str[source]¶
One-line, length-capped rendering of a settings value.
- Parameters:
value – settings value to render.
Thin public wrapper over
spacr.run_journal._render_value()so the GUI renders settings exactly the way the console report does.