spacr.hyperparam

Hyperparameter search — grids, random sweeps, UMAP criteria and grouped CV.

This module generalises spacr.core.reducer_hyperparameter_search(), which already swept UMAP/tSNE + DBSCAN/KMeans parameters and drew the resulting embeddings as a grid of small multiples. That function stays where it is (it still owns the database read + the image-glyph plotting); what lives here is the search machinery it lacked: recorded trials, failed trials that do not abort the sweep, an incremental progress callback, early stopping, reproducible sampling, and — most importantly — an explicit, named scoring criterion instead of an unscored eyeball grid.

Three things this module refuses to pretend:

UMAP has no ground truth. There is no measurement that says one embedding is correct and another is wrong. Every criterion here (trustworthiness, continuity, silhouette) rewards a different property, and they routinely disagree about which n_neighbors/min_dist wins. The scores are an aid for ranking a panel of embeddings you then look at; they are not a verdict. Anything that prints “best embedding” without naming the criterion is misleading, so format_search() always names it and always prints the caveat.

Selecting on the test split leaks. cv_search() scores every trial on cross-validation folds, never on test, and refuses to run if a caller hands it folds that touch the held-out test indices. It defaults to grouped folds (group_by='well') and reuses spacr.io.make_cv_folds() — crops from one well share focus, illumination and seeding density, so an ungrouped search picks the model that memorised wells and reports a beautiful, meaningless score.

A winner without a spread is a lie. When the top ten configurations sit inside the fold-to-fold standard deviation, the hyperparameter did not matter and the “winner” is noise. Every SearchResult reports the spread and raises a within_noise flag when that is what happened.

author:

spaCR

Classes

ActivationSearchData

The model and images one Activation sweep is scored on.

SearchData

Feature matrix (and, where applicable, labels and groups) for one search.

SearchResult

Everything one sweep produced, including what it failed to produce.

SearchSpace

A named set of parameters, each with the list of values to try.

Trial

One evaluated configuration, successful or not.

WalkAxis

One searchable direction in hyperparameter space.

Functions

activation_fit_fn(→ Callable[[Dict[str, Any]], Any])

Build the fit_fn(params) an Activation sweep evaluates.

activation_search(→ SearchResult)

Sweep attribution settings, scoring each with the faithfulness checks.

build_folds(→ Tuple[List[Tuple[Any, Any]], List[str]])

Build grouped cross-validation folds over the non-test samples.

build_sklearn_model(model_type, params[, seed, n_jobs])

Construct the classical-ML classifier model_type names.

classify_cv_fit_fn(settings, *[, criterion, n_folds, ...])

Build a fit_fn(params) that trains one deep model per fold.

cv_search(→ SearchResult)

Search hyperparameters by cross-validation, never by scoring on test.

embedding_stability(→ float)

Measure repeat-to-repeat preservation of embedding neighbours.

format_search(→ str)

Render a search result as plain text, caveats first.

grid_search(→ SearchResult)

Evaluate every configuration in space.

load_activation_data(→ ActivationSearchData)

Load the model and a handful of images an Activation sweep scores on.

load_search_data(→ SearchData)

Load the feature matrix a search needs, straight from the measurements DB.

local_direction_search(→ SearchResult)

Walk UMAP's n_neighbors x min_dist plane by 2-by-2 rounds.

random_search(→ SearchResult)

Evaluate n_trials configurations drawn at random from space.

run_search_for_app(, umap_components, seed, n_folds, ...)

Run the right search for a spaCR app. This is what the GUI calls.

sklearn_cv_fit_fn(features, labels[, model_type, ...])

Build the fit_fn(params, train_idx, val_idx) cv_search() wants.

umap_available(→ Tuple[bool, str])

Whether umap-learn can be imported.

umap_checkpoint_path(→ Optional[str])

Return the default UMAP-search checkpoint path for module settings.

umap_metrics(→ Tuple[str, ...])

Every metric the INSTALLED umap-learn will accept.

umap_objective_scores(→ Dict[str, Any])

Score neighborhood preservation, stability and cluster structure.

umap_search(, keep_embeddings, on_trial, int, int], ...)

Sweep UMAP parameters, scoring each embedding with a named criterion.

umap_walk_axes(→ List[WalkAxis])

Build Walk axes for UMAP from a starting configuration.

walk_neighbourhood(→ Tuple[List[Dict[str, Any]], bool])

The configurations one Walk round evaluates around centre.

walk_search(→ SearchResult)

Walk N-dimensional hyperparameter space toward a better score.

Module Contents

class spacr.hyperparam.ActivationSearchData[source]

The model and images one Activation sweep is scored on.

Variables:
  • model – the trained classifier, already on the right device and in eval mode.

  • images – list of per-image tensors (C, H, W).

  • masks – optional per-image boolean object masks, same spatial shape, enabling the pointing game. None when spaCR could not find them.

  • filenames – per-image provenance for the panel labels.

  • model_type – architecture name, used to make errors readable.

  • notes – provenance and warnings the caller must surface.

class spacr.hyperparam.SearchData[source]

Feature matrix (and, where applicable, labels and groups) for one search.

Variables:
  • features – 2-D numeric matrix, one row per object.

  • labels – per-row class labels, or None for an unsupervised search.

  • groups – per-row group ids used to keep folds honest, or None.

  • frame – the joined measurement table the matrix came from.

  • notes – provenance and warnings the caller must surface.

class spacr.hyperparam.SearchResult[source]

Everything one sweep produced, including what it failed to produce.

Variables:
  • trials – every trial attempted, in deterministic order, failures included.

  • best – the highest- (or lowest-) scoring successful trial, or None when nothing succeeded.

  • space – the space that was searched.

  • metric – the name of the criterion score holds.

  • notes – caveats, warnings and provenance the caller must surface.

  • partial – True when the sweep stopped before evaluating everything it was asked to. A partial sweep must never be presented as a finished one.

  • higher_is_better – direction of metric.

  • objectives – objective name to direction mapping. When populated, pareto_front() exposes non-dominated configurations.

as_rows() → List[Dict[str, Any]][source]

Flat, table-ready rendering of every trial, best-first then failures.

noise_level() → Tuple[float | None, str][source]

The yardstick used to decide whether the winner is real.

Prefers the best trial’s own fold-to-fold standard deviation, because that is the run-to-run variation of a single configuration. Falls back to the spread across trials when no fold information exists.

Returns:

(value, source_description); value is None when there is not enough information.

pareto_front() → List[Trial][source]

Return non-dominated successful trials for declared objectives.

A trial is dominated when another is at least as good on every objective and strictly better on one. The returned order follows the composite ranking so the GUI remains deterministic.

ranked() → List[Trial][source]

Successful trials best-first; ties broken by trial order.

score_stats() → Dict[str, float | None][source]

Summary statistics over the successful trials’ scores.

Returns:

dict with n, best, worst, mean, std and spread (max - min). Every value except n is None when no trial succeeded.

trials_within_noise() → List[Trial][source]

Successful trials whose score is indistinguishable from the best.

within_noise(top_n: int = 3) → bool[source]

True when the top top_n trials are within the noise level.

When this fires, the hyperparameter did not measurably matter over the range searched, and reporting a single winner hides that.

Parameters:

top_n – how many of the leading trials to compare.

Returns:

True when the leaders are statistically indistinguishable.

property failed: List[Trial][source]

Trials that raised or returned an unusable score.

property n_failed: int[source]

How many trials failed.

property ok: bool[source]

True when at least one trial produced a score.

property successful: List[Trial][source]

Trials that produced a score.

class spacr.hyperparam.SearchSpace[source]

A named set of parameters, each with the list of values to try.

Values are stored as tuples so a space cannot be mutated after the sweep that used it has been reported.

Parameters:

params – mapping of parameter name to the list of values to try.

Raises:

ValueError – if the space is empty, a name is not a string, a parameter’s values are not a list/tuple, or a parameter has no values.

__post_init__() → None[source]

Validate and freeze the parameter mapping into tuples.

describe() → str[source]

One-line human summary of the space and its size.

grid() → List[Dict[str, Any]][source]

Enumerate the full Cartesian product in a deterministic order.

Parameter names are sorted; within a name, values keep the order the caller gave them; the last name varies fastest.

Returns:

list of parameter dicts, one per configuration.

sample(rng: random.Random) → Dict[str, Any][source]

Draw one configuration uniformly at random.

Values are drawn in sorted-name order so a fixed seed always yields the same sequence regardless of the caller’s dict insertion order.

Parameters:

rng – seeded random.Random instance.

Returns:

a parameter dict.

size() → int[source]

Number of configurations in the full Cartesian product.

property is_single_point: bool[source]

True when the space contains exactly one configuration.

property names: Tuple[str, ...][source]

Parameter names in sorted order — this fixes the grid’s column order.

class spacr.hyperparam.Trial[source]

One evaluated configuration, successful or not.

Parameters:
  • params – parameter configuration that was evaluated.

  • score – primary metric value, or None when no usable score was produced.

  • extra_metrics – additional fit outputs such as fold scores, alternate criteria, embeddings, or runtime counters.

  • duration – wall-clock seconds spent evaluating the trial.

  • error – failure message, or None when evaluation succeeded.

  • index – position in deterministic trial order; -1 means an order has not yet been assigned.

label() → str[source]

Compact k=v, k=v rendering of the configuration.

property ok: bool[source]

True when the trial produced a usable score.

class spacr.hyperparam.WalkAxis[source]

One searchable direction in hyperparameter space.

An axis is either NUMERIC – it has a step and the walk moves along it by multiples of that step – or CATEGORICAL, where choices lists the values and there is no direction to move in, only other values to try.

Parameters:
  • name – the parameter name, as the fit function expects it.

  • step – numeric axes only: how far one move goes.

  • minimum – numeric axes only: inclusive lower clamp, or None.

  • maximum – numeric axes only: inclusive upper clamp, or None.

  • integer – numeric axes only: round candidates to whole numbers.

  • choices – categorical axes only: the permitted values, in the order the walk should try them.

  • resolution – how many values this axis contributes to one round, counting the centre. 2 is the classic ±step pair with no centre; 3 adds the centre back; 5 reaches two steps out. The old 2-by-2 search is exactly two numeric axes at resolution 2, which is why that number is the default and not a special case in the code.

__post_init__() → None[source]

Normalize the axis and reject unusable names, choices, or ranges.

clamp(value: Any) → Any[source]

Bring value inside this axis’s declared range.

Categorical axes clamp by membership: a value that is not a choice becomes the first choice, because a walk that steps outside its own alphabet has nowhere to come back from.

values_around(centre: Any) → List[Any][source]

The values this axis offers for one round, centred on centre.

Clamping happens here, so an axis at its boundary offers fewer distinct values rather than duplicates of the edge – which is what lets the walk keep moving along the axes that still have room.

property categorical: bool[source]

Whether this axis has choices rather than a step.

spacr.hyperparam.activation_fit_fn(data: ActivationSearchData, *, criterion: str = 'deletion_auc', n_steps: int = 12, baseline: str = 'blur', sanity_threshold: float = 0.5, run_sanity_check: bool = True, keep_maps: bool = True, attribute_fn: Callable[..., Any] | None = None) → Callable[[Dict[str, Any]], Any][source]

Build the fit_fn(params) an Activation sweep evaluates.

Every trial attributes each image once and then measures that map four ways — deletion AUC, insertion AUC, the pointing game (when masks exist) and the randomisation sanity check — so the table can be re-ranked by any criterion without re-running the sweep. All four are reported for every trial precisely because they disagree; a sweep that reported only the one it ranked by would hide the disagreement, which is the informative part.

Parameters:
  • data – the model and images to score on.

  • criterion – which of ACTIVATION_CRITERIA drives the ranking.

  • n_steps – perturbation steps in the deletion / insertion curves.

  • baseline – what removed pixels become — 'blur' (least out-of-distribution), 'zero', 'mean' or 'uniform'.

  • sanity_threshold – rank correlation below which a method passes the randomisation check.

  • run_sanity_check – run the check on the first image only. It costs one extra attribution per parameterised layer, so it is the expensive part of a trial; turning it off removes the most valuable number here.

  • keep_maps – keep each trial’s first map so the panel can draw it.

  • attribute_fn – override for the attribution call, used by tests.

Returns:

the fit function.

Raises:

ValueError – for an unknown criterion or an empty image set.

Sweep attribution settings, scoring each with the faithfulness checks.

The honest deliverable is the panel of maps plus the four scores per trial, not the top row. There is no ground truth for attribution, so this refuses to name a single “best”: ACTIVATION_NO_GROUND_TRUTH leads the notes, every criterion is computed for every trial, and the usual within-noise flag fires when the leaders are indistinguishable.

Parameters:
  • data – model + images (+ optional masks) to score on.

  • space – attribution parameters to sweep (cam_type, target_layer, smoothgrad_samples, smoothgrad_sigma, occlusion_window, occlusion_stride, ig_steps, ig_baseline).

  • criterion – which criterion ranks the trials.

  • mode – 'grid' or 'random'.

  • n_trials – configurations when mode='random'; maximum complete 2-by-2 rounds when adaptive UMAP is enabled.

  • seed – seed for random sampling.

  • n_steps – steps in the deletion / insertion curves.

  • baseline – removal baseline for those curves.

  • run_sanity_check – run the randomisation check per trial.

  • attribute_fn – override for the attribution call, used by tests.

  • on_trial – progress callback (trial, completed, total).

  • should_stop – polled before each trial.

Returns:

the SearchResult.

Raises:

ValueError – for an unknown criterion or mode.

spacr.hyperparam.build_folds(labels, n_folds: int = 5, *, groups=None, filenames: Sequence[str] | None = None, group_by: str = 'well', seed: int = 0, exclude=None) → Tuple[List[Tuple[Any, Any]], List[str]][source]

Build grouped cross-validation folds over the non-test samples.

Reuses spacr.io.make_cv_folds(), the same fold builder cross_validation_folds drives during training, so a search and the run it configures split the data the same way. Grouping defaults to 'well' because crops from one well share focus, illumination, seeding density and edge effects; letting them straddle a split lets a model recognise the well instead of the phenotype.

Parameters:
  • labels – per-sample class labels.

  • n_folds – number of folds; must be at least 2.

  • groups – explicit per-sample group ids. Takes precedence over filenames.

  • filenames – crop filenames from which group ids are parsed when groups is not given.

  • group_by – grouping level — 'cell', 'field', 'well' (default), or 'plate'. Legacy 'none' aliases 'cell'.

  • seed – RNG seed, so the folds reproduce.

  • exclude – indices to keep out of every fold — the held-out test split.

Returns:

(folds, warnings) where folds is a list of (train_idx, val_idx) index arrays into the original sample order.

Raises:

ValueError – when n_folds < 2 or every sample was excluded.

spacr.hyperparam.build_sklearn_model(model_type: str, params: Mapping[str, Any], seed: int = 42, n_jobs: int = -1)[source]

Construct the classical-ML classifier model_type names.

Mirrors the constructors in spacr.ml.ml_analysis() (ml.py, the model_type == ladder) so a search configures the same estimator the real run will fit. Unknown keyword arguments in params are dropped with a clear error rather than silently ignored.

Parameters:
  • model_type – one of the model_type_ml combo values.

  • params – hyperparameters for this trial (n_estimators, learning_rate, reg_alpha, reg_lambda, …).

  • seed – random_state.

  • n_jobs – worker count where the estimator supports it.

Returns:

an unfitted scikit-learn-compatible classifier.

Raises:
  • ValueError – for an unsupported model_type.

  • ImportError – with an install hint for optional backends.

spacr.hyperparam.classify_cv_fit_fn(settings: Mapping[str, Any], *, criterion: str = 'accuracy', n_folds: int = 5, train_fn: Callable[[Dict[str, Any]], Any] | None = None, read_fold_csv: Callable[[str], Any] | None = None)[source]

Build a fit_fn(params) that trains one deep model per fold.

The Classify (CV) app trains Torch CNNs, so the search does not roll its own folds: it hands each configuration to spacr.deep_spacr. train_test_model() with cross_validation_folds forced to at least two and cv_group_by left alone, then reads the per-fold CSV that run writes. That keeps a single implementation of grouped k-fold — spaCR’s — and means the search splits the data exactly the way the training run will.

Each trial trains n_folds models, so a grid of g configurations trains g × n_folds models. That is the honest cost; there is no cheap proxy for it.

Parameters:
  • settings – base Classify (CV) settings dict; each trial’s parameters are layered on top.

  • criterion – metric column to read from the per-fold CSV.

  • n_folds – cross-validation folds per trial; forced to at least 2.

  • train_fn – override for train_test_model (used by tests so no real CNN is trained).

  • read_fold_csv – override for the CSV reader.

Returns:

the fit function.

Raises:

ValueError – when n_folds < 2.

Search hyperparameters by cross-validation, never by scoring on test.

Every configuration is fitted once per fold on that fold’s training indices and scored on that fold’s validation indices; the trial score is the mean across folds and the fold-to-fold standard deviation becomes the noise yardstick for SearchResult.within_noise(). Indices listed in test_idx are removed before the folds are built and are never handed to fit_fn — selecting a configuration on data reserved for the final estimate makes that estimate meaningless.

Parameters:
  • fit_fn – called as fit_fn(params, train_idx, val_idx); returns the validation score, a (score, metrics) pair, or a dict with score.

  • space – the SearchSpace to search.

  • labels – per-sample class labels (used to stratify the folds).

  • groups – explicit per-sample group ids.

  • filenames – crop filenames to parse group ids from.

  • group_by – grouping level, 'well' by default.

  • n_folds – number of cross-validation folds.

  • seed – seed for the folds and for random sampling.

  • test_idx – indices of the held-out test split, excluded from every fold.

  • folds – pre-built (train_idx, val_idx) pairs, bypassing build_folds(). Validated against test_idx.

  • metric – name of the score fit_fn returns.

  • higher_is_better – direction of metric.

  • n_trials – when given, sample this many configurations at random instead of running the full grid.

  • on_trial – progress callback (trial, completed, total).

  • should_stop – polled before each trial.

Returns:

the SearchResult.

Raises:

ValueError – when supplied folds touch test_idx.

spacr.hyperparam.embedding_stability(embeddings: Sequence[Any], *, neighbourhood_k: int = 15) → float[source]

Measure repeat-to-repeat preservation of embedding neighbours.

Rotation, reflection and axis scaling do not affect this measure: for every pair of embeddings it finds each sample’s k nearest neighbours and averages the fraction shared by both fits.

Parameters:
  • embeddings – two or more aligned (n_samples, n_components) embeddings of the same rows.

  • neighbourhood_k – number of neighbours compared per row.

Returns:

mean shared-neighbour fraction in [0, 1].

Render a search result as plain text, caveats first.

The layout is deliberate: the criterion and its caveat come before the table, the spread and the within-noise flag come immediately after the winner, and a partial or all-failed sweep says so in the header rather than in a footnote nobody reads.

Parameters:
  • result – the SearchResult to render.

  • max_rows – how many ranked trials to print.

Returns:

the report as a string.

Evaluate every configuration in space.

Parameters:
  • fit_fn – called as fit_fn(params); returns a score, a (score, metrics) pair, or a dict with a score key.

  • space – the SearchSpace to enumerate.

  • metric – name of the criterion the score represents.

  • higher_is_better – direction of metric.

  • on_trial – progress callback (trial, completed, total).

  • should_stop – polled before each trial; True truncates the sweep and marks the result partial.

  • notes – extra caveats to attach.

Returns:

the SearchResult.

spacr.hyperparam.load_activation_data(settings: Mapping[str, Any], *, n_images: int = 8) → ActivationSearchData[source]

Load the model and a handful of images an Activation sweep scores on.

Two sources, in order of preference:

  • src/merged/*.npy — spaCR’s own merged arrays, which carry the image channels and the object label planes in one file. Preferred because the object mask comes free and exactly aligned, which is what makes the pointing game possible at all.

  • dataset — the crop tar the Activation run itself reads. Aligned masks do not exist for these crops, so the pointing game is unavailable and the returned notes say so rather than silently dropping the criterion.

A sweep runs every configuration over every image, so n_images is small on purpose: the cost is configurations × images × (2 curves + 1 sanity cascade) forward passes.

Parameters:
  • settings – the Activation app’s settings dict.

  • n_images – how many images to score on.

Returns:

the ActivationSearchData.

Raises:

ValueError – when neither source is usable.

spacr.hyperparam.load_search_data(app_key: str, settings: Mapping[str, Any]) → SearchData[source]

Load the feature matrix a search needs, straight from the measurements DB.

This is the same read + preprocess path spacr.core.reducer_hyperparameter_search() uses (get_db_paths → _read_and_join_tables → preprocess_data), so a search sees exactly the matrix the real run will see.

Parameters:
  • app_key – 'umap', 'ml_analyze' or 'classify'.

  • settings – the app’s settings dict; src and tables are read.

Returns:

the SearchData.

Raises:

ValueError – when src is missing, or when a supervised search finds fewer than two classes.

Walk UMAP’s n_neighbors x min_dist plane by 2-by-2 rounds.

The original two-axis Walk, kept as the name every existing caller uses. It is now walk_search() with two numeric axes at resolution 2, which produces the same four diagonal corners per round; pass axes to that function directly to search more than these two parameters.

Parameters:
  • fit_fn – called as fit_fn(params) with a clamped integer n_neighbors, a float min_dist, and every frozen key from start. A call that raises is recorded as a failed trial; a round in which all four fail ends the walk with no winner.

  • start – needs n_neighbors and min_dist – missing or non-numeric raises. Out-of-range values are clamped rather than rejected (n_neighbors up to at least 2, min_dist into [0, 1]). Any other key is held fixed and passed to every fit unchanged.

  • n_trials – maximum ROUNDS, not fits; each round costs up to four. Blank or None means 100. Zero, negative or non-numeric raises.

  • n_neighbors_step – truncated by int(), so 1.9 steps by 1 and anything below 1 becomes 0 and raises instead of stepping fractionally.

  • n_neighbors_max – upper clamp on the n_neighbors axis. None leaves it unbounded and the walk climbs until the score stops improving. Below 2 raises.

  • min_dist_step – step along min_dist; must be strictly positive.

  • min_improvement – a round must beat the running best by MORE than this to continue, so the default 0.0 still stops on an exact tie. Negative raises.

  • metric – a label recorded on the result, nothing more. It does not choose a criterion – whatever fit_fn returns is the score, whatever this names it.

  • higher_is_better – the direction the walk climbs, and the comparison that picks the winner.

  • on_trial – (trial, completed, total) after every trial, failures included. total is 4 * n_trials, an upper bound the walk normally undershoots because an already-scored configuration is never refitted.

  • should_stop – polled before each new fit; True truncates the walk and marks the result partial.

  • notes – caveats placed BEFORE the walk’s own generated notes.

  • checkpoint – resumes an interrupted walk – completed trials are replayed without refitting and the centre, round count and best score are restored. None runs without persistence.

Evaluate n_trials configurations drawn at random from space.

Reproducible: the same seed always yields the same configurations in the same order.

Parameters:
  • fit_fn – called as fit_fn(params).

  • space – the SearchSpace to sample from.

  • n_trials – how many configurations to evaluate; must be positive.

  • seed – RNG seed.

  • metric – name of the criterion the score represents.

  • higher_is_better – direction of metric.

  • on_trial – progress callback (trial, completed, total).

  • should_stop – polled before each trial.

  • notes – extra caveats to attach.

  • allow_duplicates – when False (the default) the same configuration is never evaluated twice; the sweep shrinks to the size of the space if the space is smaller than n_trials.

Returns:

the SearchResult.

Raises:

ValueError – when n_trials is not a positive integer.

spacr.hyperparam.run_search_for_app(app_key: str, settings: Mapping[str, Any], space: SearchSpace, *, criterion: str | None = None, mode: str = 'grid', n_trials: int = 12, adaptive: bool = False, walk_parameters: Sequence[str] | None = None, walk_resolutions: Mapping[str, int] | None = None, walk_steps: Mapping[str, float] | None = None, n_neighbors_step: int = 1, min_dist_step: float = 0.05, min_improvement: float = 0.0, stability_repeats: int = 3, objective_weights: Mapping[str, Any] | None = None, umap_backend: str = 'cpu', cluster_during_search: bool = False, cluster_sizes: Sequence[int] = (5, 10, 15, 25, 40), umap_components: int = 2, seed: int = 0, n_folds: int = 5, on_trial: Callable[[Trial, int, int], None] | None = None, should_stop: Callable[[], bool] | None = None, data: SearchData | None = None, checkpoint_path: str | None = None, resume: bool = False) → SearchResult[source]

Run the right search for a spaCR app. This is what the GUI calls.

  • umap — embeds the measurement features once per configuration and ranks them with a named criterion (see umap_search()).

  • ml_analyze — grouped cross-validated search over classical ML hyperparameters (see cv_search()).

  • classify — one cross-validated deep-training run per configuration (see classify_cv_fit_fn()).

  • activation — one attribution per image per configuration, scored by deletion AUC, insertion AUC, the pointing game and the randomisation sanity check (see activation_search()).

Parameters:
  • app_key – which app is asking.

  • settings – that app’s settings dict.

  • space – the SearchSpace to search.

  • criterion – metric name; defaults to the app’s first APP_CRITERIA entry.

  • mode – 'grid' or 'random'.

  • n_trials – configurations to evaluate when mode='random'.

  • adaptive – for UMAP only, run a Walk from one starting point.

  • walk_parameters – which UMAP parameters the Walk searches; the default is n_neighbors and min_dist. Any subset of UMAP_WALK_PARAMETERS.

  • walk_resolutions – per-axis Walk grid resolution, {name: n}.

  • walk_steps – per-axis Walk step override, {name: size}.

  • n_neighbors_step – Walk integer neighborhood increment.

  • min_dist_step – Walk min_dist increment.

  • min_improvement – Walk score-gain stopping threshold.

  • stability_repeats – repeated seeded embeddings per multi-objective UMAP configuration.

  • objective_weights – weights for neighborhood preservation, stability and cluster structure in multi-objective UMAP mode.

  • umap_backend – 'cpu' or the explicitly requested 'cuml'.

  • cluster_during_search – cluster each UMAP trial as it is completed.

  • cluster_sizes – HDBSCAN scales searched for each UMAP trial.

  • umap_components – fixed 2-D or 3-D output for UMAP trials.

  • seed – seed for sampling, folds and reducers.

  • n_folds – cross-validation folds for the supervised apps.

  • on_trial – progress callback (trial, completed, total).

  • should_stop – polled before each trial.

  • data – pre-loaded SearchData (or ActivationSearchData for 'activation'), skipping the database / model read.

  • checkpoint_path – UMAP checkpoint path; when omitted the UMAP project path is derived by umap_checkpoint_path().

  • resume – continue a compatible UMAP search checkpoint.

Returns:

the SearchResult.

Raises:

ValueError – for an unknown app_key or mode.

spacr.hyperparam.sklearn_cv_fit_fn(features, labels, model_type: str = 'xgboost', *, criterion: str = 'accuracy', seed: int = 42, n_jobs: int = -1)[source]

Build the fit_fn(params, train_idx, val_idx) cv_search() wants.

The estimator is fitted on the fold’s training indices and scored on the fold’s validation indices — the function is never given any other indices, so it structurally cannot score on the held-out test split.

Parameters:
  • features – 2-D numeric feature matrix.

  • labels – per-row class labels.

  • model_type – which classifier to build (see build_sklearn_model()).

  • criterion – 'accuracy', 'roc_auc' or 'f1'.

  • seed – random_state for the estimator.

  • n_jobs – worker count where supported.

Returns:

the fit function.

spacr.hyperparam.umap_available() → Tuple[bool, str][source]

Whether umap-learn can be imported.

Returns:

(True, "") when available, otherwise (False, message) carrying UMAP_MISSING_MESSAGE.

spacr.hyperparam.umap_checkpoint_path(settings: Mapping[str, Any]) → str | None[source]

Return the default UMAP-search checkpoint path for module settings.

checkpoint_path wins when explicitly supplied. Otherwise the path is <project>/results/.spacr_checkpoints/umap_search.json, with database and measurements/ inputs normalised back to their project root.

Parameters:

settings – Image UMAP module settings.

Returns:

absolute path, or None when no source/project can be inferred.

spacr.hyperparam.umap_metrics() → Tuple[str, ...][source]

Every metric the INSTALLED umap-learn will accept.

Falls back to UMAP_METRICS when umap-learn is absent, because the settings panel has to build on a machine that cannot run UMAP – a user configuring a run on a laptop and executing it elsewhere is an ordinary thing to do.

spacr.hyperparam.umap_objective_scores(features: Any, embeddings: Sequence[Any], *, labels: Any = None, neighbourhood_k: int = 15, weights: Mapping[str, Any] | None = None, seed: int = 0) → Dict[str, Any][source]

Score neighborhood preservation, stability and cluster structure.

The returned multi_objective value is a weighted geometric mean used to guide grid/adaptive search. The individual objective values remain the primary result and define SearchResult.pareto_front().

Parameters:
  • features – read only for trustworthiness and continuity, hence only for neighborhood_preservation. Stability and cluster structure come from the embeddings alone and do not move with it.

  • embeddings – two or more fits of the same rows; fewer than two raises, because stability is a repeat-to-repeat measure. How they were produced is the caller’s business – only their count and geometry are used here.

  • labels – optional, and consumed in two places: the silhouette entry (present only when every repeat could compute it) and the cluster-structure partition. Labels whose length does not match the rows, or carrying fewer than two classes, are silently ignored and K-means discovery runs instead, so read cluster_structure_method rather than assuming they were used.

  • neighbourhood_k – one value, two different caps. Trustworthiness and continuity clamp it to (n_samples - 1) // 2 and report that clamped number back as neighbourhood_k; stability clamps only to n_samples - 1. A k near the sample count therefore drives stability to a meaningless 1.0 while the reported k still looks reasonable.

  • weights – merged over DEFAULT_UMAP_OBJECTIVE_WEIGHTS and then renormalized, so naming one objective does not zero the others – {'stability': 1.0} ends up near 0.59, not 1.0. Unknown, negative, non-finite, non-numeric or all-zero raises.

  • seed – reaches only the K-means discovery path, offset by the repeat index so the repeats are deliberately not identical. It has no effect at all when usable labels are supplied.

Returns:

the three objectives plus multi_objective, the component trustworthiness/continuity, the normalized weights, and the provenance fields cluster_structure_method and cluster_counts.

Sweep UMAP parameters, scoring each embedding with a named criterion.

The honest deliverable is the panel of embeddings, not the top row of the table. Every trial keeps its embedding (unless keep_embeddings is off) so the caller can draw small multiples, and every criterion is computed for every trial so the caller can see how the ranking changes when the criterion does.

Parameters:
  • features – 2-D numeric feature matrix, (n_samples, n_features).

  • space – UMAP parameters to sweep (n_neighbors, min_dist, metric, …).

  • metric – which criterion drives the ranking; one of UMAP_CRITERIA.

  • labels – optional class labels; required for 'silhouette'.

  • seed – random_state for the reducer, so the sweep reproduces.

  • n_components – fixed display dimensionality, 2 or 3. It is retained on every row even when it is not one of the searched axes.

  • neighbourhood_k – neighbourhood size for trustworthiness/continuity.

  • adaptive – run a Walk – iterative local optimization from one starting point – instead of a grid. See walk_search().

  • walk_parameters – which UMAP parameters the Walk searches. Defaults to n_neighbors and min_dist, the two it has always used; any subset of UMAP_WALK_PARAMETERS is accepted.

  • walk_resolutions – per-axis grid resolution, {name: n}. 2 is the classic +/-step pair. Missing axes get 2.

  • walk_steps – per-axis step override, {name: size}.

  • n_trials – maximum complete Walk rounds; blank or None means 100.

  • n_neighbors_step – local step along the n_neighbors axis.

  • min_dist_step – local step along the min_dist axis.

  • min_improvement – score gain required to continue after a round.

  • stability_repeats – reproducible UMAP fits per configuration when metric='multi_objective'; must be at least 2.

  • objective_weights – optional weights for neighborhood_preservation, stability and cluster_structure. Values are normalized to sum to one.

  • embed_fn – embed_fn(features, params) -> embedding override; when omitted, umap-learn is used.

  • backend – 'cpu' for umap-learn or 'cuml' for the optional RAPIDS implementation. Every successful trial records the backend that actually drew it because the two libraries make different maps.

  • cluster_during_search – run the HDBSCAN scale walk against every completed embedding and retain its best labels beside the coordinates.

  • cluster_sizes – min_cluster_size values for that walk.

  • keep_embeddings – store each trial’s embedding in its extra metrics.

  • on_trial – progress callback (trial, completed, total).

  • should_stop – polled before each trial.

  • checkpoint_path – optional atomic checkpoint JSON. Embeddings are stored as adjacent .npy artifacts after each completed trial.

  • resume – load a compatible checkpoint. Input features, labels, search space, criterion, seed and material search settings must match.

Returns:

the SearchResult. When umap-learn is missing and no embed_fn was given, this returns an empty result whose notes lead with UMAP_MISSING_MESSAGE rather than raising ImportError.

Raises:

ValueError – when metric is not a known criterion, or 'silhouette' is requested without labels.

spacr.hyperparam.umap_walk_axes(start: Mapping[str, Any], *, parameters: Sequence[str] | None = None, steps: Mapping[str, float] | None = None, resolutions: Mapping[str, int] | None = None, n_neighbors_max: int | None = None) → List[WalkAxis][source]

Build Walk axes for UMAP from a starting configuration.

parameters names which of UMAP’s structural parameters take part. The default is the two the search has always used, so an existing call is unchanged; the panel passes the user’s selection.

Every axis carries the range UMAP itself requires – n_neighbors at least 2, min_dist within [0, 1], set_op_mix_ratio within [0, 1] – because a walk is the one search that generates values that were never typed by anyone, and an out-of-range one fails inside the fit rather than at the edge.

Parameters:
  • start – only its KEYS are read. Every searched name must be present or this raises; the starting values never reach the axes, so two different starting points build identical axes.

  • parameters – which names take part. Empty or None falls back to the default pair, and a name outside UMAP_WALK_PARAMETERS raises. A name listed twice is passed through here and only rejected later, by walk_search().

  • steps – per-axis step override. Discarded on the categorical axes (init, metric), which move by choice rather than by step – a step given for metric is even stored on the axis, but nothing ever reads it.

  • resolutions – per-axis values per round, counting the centre; axes not named get 2. Below 2, or not a whole number, raises.

  • n_neighbors_max – upper clamp for the n_neighbors axis only, and silently inert when that axis is not searched. Unvalidated here, unlike in local_direction_search(): below 2 it collides with the fixed minimum and raises from WalkAxis, and a fractional cap rounds up (7.9 admits 8).

spacr.hyperparam.walk_neighbourhood(axes: Sequence[WalkAxis], centre: Mapping[str, Any], *, max_candidates: int = MAX_WALK_CANDIDATES_PER_ROUND) → Tuple[List[Dict[str, Any]], bool][source]

The configurations one Walk round evaluates around centre.

The full neighbourhood is the Cartesian product of every axis’s WalkAxis.values_around(), minus the centre itself. With two numeric axes at resolution 2 that is the four diagonal corners the original 2-by-2 search used, which is the point: the old behaviour is this function’s two-axis case and not a separate code path.

The product is exponential in the number of axes, so when it exceeds max_candidates the round falls back to varying one axis at a time – linear in the axis count, and still enough to choose a direction, at the cost of not seeing interactions between axes.

Returns:

(candidates, full_factorial). The flag is False when the fallback was used, and the caller is expected to say so in the result notes rather than quietly search less than it claimed.

Walk N-dimensional hyperparameter space toward a better score.

Each round scores the neighbourhood around the current centre – see walk_neighbourhood() – and moves to the best configuration found, stopping when a round fails to improve on the best score by more than min_improvement. The starting point is a centre, not a trial: it is never fitted, because a walk asks “which way is better from here”, and “here” is where the user already is.

With two numeric axes at resolution 2 this is the 2-by-2 search that preceded it, corner for corner. It is not restricted to two: any parameter the fit function accepts can be an axis, and each axis carries its own step and resolution.

Parameters:
  • fit_fn – callable invoked with each candidate parameter mapping. It returns a number, (score, metrics) pair, or mapping containing a score key, exactly as _normalise_outcome() accepts.

  • start – one value per axis, plus any parameters held fixed – anything not named by an axis is passed to every fit unchanged.

  • axes – the search space. Empty raises.

  • n_trials – maximum complete rounds (100 when blank/None). NOT the number of fits, which is rounds times the neighbourhood size.

  • max_candidates_per_round – past this, a round varies one axis at a time instead of taking the full product. Recorded in the notes.

Nested helpers

activation_fit_fn._attribute(params: Mapping[str, Any], image)

Attribute one image with one trial’s configuration.

spacr/hyperparam.py:2660

activation_fit_fn._fit(params: Dict[str, Any]) → Tuple[float, Dict[str, Any]]

Score one configuration on every image, reporting every criterion.

spacr/hyperparam.py:2672

classify_cv_fit_fn._default_read(path)

Read the per-fold CSV a cross-validated training run wrote.

spacr/hyperparam.py:3637

classify_cv_fit_fn._default_train(cfg)

Run spaCR’s own cross-validated training for one configuration.

spacr/hyperparam.py:3632

classify_cv_fit_fn._fit(params)

Train one configuration across the folds and average the metric.

spacr/hyperparam.py:3645

cv_search._call(fn, params: Dict[str, Any]) → Tuple[float, Dict[str, Any]]

Fan one configuration out over the folds and average the scores.

spacr/hyperparam.py:3152

sklearn_cv_fit_fn._fit(params, train_idx, val_idx)

Fit on the fold’s training rows, score on its validation rows.

spacr/hyperparam.py:3572

umap_search._fit(params: Dict[str, Any]) → Tuple[float, Dict[str, Any]]

Embed one configuration and score it with every criterion.

spacr/hyperparam.py:2360

umap_search.embed_fn(feats, params, _seed=seed)

Default embedder — umap-learn with a pinned random_state.

cuML embedder; no silent CPU downgrade for a checked GPU run.

spacr/hyperparam.py:2259 spacr/hyperparam.py:2270

walk_search._persisted_state() → Dict[str, Any]

Return a fresh walk checkpoint with legacy two-axis centre keys.

spacr/hyperparam.py:1496