spacr.umap_search¶
Represent Image UMAP searches as reproducible embedding recipes.
Each search row stores the parameters and score needed to redraw its two- or three-dimensional embedding and continue with clustering.
Each stored recipe includes its selected columns, random state and backend so the associated score remains bound to the embedding that produced it. cuML and umap-learn can produce different embeddings from the same data and hyperparameters; recording the backend therefore preserves provenance.
The module is independent of Qt and can be tested without a display.
Classes¶
One clustering tried against one fixed embedding. |
|
One trial: its recipe, its scores, and the embedding it produced. |
|
The rows a search produced, in the order they were scored. |
|
Everything needed to redraw one embedding, and nothing else. |
Functions¶
|
Cluster a fixed 2-D or 3-D embedding with HDBSCAN. |
|
Try HDBSCAN scales on one map and return them best-first. |
|
The recipes a walk would try, worked out before any of them runs. |
Module Contents¶
- class spacr.umap_search.ClusterWalkRow[source]¶
One clustering tried against one fixed embedding.
The coordinates are deliberately not stored here: a clustering walk changes the partition, not the map. Keeping that distinction explicit prevents a cluster button from quietly refitting UMAP and making the row the user selected cease to be the row they are looking at.
- Parameters:
min_cluster_size – HDBSCAN minimum cluster size used for this trial.
labels – cluster label for each embedding row, with noise represented by
-1.silhouette – silhouette score over assigned points when defined; a non-finite value makes
scorerank below every measured trial.n_clusters – number of non-noise clusters found.
noise_fraction – fraction of embedding rows assigned to noise, used to discount
score.
- class spacr.umap_search.SearchRow[source]¶
One trial: its recipe, its scores, and the embedding it produced.
embeddingretains the exact array used to calculatescores. Recomputing an embedding when a row is selected could produce different coordinates with a non-deterministic backend.- Parameters:
recipe – complete embedding recipe evaluated by this trial.
scores – named quality measurements calculated from this embedding.
embedding – exact coordinates those scores describe, retained so selecting the row never silently refits a different map.
labels – optional cluster assignment aligned with the embedded rows; negative labels represent noise.
note – warning or explanatory text retained and exported with the trial.
- class spacr.umap_search.SearchTable[source]¶
The rows a search produced, in the order they were scored.
Deliberately not a DataFrame: a row owns an ndarray and a recipe, and putting those in cells makes every operation on the table a chance to lose the pairing between a score and the embedding it describes.
Initialize an empty insertion-ordered search-result table.
- __getitem__(index: int | slice) SearchRow | List[SearchRow][source]¶
Return one row or a list slice using ordinary list semantics.
- add(row: SearchRow) SearchRow[source]¶
Append and return one search result row.
- Parameters:
row – search result to retain in insertion order.
- Returns:
rowafter appending it.
- backends() Tuple[str, ...][source]¶
Which backends drew these rows.
More than one is worth saying out loud: a table mixing cuML and umap-learn rows is comparing two libraries as well as the settings.
- Returns:
sorted distinct backend names from retained recipes.
- best() SearchRow | None[source]¶
The highest-scoring row, or None when nothing scored.
Rows with a NaN score are excluded from the comparison.
- Returns:
highest-scoring finite row, or
Nonewhen none exists.
- class spacr.umap_search.UmapRecipe[source]¶
Everything needed to redraw one embedding, and nothing else.
Frozen and round-tripping, so a row saved to disk today rebuilds the same requested configuration tomorrow.
columnsis part of it: a recipe that recorded only the hyperparameters would request a different map the moment the column selection changed. The exact scored coordinates remain onSearchRow, because nondeterministic backends and dependency changes can produce a different map from the same recipe.- Variables:
n_neighbors – neighbourhood size balancing local detail against global structure in the embedding.
min_dist – minimum separation between embedded points, controlling how tightly local clusters may pack.
n_components – drawable output dimensions, clamped to two or three.
metric – distance function used to compare input feature vectors.
random_state – seed retained so the CPU embedding can be reproduced.
scale – standardize selected features before fitting when true.
columns – exact input feature columns scored by this recipe.
backend – implementation used to build the map, such as CPU UMAP or cuML; different backends are treated as different recipes.
- classmethod from_dict(payload: Dict[str, Any]) UmapRecipe[source]¶
Build a recipe from known fields in a stored mapping.
- Parameters:
payload – serialized recipe mapping, possibly with newer fields.
- Returns:
normalized recipe containing only fields this version knows.
- spacr.umap_search.cluster_embedding(embedding: Any, *, min_cluster_size: int = 15, min_samples: int | None = None) numpy.ndarray[source]¶
Cluster a fixed 2-D or 3-D embedding with HDBSCAN.
scikit-learn’s implementation is used because it is already a spaCR core dependency (spaCR requires a version new enough to provide HDBSCAN). No DBSCAN substitution is made: changing the algorithm while keeping the HDBSCAN label would make the cluster count beside a map false provenance.
- Parameters:
embedding – finite coordinate array shaped
(rows, 2)or(rows, 3).min_cluster_size – smallest group HDBSCAN may call a cluster; at least two and smaller than the embedding row count.
min_samples – optional HDBSCAN core-sample threshold;
Noneand zero leave it unset.
- Returns:
one integer label per embedding row, with noise labelled
-1.- Raises:
ValueError – when the coordinates or clustering thresholds cannot describe a valid partition.
RuntimeError – when HDBSCAN returns the wrong number of labels.
- spacr.umap_search.walk_clusters(embedding: Any, *, min_cluster_sizes: Sequence[int] = (5, 10, 15, 25, 40), min_samples: int | None = None) List[ClusterWalkRow][source]¶
Try HDBSCAN scales on one map and return them best-first.
This is the clustering half of the Starplast-style walk. It can run for every UMAP trial as that trial arrives, or later against the table row the user chose. Duplicate and out-of-range candidate sizes are skipped; if no size is meaningful the call is refused rather than fabricating a one-cluster winner. A clustering failure propagates, while an undefined silhouette is retained as
nanand ranks below every measured score.- Parameters:
embedding – fixed finite 2-D or 3-D coordinates to cluster at each candidate scale.
min_cluster_sizes – candidate HDBSCAN minimum cluster sizes.
min_samples – optional HDBSCAN core-sample threshold passed to every candidate.
- Returns:
scored partitions ordered best-first, then by cluster size.
- Raises:
ValueError – when the map or candidate sequence is invalid.
- spacr.umap_search.walk_recipes(base: UmapRecipe, *, steps: int = 12, neighbors: Sequence[int] = (), min_dists: Sequence[float] = (), components: Sequence[int] = ()) List[UmapRecipe][source]¶
The recipes a walk would try, worked out before any of them runs.
The returned list lets the panel report the total trial count before the first trial starts.
- Parameters:
base – recipe cloned for every candidate. Its values fill dimensions with no explicit grid, and its neighbour count scales the default neighbour walk.
steps – target sample count when no explicit grid is given. At least two neighbor values are attempted, and integer rounding plus deduplication may change the final count.
neighbors – explicit neighborhood-size values, or empty to derive a walk from
baseandsteps.min_dists – explicit minimum-distance values, or empty to retain the base value.
components – explicit dimensionalities, or empty to retain the base value.
- Returns:
distinct recipes in Cartesian-product order.