spaCR nightly · Main documentation · Nightly preview
Contents Menu Expand Light mode Dark mode Auto light/dark, in light mode Auto light/dark, in dark mode Skip to content
spaCR 1.5.1.4 documentation
Logo
spaCR 1.5.1.4 documentation
  • Installer guide
  • System requirements
  • Choose a workflow after installation
  • Combine measurement tables for plots and gates
  • Installer archive
  • Capabilities
  • Keyboard shortcuts and settings templates
  • Make Masks: editing, detection and measurement
  • Measure: preview checked images and verify mask planes
  • Train a Cellpose model
  • Process images with a point-spread function
  • Recruitment: compartment ratios and channel identity
  • Image quality before segmentation
  • Host–Pathogen Analysis
  • Plaque Assay: fields, figures and reviewed conditions
  • Python API quickstart
  • Export measurements to AnnData
  • Where a setting goes
  • Model zoo
  • Language
  • Setting animation gallery
  • Checkpoint and resume
  • Reproducibility manifests
  • Unified run history
  • Plate and batch-effect correction
  • Classifier evaluation workbench
  • Plate-aware guide permutation analysis
  • Explain CV models and investigate hits
  • Resumable multi-objective UMAP search
  • Remote and distributed execution
  • spaCR plugin SDK
  • Train/test leakage audit
  • Threading and cancellation audit
  • Database concurrency audit
  • API reference
    • spacr
      • spacr.api
      • spacr.core
      • spacr.measure
      • spacr.deep_spacr
      • spacr.sequencing
      • spacr.ml
      • spacr.artifacts
      • spacr.settings
Back to top
View this page
Edit this page

spacr.umap_search¶

Represent Image UMAP searches as reproducible embedding recipes.

Each search row stores the parameters and score needed to redraw its two- or three-dimensional embedding and continue with clustering.

Each stored recipe includes its selected columns, random state and backend so the associated score remains bound to the embedding that produced it. cuML and umap-learn can produce different embeddings from the same data and hyperparameters; recording the backend therefore preserves provenance.

The module is independent of Qt and can be tested without a display.

Classes¶

ClusterWalkRow

One clustering tried against one fixed embedding.

SearchRow

One trial: its recipe, its scores, and the embedding it produced.

SearchTable

The rows a search produced, in the order they were scored.

UmapRecipe

Everything needed to redraw one embedding, and nothing else.

Functions¶

cluster_embedding(→ numpy.ndarray)

Cluster a fixed 2-D or 3-D embedding with HDBSCAN.

walk_clusters(, min_samples)

Try HDBSCAN scales on one map and return them best-first.

walk_recipes(, min_dists, components)

The recipes a walk would try, worked out before any of them runs.

Module Contents¶

class spacr.umap_search.ClusterWalkRow[source]¶

One clustering tried against one fixed embedding.

The coordinates are deliberately not stored here: a clustering walk changes the partition, not the map. Keeping that distinction explicit prevents a cluster button from quietly refitting UMAP and making the row the user selected cease to be the row they are looking at.

Parameters:
  • min_cluster_size – HDBSCAN minimum cluster size used for this trial.

  • labels – cluster label for each embedding row, with noise represented by -1.

  • silhouette – silhouette score over assigned points when defined; a non-finite value makes score rank below every measured trial.

  • n_clusters – number of non-noise clusters found.

  • noise_fraction – fraction of embedding rows assigned to noise, used to discount score.

property score: float[source]¶

Return separation discounted by the unassigned-point fraction.

class spacr.umap_search.SearchRow[source]¶

One trial: its recipe, its scores, and the embedding it produced.

embedding retains the exact array used to calculate scores. Recomputing an embedding when a row is selected could produce different coordinates with a non-deterministic backend.

Parameters:
  • recipe – complete embedding recipe evaluated by this trial.

  • scores – named quality measurements calculated from this embedding.

  • embedding – exact coordinates those scores describe, retained so selecting the row never silently refits a different map.

  • labels – optional cluster assignment aligned with the embedded rows; negative labels represent noise.

  • note – warning or explanatory text retained and exported with the trial.

cluster_count() → int[source]¶

How many clusters this row’s labels describe, noise excluded.

property score: float[source]¶

The headline number the table sorts on.

class spacr.umap_search.SearchTable[source]¶

The rows a search produced, in the order they were scored.

Deliberately not a DataFrame: a row owns an ndarray and a recipe, and putting those in cells makes every operation on the table a chance to lose the pairing between a score and the embedding it describes.

Initialize an empty insertion-ordered search-result table.

__getitem__(index: int | slice) → SearchRow | List[SearchRow][source]¶

Return one row or a list slice using ordinary list semantics.

__iter__()[source]¶

Iterate over retained rows in insertion order.

__len__() → int[source]¶

Return the number of retained search rows.

add(row: SearchRow) → SearchRow[source]¶

Append and return one search result row.

Parameters:

row – search result to retain in insertion order.

Returns:

row after appending it.

backends() → Tuple[str, ...][source]¶

Which backends drew these rows.

More than one is worth saying out loud: a table mixing cuML and umap-learn rows is comparing two libraries as well as the settings.

Returns:

sorted distinct backend names from retained recipes.

best() → SearchRow | None[source]¶

The highest-scoring row, or None when nothing scored.

Rows with a NaN score are excluded from the comparison.

Returns:

highest-scoring finite row, or None when none exists.

mixed_backends() → bool[source]¶

Return whether retained recipes name multiple backends.

to_dicts() → List[Dict[str, Any]][source]¶

Return serializable rows without embedding or label arrays.

property rows: List[SearchRow][source]¶

Return a shallow outer-list copy in insertion order.

class spacr.umap_search.UmapRecipe[source]¶

Everything needed to redraw one embedding, and nothing else.

Frozen and round-tripping, so a row saved to disk today rebuilds the same requested configuration tomorrow. columns is part of it: a recipe that recorded only the hyperparameters would request a different map the moment the column selection changed. The exact scored coordinates remain on SearchRow, because nondeterministic backends and dependency changes can produce a different map from the same recipe.

Variables:
  • n_neighbors – neighbourhood size balancing local detail against global structure in the embedding.

  • min_dist – minimum separation between embedded points, controlling how tightly local clusters may pack.

  • n_components – drawable output dimensions, clamped to two or three.

  • metric – distance function used to compare input feature vectors.

  • random_state – seed retained so the CPU embedding can be reproduced.

  • scale – standardize selected features before fitting when true.

  • columns – exact input feature columns scored by this recipe.

  • backend – implementation used to build the map, such as CPU UMAP or cuML; different backends are treated as different recipes.

__post_init__() → None[source]¶

Normalize columns and clamp dimensions to the supported 2–3.

classmethod from_dict(payload: Dict[str, Any]) → UmapRecipe[source]¶

Build a recipe from known fields in a stored mapping.

Parameters:

payload – serialized recipe mapping, possibly with newer fields.

Returns:

normalized recipe containing only fields this version knows.

label() → str[source]¶

Return the compact configuration label shown in a table cell.

to_dict() → Dict[str, Any][source]¶

Return a storable field mapping with columns represented as a list.

property is_3d: bool[source]¶

Return whether this recipe requests a three-dimensional map.

spacr.umap_search.cluster_embedding(embedding: Any, *, min_cluster_size: int = 15, min_samples: int | None = None) → numpy.ndarray[source]¶

Cluster a fixed 2-D or 3-D embedding with HDBSCAN.

scikit-learn’s implementation is used because it is already a spaCR core dependency (spaCR requires a version new enough to provide HDBSCAN). No DBSCAN substitution is made: changing the algorithm while keeping the HDBSCAN label would make the cluster count beside a map false provenance.

Parameters:
  • embedding – finite coordinate array shaped (rows, 2) or (rows, 3).

  • min_cluster_size – smallest group HDBSCAN may call a cluster; at least two and smaller than the embedding row count.

  • min_samples – optional HDBSCAN core-sample threshold; None and zero leave it unset.

Returns:

one integer label per embedding row, with noise labelled -1.

Raises:
  • ValueError – when the coordinates or clustering thresholds cannot describe a valid partition.

  • RuntimeError – when HDBSCAN returns the wrong number of labels.

spacr.umap_search.walk_clusters(embedding: Any, *, min_cluster_sizes: Sequence[int] = (5, 10, 15, 25, 40), min_samples: int | None = None) → List[ClusterWalkRow][source]¶

Try HDBSCAN scales on one map and return them best-first.

This is the clustering half of the Starplast-style walk. It can run for every UMAP trial as that trial arrives, or later against the table row the user chose. Duplicate and out-of-range candidate sizes are skipped; if no size is meaningful the call is refused rather than fabricating a one-cluster winner. A clustering failure propagates, while an undefined silhouette is retained as nan and ranks below every measured score.

Parameters:
  • embedding – fixed finite 2-D or 3-D coordinates to cluster at each candidate scale.

  • min_cluster_sizes – candidate HDBSCAN minimum cluster sizes.

  • min_samples – optional HDBSCAN core-sample threshold passed to every candidate.

Returns:

scored partitions ordered best-first, then by cluster size.

Raises:

ValueError – when the map or candidate sequence is invalid.

spacr.umap_search.walk_recipes(base: UmapRecipe, *, steps: int = 12, neighbors: Sequence[int] = (), min_dists: Sequence[float] = (), components: Sequence[int] = ()) → List[UmapRecipe][source]¶

The recipes a walk would try, worked out before any of them runs.

The returned list lets the panel report the total trial count before the first trial starts.

Parameters:
  • base – recipe cloned for every candidate. Its values fill dimensions with no explicit grid, and its neighbour count scales the default neighbour walk.

  • steps – target sample count when no explicit grid is given. At least two neighbor values are attempted, and integer rounding plus deduplication may change the final count.

  • neighbors – explicit neighborhood-size values, or empty to derive a walk from base and steps.

  • min_dists – explicit minimum-distance values, or empty to retain the base value.

  • components – explicit dimensionalities, or empty to retain the base value.

Returns:

distinct recipes in Cartesian-product order.

Copyright © 2025-2026, Einar Birnir Olafsson
Made with Sphinx and @pradyunsg's Furo
On this page
  • spacr.umap_search
    • Classes
    • Functions
    • Module Contents
      • ClusterWalkRow
        • ClusterWalkRow.score
      • SearchRow
        • SearchRow.cluster_count()
        • SearchRow.score
      • SearchTable
        • SearchTable.__getitem__()
        • SearchTable.__iter__()
        • SearchTable.__len__()
        • SearchTable.add()
        • SearchTable.backends()
        • SearchTable.best()
        • SearchTable.mixed_backends()
        • SearchTable.to_dicts()
        • SearchTable.rows
      • UmapRecipe
        • UmapRecipe.__post_init__()
        • UmapRecipe.from_dict()
        • UmapRecipe.label()
        • UmapRecipe.to_dict()
        • UmapRecipe.is_3d
      • cluster_embedding()
      • walk_clusters()
      • walk_recipes()