spacr.embeddings

A label-free embedding for every object, beside the measured features.

WHY THIS EXISTS. spaCR’s two existing ways of turning an image into numbers both require you to already know what you are looking for. The measured panel in spacr.measure computes a fixed, hand-designed set – intensity, shape, texture, distances – so it finds only what somebody thought to name. The supervised path in spacr.deep_spacr fine-tunes a backbone against annotations, so it finds only what somebody thought to annotate. A screen whose real hit is “the vacuole is subtly mislocalised in a way we have no column for” returns nothing, and returns it confidently. The cost of that gap is invisible, which is exactly why it is worth a module.

THE CHANNEL DECISION, SETTLED HERE RATHER THAN IMPLICITLY. Pretrained encoders expect three channels. spaCR images routinely have four or five, with no RGB meaning at all – a DAPI channel is not “red”. Two ways out, and the choice is recorded here rather than inherited from whatever the first backbone happened to want:

PER-CHANNEL, THEN CONCATENATE (CHANNEL_PER_CHANNEL, the default). Each channel is encoded on its own and the vectors are joined. Channel identity SURVIVES: dimension 517 belongs to channel 2 and to nothing else, so spacr.attribution_columns can say which stain a hit came from and a reader can drop a channel without re-embedding. Costs N forward passes.

PROJECT TO THREE (CHANNEL_PROJECT). One pass over a fixed channel-to-RGB mapping. Cheaper, and irreversible: once channels are mixed, no downstream column can be attributed to a stain, and a screen cannot ask “was this the parasite channel or the host one?”.

The default is per-channel because THE ATTRIBUTION IS THE POINT. An embedding that finds an unnamed phenotype and cannot say which stain carried it has moved the problem rather than solved it. Projection stays available for the case where forward passes actually bind.

WHAT THIS MODULE DOES NOT DO. It does not crop – it takes the crops spacr.crops already produces. It does not write to the database. It returns a float32 matrix keyed by object id, which is exactly the shape the downstream already consumes: UMAP, the regression family, hit calling and the AnnData export do not care where a numeric column came from.

Exceptions

EmbeddingError

Raised when embedding is refused, with the reason a caller can show.

Classes

EmbeddingResult

The matrix, its column names, and how it was made.

EmbeddingSpec

How to embed, recorded so a matrix can say where it came from.

Functions

channel_of_column(→ Optional[int])

Which channel a column came from, or None if it cannot be said.

embed_array(→ EmbeddingResult)

Embed a stack of crops.

embedding_column_names(→ Tuple[str, ...])

Names for the columns a run produces.

encoder_entry([spec, scorecard])

Describe this encoder as a spacr.model_zoo.ModelEntry.

encoder_key(→ str)

The stable id for one encoder configuration.

Module Contents

exception spacr.embeddings.EmbeddingError[source]

Bases: ValueError

Raised when embedding is refused, with the reason a caller can show.

Initialize self. See help(type(self)) for accurate signature.

class spacr.embeddings.EmbeddingResult[source]

The matrix, its column names, and how it was made.

Parameters:
  • values – (n_objects, n_dims) float32.

  • columns – one name per column, all starting with EMBEDDING_PREFIX.

  • spec – the EmbeddingSpec that produced it.

  • n_channels – how many channels were encoded.

  • per_channel_dims – dimensions contributed by ONE channel.

to_frame(object_ids: Sequence[Any])[source]

A DataFrame keyed by object id, ready to join to measurements.

Parameters:

object_ids – one id per row, in the order the crops were given.

Raises:

EmbeddingError – when the count does not match the matrix, because a silent misalignment here corrupts every downstream join and is unrecoverable afterwards.

class spacr.embeddings.EmbeddingSpec[source]

How to embed, recorded so a matrix can say where it came from.

For subcell_rybg, choose CHANNEL_PROJECT, supply four explicit, distinct channels in microtubules (R), ER (Y), DNA (B), protein (G) order, and set normalize=False. Crops must be at least 16 pixels in each spatial dimension. Native dimensions are preserved; the encoder applies the authors’ whole-crop min-max normalization.

Parameters:
  • backbone – Name of a timm or registered foundation encoder.

  • channel_policy – CHANNEL_PER_CHANNEL or CHANNEL_PROJECT.

  • channels – which channel indices to encode, in order. None means every channel the array has.

  • batch_size – crops per forward pass.

  • device – torch device string, or None to choose one.

  • normalize – divide each channel by a fixed scale before encoding, clipped to [0, 1]. Microscopy dynamic range varies by orders of magnitude between stains, and an encoder trained on photographs will otherwise see one channel as noise.

  • channel_scale – that scale, one value per encoded channel in encoding order – normally each channel’s 99th percentile over a random sample of the plate. The same numbers for every crop mean an object’s embedding does not depend on which crops share its batch. None lets embed_array() estimate it from the crops it is given.

  • cell_dino_factory – pinned official Cell-DINO hub factory.

  • checkpoint_path – local official Cell-DINO weights file.

  • checkpoint_sha256 – expected SHA-256 of those weights.

__post_init__() → None[source]

Reject a specification that could not produce a matrix.

The check is made where the spec is built rather than where it is used, because one spec encodes every crop of a run: an unknown channel policy or a batch size below one is a typing mistake, and discovering it after the images are loaded wastes the load.

Raises:

EmbeddingError – when channel_policy is not one of CHANNEL_POLICIES, batch_size is below one, or channel_scale is set without normalize, holds a negative or non-finite value, or does not give one value per channel in channels.

fingerprint() → str[source]

A short digest of everything that changes the numbers.

Two matrices with the same fingerprint were produced the same way. A scorecard can quote it, and a cached matrix can be invalidated by it rather than by a timestamp.

spacr.embeddings.channel_of_column(column: str) → int | None[source]

Which channel a column came from, or None if it cannot be said.

None for a projected column is the honest answer rather than a guess: the channels were mixed and no dimension belongs to one stain.

Parameters:

column – an embedding column name.

spacr.embeddings.embed_array(crops: numpy.ndarray, spec: EmbeddingSpec | None = None, *, encoder: Callable[[numpy.ndarray], numpy.ndarray] | None = None) → EmbeddingResult[source]

Embed a stack of crops.

Every crop is divided by the same per-channel scale, spec.channel_scale, so a crop’s embedding does not depend on the other crops in crops. When the spec carries none, one is estimated from a random sample of crops – the stack is treated as the whole plate – and a warning is logged. To embed one plate in several calls, pass the same scale to each: the returned spec carries it.

Parameters:
  • crops – (n, height, width, channels), the shape spacr.crops already produces.

  • spec – how to embed; the default is per-channel resnet18.

  • encoder – a callable taking (n, height, width, 3) float32 in [0, 1] and returning (n, dims). An encoder with an in_channels attribute takes that many planes instead, and one whose in_channels is None takes any number: each channel alone under the per-channel policy, every encoded channel in one pass under the projection policy. Injected by the tests so the wiring can be exercised without downloading a backbone; production callers leave it None.

Returns:

an EmbeddingResult.

Raises:

EmbeddingError – on a shape or channel problem, before any model is loaded – a download is a slow way to find out the array was wrong.

spacr.embeddings.embedding_column_names(per_channel_dims: int, channels: Sequence[int], policy: str = CHANNEL_PER_CHANNEL) → Tuple[str, ...][source]

Names for the columns a run produces.

THE CHANNEL IS IN THE NAME under the per-channel policy – emb_c2_017 – which is what makes attribution possible later without carrying a separate map beside the matrix. Under projection there is no channel to name, and the name says so rather than implying one.

Parameters:
  • per_channel_dims – dimensions one forward pass produces.

  • channels – channel indices, in encoding order.

  • policy – which channel policy produced them.

Returns:

column names, in matrix order.

spacr.embeddings.encoder_entry(spec: EmbeddingSpec | None = None, *, scorecard: Mapping[str, Any] | None = None)[source]

Describe this encoder as a spacr.model_zoo.ModelEntry.

The entry combines encoder configuration, available checkpoint provenance and optional retrieval metrics. Compare entries only when their channel policies are compatible.

The entry records the backbone, channel policy and digest of readable local checkpoint bytes without downloading or loading a model. Public backbones use timm cache metadata; OpenPhenom and ChAda-ViT use pinned Hugging Face revisions; SubCell uses its official torch hub checkpoint; a local DINO encoder uses its configured checkpoint path. Cell-DINO uses the declared local checkpoint only when stable regular-file bytes match its expected SHA-256; changed or mismatched files leave no digest. Foundation entries name the original provider and checkpoint URL. A checksum alone does not verify a model or establish its training-data provenance.

If no readable local checkpoint is found, the entry keeps an empty digest and explains this in its notes. spacr.model_zoo.fetch() refuses entries without a digest.

Parameters:
  • spec – Encoder configuration; uses EmbeddingSpec() when omitted.

  • scorecard – Optional retrieval metrics measured on labelled controls, stored in ModelEntry.metrics for display by scorecard_lines.

Returns:

A ModelEntry with kind 'encoder'.

spacr.embeddings.encoder_key(spec: EmbeddingSpec) → str[source]

The stable id for one encoder configuration.

BACKBONE AND POLICY TOGETHER, because they are jointly what makes two runs’ dimensions mean the same thing. The same backbone under CHANNEL_PER_CHANNEL and CHANNEL_PROJECT produces columns that are the same width, the same dtype and not remotely the same quantity – one is per-stain, the other is a mixture. A key that named only the backbone would let those two be compared silently.

Parameters:

spec – the run’s configuration. Only backbone and channel_policy reach the key: everything else on the spec – batch size, device, crop size – changes how long the run takes and not what a dimension means.

Nested helpers

_cell_dino_model.forward(net, x)

Return the official DINOv2 head output for each declared crop.

spacr/embeddings.py:883

_dino_encoder.run(stack: np.ndarray) → np.ndarray

Encode (n, h, w, k) crops in batches at the training size.

Parameters:

stack – float32 in [0, 1], channels last.

Returns:

(n, dims) float32 features.

spacr/embeddings.py:1382

_dino_network.Head.__init__(self)

Build the MLP and the prototype layer.

spacr/embeddings.py:1177

_dino_network.Head.forward(self, x)

(n, dim) features to (n, out_dim) prototype scores.

spacr/embeddings.py:1185

_dino_views.uniform(low, high, *shape)

Uniform draws from the run’s generator, on the batch’s device.

spacr/embeddings.py:1126

_foundation_encoder.run(stack: np.ndarray) → np.ndarray

Encode declared Cell-DINO planes at the official crop size.

Encode (n, h, w, k) crops in model-appropriate batches.

Parameters:

stack – float32 crops, channels last. SubCell R/Y/B/G uses native intensity values; other foundations expect [0, 1].

Returns:

(n, dims) float32 features.

spacr/embeddings.py:783 spacr/embeddings.py:809

_hub_model.forward(net, x)

OpenPhenom’s mean patch token; it rescales 0-255 itself.

ChAda-ViT’s class token, each crop’s channels as one sequence.

spacr/embeddings.py:907 spacr/embeddings.py:914

_mil_fit._Attention.__init__(self)

Build the cell projection, attention gate and binary well head.

spacr/embeddings.py:1565

_mil_fit._Attention.forward(self, x, mask)

Return well logits, attention and cell evidence, excluding padding.

spacr/embeddings.py:1575

_subcell_model.Pooled.__init__(self)

Build the ViT and the pooler the checkpoint expects.

spacr/embeddings.py:955

_subcell_model.Pooled.forward(self, x)

(n, k, h, w) to (n, 1536), with the two heads joined.

The checkpoint determines whether k is two or four channels.

spacr/embeddings.py:963

_subcell_model.forward(net, x)

Min-max scale the expected model planes, then encode.

spacr/embeddings.py:980

_timm_encoder.run(stack: np.ndarray) → np.ndarray

Encode one three-channel stack, in batches, under no_grad.

Batched here rather than by the caller because the batch size belongs to the DEVICE, not to the channel policy: the per-channel policy calls this once per channel and each call must fit the same GPU.

Parameters:

stack – (n, height, width, 3) float32 in [0, 1].

Returns:

(n, dims) float32 features.

spacr/embeddings.py:657