spacr.embeddings¶
A label-free embedding for every object, beside the measured features.
WHY THIS EXISTS. spaCR’s two existing ways of turning an image into numbers
both require you to already know what you are looking for. The measured panel
in spacr.measure computes a fixed, hand-designed set – intensity,
shape, texture, distances – so it finds only what somebody thought to name.
The supervised path in spacr.deep_spacr fine-tunes a backbone against
annotations, so it finds only what somebody thought to annotate. A screen
whose real hit is “the vacuole is subtly mislocalised in a way we have no
column for” returns nothing, and returns it confidently. The cost of that gap
is invisible, which is exactly why it is worth a module.
THE CHANNEL DECISION, SETTLED HERE RATHER THAN IMPLICITLY. Pretrained encoders expect three channels. spaCR images routinely have four or five, with no RGB meaning at all – a DAPI channel is not “red”. Two ways out, and the choice is recorded here rather than inherited from whatever the first backbone happened to want:
PER-CHANNEL, THEN CONCATENATE (
CHANNEL_PER_CHANNEL, the default). Each channel is encoded on its own and the vectors are joined. Channel identity SURVIVES: dimension 517 belongs to channel 2 and to nothing else, sospacr.attribution_columnscan say which stain a hit came from and a reader can drop a channel without re-embedding. Costs N forward passes.PROJECT TO THREE (
CHANNEL_PROJECT). One pass over a fixed channel-to-RGB mapping. Cheaper, and irreversible: once channels are mixed, no downstream column can be attributed to a stain, and a screen cannot ask “was this the parasite channel or the host one?”.
The default is per-channel because THE ATTRIBUTION IS THE POINT. An embedding that finds an unnamed phenotype and cannot say which stain carried it has moved the problem rather than solved it. Projection stays available for the case where forward passes actually bind.
WHAT THIS MODULE DOES NOT DO. It does not crop – it takes the crops
spacr.crops already produces. It does not write to the database. It
returns a float32 matrix keyed by object id, which is exactly the shape the
downstream already consumes: UMAP, the regression family, hit calling and
the AnnData export do not care where a numeric column came from.
Exceptions¶
Raised when embedding is refused, with the reason a caller can show. |
Classes¶
The matrix, its column names, and how it was made. |
|
How to embed, recorded so a matrix can say where it came from. |
Functions¶
|
Which channel a column came from, or |
|
Embed a stack of crops. |
|
Names for the columns a run produces. |
|
Describe this encoder as a |
|
The stable id for one encoder configuration. |
Module Contents¶
- exception spacr.embeddings.EmbeddingError[source]¶
Bases:
ValueErrorRaised when embedding is refused, with the reason a caller can show.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.embeddings.EmbeddingResult[source]¶
The matrix, its column names, and how it was made.
- Parameters:
values –
(n_objects, n_dims)float32.columns – one name per column, all starting with
EMBEDDING_PREFIX.spec – the
EmbeddingSpecthat produced it.n_channels – how many channels were encoded.
per_channel_dims – dimensions contributed by ONE channel.
- to_frame(object_ids: Sequence[Any])[source]¶
A DataFrame keyed by object id, ready to join to measurements.
- Parameters:
object_ids – one id per row, in the order the crops were given.
- Raises:
EmbeddingError – when the count does not match the matrix, because a silent misalignment here corrupts every downstream join and is unrecoverable afterwards.
- class spacr.embeddings.EmbeddingSpec[source]¶
How to embed, recorded so a matrix can say where it came from.
For
subcell_rybg, chooseCHANNEL_PROJECT, supply four explicit, distinctchannelsin microtubules (R), ER (Y), DNA (B), protein (G) order, and setnormalize=False. Crops must be at least 16 pixels in each spatial dimension. Native dimensions are preserved; the encoder applies the authors’ whole-crop min-max normalization.- Parameters:
backbone – Name of a timm or registered foundation encoder.
channel_policy –
CHANNEL_PER_CHANNELorCHANNEL_PROJECT.channels – which channel indices to encode, in order.
Nonemeans every channel the array has.batch_size – crops per forward pass.
device – torch device string, or
Noneto choose one.normalize – divide each channel by a fixed scale before encoding, clipped to [0, 1]. Microscopy dynamic range varies by orders of magnitude between stains, and an encoder trained on photographs will otherwise see one channel as noise.
channel_scale – that scale, one value per encoded channel in encoding order – normally each channel’s 99th percentile over a random sample of the plate. The same numbers for every crop mean an object’s embedding does not depend on which crops share its batch.
Noneletsembed_array()estimate it from the crops it is given.cell_dino_factory – pinned official Cell-DINO hub factory.
checkpoint_path – local official Cell-DINO weights file.
checkpoint_sha256 – expected SHA-256 of those weights.
- __post_init__() None[source]¶
Reject a specification that could not produce a matrix.
The check is made where the spec is built rather than where it is used, because one spec encodes every crop of a run: an unknown channel policy or a batch size below one is a typing mistake, and discovering it after the images are loaded wastes the load.
- Raises:
EmbeddingError – when
channel_policyis not one ofCHANNEL_POLICIES,batch_sizeis below one, orchannel_scaleis set withoutnormalize, holds a negative or non-finite value, or does not give one value per channel inchannels.
- spacr.embeddings.channel_of_column(column: str) int | None[source]¶
Which channel a column came from, or
Noneif it cannot be said.Nonefor a projected column is the honest answer rather than a guess: the channels were mixed and no dimension belongs to one stain.- Parameters:
column – an embedding column name.
- spacr.embeddings.embed_array(crops: numpy.ndarray, spec: EmbeddingSpec | None = None, *, encoder: Callable[[numpy.ndarray], numpy.ndarray] | None = None) EmbeddingResult[source]¶
Embed a stack of crops.
Every crop is divided by the same per-channel scale,
spec.channel_scale, so a crop’s embedding does not depend on the other crops incrops. When the spec carries none, one is estimated from a random sample ofcrops– the stack is treated as the whole plate – and a warning is logged. To embed one plate in several calls, pass the same scale to each: the returnedspeccarries it.- Parameters:
crops –
(n, height, width, channels), the shapespacr.cropsalready produces.spec – how to embed; the default is per-channel resnet18.
encoder – a callable taking
(n, height, width, 3)float32 in [0, 1] and returning(n, dims). An encoder with anin_channelsattribute takes that many planes instead, and one whosein_channelsisNonetakes any number: each channel alone under the per-channel policy, every encoded channel in one pass under the projection policy. Injected by the tests so the wiring can be exercised without downloading a backbone; production callers leave itNone.
- Returns:
an
EmbeddingResult.- Raises:
EmbeddingError – on a shape or channel problem, before any model is loaded – a download is a slow way to find out the array was wrong.
- spacr.embeddings.embedding_column_names(per_channel_dims: int, channels: Sequence[int], policy: str = CHANNEL_PER_CHANNEL) Tuple[str, ...][source]¶
Names for the columns a run produces.
THE CHANNEL IS IN THE NAME under the per-channel policy –
emb_c2_017– which is what makes attribution possible later without carrying a separate map beside the matrix. Under projection there is no channel to name, and the name says so rather than implying one.- Parameters:
per_channel_dims – dimensions one forward pass produces.
channels – channel indices, in encoding order.
policy – which channel policy produced them.
- Returns:
column names, in matrix order.
- spacr.embeddings.encoder_entry(spec: EmbeddingSpec | None = None, *, scorecard: Mapping[str, Any] | None = None)[source]¶
Describe this encoder as a
spacr.model_zoo.ModelEntry.The entry combines encoder configuration, available checkpoint provenance and optional retrieval metrics. Compare entries only when their channel policies are compatible.
The entry records the backbone, channel policy and digest of readable local checkpoint bytes without downloading or loading a model. Public backbones use timm cache metadata; OpenPhenom and ChAda-ViT use pinned Hugging Face revisions; SubCell uses its official torch hub checkpoint; a local DINO encoder uses its configured checkpoint path. Cell-DINO uses the declared local checkpoint only when stable regular-file bytes match its expected SHA-256; changed or mismatched files leave no digest. Foundation entries name the original provider and checkpoint URL. A checksum alone does not verify a model or establish its training-data provenance.
If no readable local checkpoint is found, the entry keeps an empty digest and explains this in its notes.
spacr.model_zoo.fetch()refuses entries without a digest.- Parameters:
spec – Encoder configuration; uses
EmbeddingSpec()when omitted.scorecard – Optional retrieval metrics measured on labelled controls, stored in
ModelEntry.metricsfor display byscorecard_lines.
- Returns:
A
ModelEntrywith kind'encoder'.
- spacr.embeddings.encoder_key(spec: EmbeddingSpec) str[source]¶
The stable id for one encoder configuration.
BACKBONE AND POLICY TOGETHER, because they are jointly what makes two runs’ dimensions mean the same thing. The same backbone under
CHANNEL_PER_CHANNELandCHANNEL_PROJECTproduces columns that are the same width, the same dtype and not remotely the same quantity – one is per-stain, the other is a mixture. A key that named only the backbone would let those two be compared silently.- Parameters:
spec – the run’s configuration. Only
backboneandchannel_policyreach the key: everything else on the spec – batch size, device, crop size – changes how long the run takes and not what a dimension means.
Nested helpers¶
- _cell_dino_model.forward(net, x)¶
Return the official DINOv2 head output for each declared crop.
spacr/embeddings.py:883
- _dino_encoder.run(stack: np.ndarray) np.ndarray¶
Encode
(n, h, w, k)crops in batches at the training size.- Parameters:
stack – float32 in [0, 1], channels last.
- Returns:
(n, dims)float32 features.
spacr/embeddings.py:1382
- _dino_network.Head.__init__(self)¶
Build the MLP and the prototype layer.
spacr/embeddings.py:1177
- _dino_network.Head.forward(self, x)¶
(n, dim)features to(n, out_dim)prototype scores.spacr/embeddings.py:1185
- _dino_views.uniform(low, high, *shape)¶
Uniform draws from the run’s generator, on the batch’s device.
spacr/embeddings.py:1126
- _foundation_encoder.run(stack: np.ndarray) np.ndarray¶
Encode declared Cell-DINO planes at the official crop size.
Encode
(n, h, w, k)crops in model-appropriate batches.- Parameters:
stack – float32 crops, channels last. SubCell R/Y/B/G uses native intensity values; other foundations expect [0, 1].
- Returns:
(n, dims)float32 features.
spacr/embeddings.py:783spacr/embeddings.py:809
- _hub_model.forward(net, x)¶
OpenPhenom’s mean patch token; it rescales 0-255 itself.
ChAda-ViT’s class token, each crop’s channels as one sequence.
spacr/embeddings.py:907spacr/embeddings.py:914
- _mil_fit._Attention.__init__(self)¶
Build the cell projection, attention gate and binary well head.
spacr/embeddings.py:1565
- _mil_fit._Attention.forward(self, x, mask)¶
Return well logits, attention and cell evidence, excluding padding.
spacr/embeddings.py:1575
- _subcell_model.Pooled.__init__(self)¶
Build the ViT and the pooler the checkpoint expects.
spacr/embeddings.py:955
- _subcell_model.Pooled.forward(self, x)¶
(n, k, h, w)to(n, 1536), with the two heads joined.The checkpoint determines whether
kis two or four channels.spacr/embeddings.py:963
- _subcell_model.forward(net, x)¶
Min-max scale the expected model planes, then encode.
spacr/embeddings.py:980
- _timm_encoder.run(stack: np.ndarray) np.ndarray¶
Encode one three-channel stack, in batches, under no_grad.
Batched here rather than by the caller because the batch size belongs to the DEVICE, not to the channel policy: the per-channel policy calls this once per channel and each call must fit the same GPU.
- Parameters:
stack –
(n, height, width, 3)float32 in [0, 1].- Returns:
(n, dims)float32 features.
spacr/embeddings.py:657