spacr.ops_merge

Match cells between two magnifications by their arrangement, not their pixels.

THE PROBLEM. Optical pooled screening images the same well twice: once at low magnification to read the barcodes, once at high magnification to measure the phenotype. Every cell must be matched between the two, and the obvious method – correlate the images – does not work. The pixels do not correspond: the scale differs, the stains differ, and often the microscope differs too.

THE METHOD. What survives all of that is the ARRANGEMENT of the cells. Take three neighbouring cells and the SHAPE of the triangle they form is unchanged by magnification, rotation and translation. So the cells are triangulated, each triangle is reduced to a descriptor that depends only on its shape, and the two sets are matched on those descriptors. The transform between the acquisitions falls out of the matched triangles, and the cells are paired under it.

WHY THE PAIRING REFUSES AMBIGUITY. Matching is MUTUAL nearest neighbour with a distance limit: a pair is kept only when each cell is the other’s nearest, and anything doubtful is dropped rather than assigned. A wrongly matched cell gives a real phenotype the wrong barcode, which silently corrupts a screen; a dropped one costs only statistical power, which more cells can buy back.

WHERE IT STOPS WORKING, measured on 90 cells over a 500 px field related by a known scale of 2.0, an 11 degree rotation and a (137.5, -92.25) shift:

centroid jitter cells dropped scale angle shift error none none 2.000 11.00 0.00 px none 8 2.000 11.00 0.00 px 1.5 px none 1.995 10.77 3.45 px 1.5 px 8 1.993 10.71 4.45 px 4.0 px 25 2.050 8.83 41.29 px

The last row is a failure and is recorded as one. Triangle shape is what carries the signal, and jitter of a few pixels on a cell a few tens of pixels across changes that shape materially – so the method degrades with centroid noise rather than with cell count. If a real acquisition sits nearer the last row than the fourth, the answer is better segmentation centroids, not a looser tolerance: widening the tolerance admits more coincidental shape matches and moves the median onto them.

Follows brieflow (Cheeseman lab; github.com/cheeseman-lab/brieflow, MIT, Copyright 2025 Massachusetts Institute of Technology), whose merge stage was read for the approach.

Functions

align_by_triangles(source, target, *[, tolerance, ...])

Recover the transform between two cell sets, from triangle shapes alone.

match_cells(→ numpy.ndarray)

Pair cells one-to-one, keeping only mutual nearest neighbours.

triangle_descriptors(→ Tuple[numpy.ndarray, numpy.ndarray])

Delaunay-triangulate points and describe each triangle by its shape.

Module Contents

spacr.ops_merge.align_by_triangles(source: numpy.ndarray, target: numpy.ndarray, *, tolerance: float = 0.02, min_votes: int = 3)[source]

Recover the transform between two cell sets, from triangle shapes alone.

Every triangle in source is matched to the target triangles whose shape descriptor is within tolerance, each match proposes a transform, and the proposals are pooled. Pooling is what makes this robust: a single coincidental shape match proposes a wrong transform, but wrong proposals disagree with each other while right ones agree, so the median survives and the outliers do not.

Parameters:
  • source – (N, 2) centroids in the frame to be moved.

  • target – (M, 2) centroids in the frame to move to.

  • tolerance – how close two shape descriptors must be to be considered the same triangle.

  • min_votes – how many agreeing triangle matches are required before a transform is returned at all. Below this the answer is not evidence.

Returns:

(scale, rotation, translation), or None if too few triangles agreed.

spacr.ops_merge.match_cells(source: numpy.ndarray, target: numpy.ndarray, *, transform=None, threshold: float = 10.0, gpu: bool = True) → numpy.ndarray[source]

Pair cells one-to-one, keeping only mutual nearest neighbours.

AMBIGUITY IS DROPPED, NOT ASSIGNED. A pair survives only when each cell is the other’s nearest AND they are within threshold. Two cells competing for the same partner both lose. That asymmetry is deliberate: a wrongly matched cell hands a real phenotype somebody else’s barcode and quietly corrupts every statistic computed from it, whereas a dropped cell costs only statistical power that more cells can recover.

Parameters:
  • source – (N, 2) centroids.

  • target – (M, 2) centroids.

  • transform – (scale, rotation, translation) from align_by_triangles(), applied to source first. None means the two sets are already in the same frame.

  • threshold – the largest distance, in TARGET units, that may still be called the same cell.

  • gpu – run the two nearest-neighbour searches on the card where there is a usable one. A well carries tens of thousands of nuclei on each side and the search is done twice, once in each direction; spacr.ops_accel.nearest_neighbours chunks it so the working set does not grow with the well.

Returns:

(K, 2) array of (source_index, target_index) pairs.

spacr.ops_merge.triangle_descriptors(points: numpy.ndarray) → Tuple[numpy.ndarray, numpy.ndarray][source]

Delaunay-triangulate points and describe each triangle by its shape.

THE DESCRIPTOR IS TWO NUMBERS, and it has to be: the ratios of the two shorter sides to the longest. Scaling a triangle multiplies every side by the same factor, so the ratios do not move; rotating and translating it does not change side lengths at all. That is exactly the invariance the problem needs, and nothing about the absolute size or position survives – which is the point, because those are what differ between the two acquisitions.

Parameters:

points – (N, 2) cell centroids.

Returns:

(descriptors, vertex_indices) – an (M, 2) array and the (M, 3) indices of each triangle’s vertices, ordered so that vertex 0 is opposite the shortest side and vertex 2 opposite the longest. That ordering is what makes two matched descriptors also give matched POINTS, without which the transform cannot be solved.