spacr.ops_accel¶
The three array operations OPS is made of, on whatever hardware there is.
The hardware question has a short answer, and it names what to put on the card: “the per-window affine warp; peak finding and the unmixing matrix multiply; cellpose segmentation, which already has a GPU path” – and, for the merge, “ORB detect + Hamming descriptor matching” and “match cells: mutual nearest neighbour”. Strip the names away and that is three primitives: a windowed maximum, a matrix multiply, and a nearest-neighbour search. They are collected here rather than spelled out at each call site so there is ONE place that knows about devices, and the modules that use them go on reading as the science they are.
BACKEND ORDER: PyTorch first, CuPy second, NumPy always. Torch is already a spaCR dependency through Cellpose, so nothing new is installed, and one code path reaches CUDA, ROCm and Apple MPS.
EVERY FUNCTION TAKES AND RETURNS NUMPY, so no caller learns which backend ran and no caller has to change when one is added. Every one has a test that runs BOTH paths on the same input and asserts they agree – not “the GPU version passes its own test”, which is how a fallback quietly becomes a second implementation nobody compares.
AND IT NEVER ASSUMES THE CARD IS FREE. A shared GPU is the common case –
this one also runs structure prediction – so ops_gpu=False refuses a
device that exists, and a backend that raises at runtime (an
out-of-memory, a driver that went away) falls through to the next rather
than failing the plate.
THE CPU PATH IS NOT AN ERROR PATH. It is what runs on a laptop, in CI, and on the machine whose card is busy, which between them is most of the time this code will ever run.
Functions¶
|
Which backends can run here, best first, always ending in "numpy". |
|
|
|
The maximum of each |
|
For each row of |
Module Contents¶
- spacr.ops_accel.accelerated_backends(gpu: bool = True) Tuple[str, ...][source]¶
Which backends can run here, best first, always ending in “numpy”.
- Parameters:
gpu – False refuses the accelerated ones outright.
- Returns:
the usable names.
- spacr.ops_accel.matmul(first: numpy.ndarray, second: numpy.ndarray, *, gpu: bool = True, backend: str | None = None) numpy.ndarray[source]¶
first @ second, on the fastest thing available.THE UNMIXING IS THIS. Correcting dye bleed-through is one small matrix applied to every spot of every cycle, which is a large multiply by a tiny operand – the shape that goes to a card well and the shape a Python loop would be absurd for.
- Parameters:
first – left operand.
second – right operand.
gpu – False keeps it on the CPU.
backend – force one.
- Returns:
the product, as float32 NumPy.
- spacr.ops_accel.maximum_filter(field: numpy.ndarray, size: int, *, gpu: bool = True, backend: str | None = None) numpy.ndarray[source]¶
The maximum of each
size x sizewindow, centred on every pixel.WHAT PEAK FINDING IS MADE OF. A read is a local maximum of the spot score, and “local” means this. On a well-sized field it is also the most expensive thing in the decode chain, and it is a max-pool – which is the operation a GPU exists to do.
Edges repeat the border pixel, matching
scipy.ndimage.maximum_filterwithmode="nearest", because the alternative is a rim of false maxima all the way round the field.- Parameters:
field – a 2-D array.
size – the window edge, in pixels. Even sizes are raised to the next odd one: a window with no centre cannot be centred.
gpu – False keeps it on the CPU.
backend – force one.
- Returns:
the filtered field, same shape and dtype-compatible.
- spacr.ops_accel.nearest_neighbours(source: numpy.ndarray, target: numpy.ndarray, *, gpu: bool = True, backend: str | None = None, chunk: int = CHUNK) Tuple[numpy.ndarray, numpy.ndarray][source]¶
For each row of
source, the closest row oftarget.CHUNKED, ALWAYS. A well can carry a hundred thousand nuclei, and the full distance matrix between two such sets is 40 GB – more than the card and more than the host. Chunking the first set bounds the working set at
chunk x Mhowever large the well is, which is the same argument PART 5 makes about the mosaic: the ceiling must not depend on how much data there is.Ties go to the lower index, on every backend, so a caller that pairs mutual nearest neighbours gets the same pairs whatever ran.
- Parameters:
source –
(N, D)points.target –
(M, D)points.gpu – False keeps it on the CPU.
backend – force one.
chunk – rows of
sourceper pass.
- Returns:
(indices, distances), both length N.
Nested helpers¶
- matmul.run(name)¶
The matrix product on one named backend.
- Parameters:
name – “torch”, “cupy” or “numpy”.
- Returns:
the product as float32 NumPy, wherever it was computed.
- Raises:
RuntimeError – when that backend is not installed.
spacr/ops_accel.py:221
- maximum_filter.run(name)¶
The windowed maximum on one named backend.
- Parameters:
name – “torch”, “cupy” or “numpy”.
- Returns:
the filtered field as a 2-D array.
- Raises:
RuntimeError – when that backend is not installed, which is what
_tryreads to fall through to the next one.
spacr/ops_accel.py:166
- nearest_neighbours.run(name)¶
Nearest neighbours on one named backend.
- Parameters:
name – “torch”, “cupy” or “numpy”.
- Returns:
(indices, distances), one row per source point.- Raises:
RuntimeError – when that backend is not installed.
Chunked at
CHUNKon every backend, because the pairwise distance matrix is the memory cost here and it is quadratic.spacr/ops_accel.py:281