spacr.accelerator¶
One answer to “what should this run on, and what is it called”.
spaCR grew up on NVIDIA, so roughly forty-five call sites
ask torch.cuda.is_available() and most of them mean “is there a GPU?”.
Those are DIFFERENT QUESTIONS on every machine that is not NVIDIA, and
conflating them is what makes a perfectly good card read as “no GPU”:
ROCm answers
torch.cuda.is_available()True and wantsdevice="cuda". Code that then prints “CUDA” is naming a vendor that is not there; code that prints “no NVIDIA GPU” is denying one that is.Apple’s Metal answers it False and wants
device="mps". MPS is Metal, not an Apple-Silicon feature. On the Intel iMac this was written on it drives an AMD Radeon Pro 5300 – measured 16x on conv2d inference and 2.7x on a training step against the same machine’s CPU. ROCm has no macOS build at all, so Metal is the ONLY route to an AMD card on a Mac.
So this module answers both questions separately, and a third that matters more than either: what is this device unable to do? Re-pointing a tensor at MPS is not enough, because three things fail rather than degrade, and each was measured rather than assumed:
float64raisesTypeErroroutright – not a slow path, a hard stop. Anything wanting double precision must stay on the CPU.autocast(device_type="mps")raisesRuntimeError. Mixed precision has to be branched off, not merely re-pointed.Some operators are simply missing (
aten::linalg_qr.outamong them).PYTORCH_ENABLE_MPS_FALLBACK=1turns those from a crash into a quiet CPU detour, so this module sets it when it selects MPS.
DETECTED AND USABLE ARE REPORTED SEPARATELY. A setup screen announcing an accelerator spaCR will never dispatch to is worse than one that says nothing, because the user then blames their hardware for CPU speed. Neural engines are the whole reason that distinction exists here: the Apple Neural Engine, Intel AI Boost and Qualcomm Hexagon are all real silicon with no portable torch device, so they are named as FOUND and never selected.
NOTHING HERE RAISES. A half-installed ROCm, a broken driver, a torch built without a backend it advertises – all resolve to the CPU with a note. The machines most likely to have a strange accelerator are the least able to afford a traceback at import time.
Classes¶
What spaCR will compute on, and what it can be told about it. |
Functions¶
|
The string for |
|
|
|
What to pass cellpose as |
|
Consistent device and weight precision for |
|
One line for a log or a console: what was found and whether it runs. |
|
The resolved device as a string, e.g. |
|
Hand the driver back whatever this backend caches. Never raises. |
|
Resolve against a SPECIFIC torch module, without touching the cache. |
|
Is the resolved accelerator genuinely NVIDIA CUDA. |
|
Is there a usable accelerator of any vendor. |
|
Inference silicon that is present and that spaCR will NOT use. |
|
The accelerator spaCR will use. Cached; never raises. |
|
Whether |
|
Whether |
|
Whether double precision may be sent to the device. |
|
Module Contents¶
- class spacr.accelerator.Accelerator[source]¶
What spaCR will compute on, and what it can be told about it.
- Parameters:
kind – one of
KINDS.rocmis reported distinctly fromcudaeven though both dispatch todevice="cuda", because the only thing a user can act on is the true vendor.device – the string to hand
torch.device.label – human text for the setup slide and the doctor.
name – undecorated hardware marketing name shown to users, without the backend or version suffix carried by
label; empty when no device name was discovered.detected – the hardware is present.
usable – spaCR will actually dispatch to it. Never true unless
detected; false for accelerators with no torch device.note – why it is not usable, when it is not. Empty otherwise.
float64 – the device accepts double precision.
autocast –
torch.autocastaccepts this device type.fallback – missing operators silently run on the CPU instead of raising, because this module set the backend’s fallback flag.
bfloat16 – the device accepts
torch.bfloat16. PROBED rather than assumed from the backend name, because it is the one capability that moves with the torch version: Metal gained it after 2.2, and 2.2.2 is the last x86-64 macOS wheel, so the same backend answers differently on an Intel Mac and an Apple Silicon one. Cellpose’s cpsam loads its weights in bfloat16 by default, so a wrong answer here is a TypeError at model construction.
- property is_cuda: bool[source]¶
Is this genuinely NVIDIA CUDA.
Narrower than
is_gpuand deliberately so: anything readingnvidia-smi, a CUDA version, or NVIDIA-specific memory interfaces wants this one and would be wrong on ROCm.
- spacr.accelerator.autocast_device_type() str | None[source]¶
The string for
torch.autocast(device_type=...), or None.None means “do not use mixed precision here”, which is a different answer from “use it on the CPU” and has to stay distinguishable.
- spacr.accelerator.capabilities(found: Accelerator | None = None) Tuple[Tuple[str, bool, str], ...][source]¶
(task, accelerated, detail)for what this machine can actually do.WHAT THE SETUP SCREEN IS FOR. “Compatible GPU” on its own answers a question nobody asked: users want to know whether THE SLOW STEP will be slow. So this reports per task, and it reports the truth per backend rather than one verdict for all of them – on Metal the segmentation and the classifier are accelerated while the cuML reductions are not, and a single green tick would be a lie about the second.
Ordered by how much the acceleration is worth: segmentation on the CPU took 444 s for one 256x256 image on the machine this was written on, and 3.2 s on its Radeon.
- Parameters:
found – optional existing accelerator snapshot. Providing it avoids another resolution and any dtype-probe allocations, as required by the doctor’s metadata-only check.
- Returns:
task, accelerated flag and explanation for each capability.
- spacr.accelerator.cellpose_gpu() bool[source]¶
What to pass cellpose as
gpu=.Cellpose does its own device resolution and already knows about MPS –
assign_device(gpu=True)answersmpson a Metal machine. What it cannot do is guess, so it must be TOLD there is a GPU. Passingtorch.cuda.is_available()here pins a Mac to the CPU when no explicit device is supplied. An explicit Cellpose device takes precedence.
- spacr.accelerator.cellpose_kwargs() dict[source]¶
Consistent device and weight precision for
CellposeModelinference.THREE ARGUMENTS THAT HAVE TO AGREE, which is why they are produced together rather than spelled out at six call sites:
gpu– controls Cellpose’s selection when no explicit device is supplied, including callers that deliberately dropdevicebelow.device– from the one resolver, so cellpose and spaCR cannot disagree about the same machine; an explicit device takes precedence.use_bfloat16– cpsam loads its weights in bfloat16 by default and Metal on torch 2.2 has no bfloat16, so the default is aTypeError: BFloat16 is not supported on MPSat construction. Measured on the reporting iMac; float32 weights work there and cost VRAM, which is the right trade for a card that otherwise sits idle. CPU inference also uses float32: bfloat16 support does not imply native arithmetic, and emulation can be substantially slower. This matches Make Masks; float32 arithmetic need not produce identical predictions to bfloat16.
Callers that pass
device=Noneon purpose – letting cellpose resolve it – should takegpuanduse_bfloat16from here and dropdevice.
- spacr.accelerator.describe() str[source]¶
One line for a log or a console: what was found and whether it runs.
- spacr.accelerator.device_string() str[source]¶
The resolved device as a string, e.g.
"cuda:0","mps".
- spacr.accelerator.empty_cache(torch_module=None) str[source]¶
Hand the driver back whatever this backend caches. Never raises.
- Parameters:
torch_module – ask about THIS torch instead of the cached answer for the machine.
spacr.qt.resource_cleanupholds its own handle – it deliberately does not import torch just to free memory – and a cached global would send the call to the wrong backend when that handle is a stand-in.
- spacr.accelerator.inspect_torch(torch, *, device_names=True, include_cuda=True) Accelerator[source]¶
Resolve against a SPECIFIC torch module, without touching the cache.
For callers that already hold a torch handle and must be answered about that one –
spacr.doctor.check_gpu()is the case: it imports torch through its own indirection so the diagnosis can be exercised against a stand-in, and a cached answer about the real machine would defeat that entirely. Same probes and same order asresolve(), so the two cannot drift.- Parameters:
torch – torch module whose availability metadata should be inspected.
device_names – False avoids CUDA/XPU property queries that can initialize their runtime. Availability queries still run.
include_cuda – False skips CUDA/ROCm entirely, for a process whose environment deliberately hides those devices. Other backends remain eligible. This function does not allocate dtype-probe tensors.
- Returns:
accelerator described by the supplied torch module.
- spacr.accelerator.is_cuda() bool[source]¶
Is the resolved accelerator genuinely NVIDIA CUDA.
For the places that legitimately need CUDA specifically – memory interrogation,
nvidia-smiadvice, CUDA version reporting.
- spacr.accelerator.is_gpu() bool[source]¶
Is there a usable accelerator of any vendor.
What a call site means when it writes
torch.cuda.is_available()to decide whether to use a GPU at all.
- spacr.accelerator.neural_engines() Tuple[str, ...][source]¶
Inference silicon that is present and that spaCR will NOT use.
Reported so the setup slide can say “found, not used” rather than leaving a user to wonder why their Neural Engine is idle. There is no portable torch device for any of these – the ANE is reachable only through CoreML, Intel’s AI Boost only through OpenVINO – so naming one as a compute device would be a promise nothing behind it can keep.
- spacr.accelerator.resolve(refresh: bool = False) Accelerator[source]¶
The accelerator spaCR will use. Cached; never raises.
- Parameters:
refresh – probe again instead of answering from the cache. For tests, which fake the probes.
- spacr.accelerator.supports_autocast() bool[source]¶
Whether
torch.autocastaccepts this device type.
- spacr.accelerator.supports_bfloat16() bool[source]¶
Whether
torch.bfloat16tensors may be put on the device.
Nested helpers¶
- _keep_cellpose_flows_off_metal.on_the_cpu(masks, flows, threshold=0.4, device=None)¶
Run Cellpose flow-error filtering on the CPU for Metal.
- Parameters:
masks – label masks forwarded unchanged to Cellpose.
flows – flow field forwarded unchanged to Cellpose.
threshold – flow-error threshold forwarded to Cellpose.
device – ignored compatibility argument; the wrapped call always receives
torch.device("cpu").
- Returns:
the original
remove_bad_flow_masksreturn value.
spacr/accelerator.py:736
- _measure_dtypes.accepts(name: str) bool | None¶
Probe whether the resolved device accepts one named torch dtype.
- Parameters:
name – torch dtype attribute name, such as
"float64"or"bfloat16".- Returns:
Truewhen a two-element tensor allocates on the device,Falsefor the errors used by unsupported dtypes, orNonewhen torch lacks the dtype or the probe fails unexpectedly.
spacr/accelerator.py:466