spacr.accelerator

One answer to “what should this run on, and what is it called”.

spaCR grew up on NVIDIA, so roughly forty-five call sites ask torch.cuda.is_available() and most of them mean “is there a GPU?”. Those are DIFFERENT QUESTIONS on every machine that is not NVIDIA, and conflating them is what makes a perfectly good card read as “no GPU”:

  • ROCm answers torch.cuda.is_available() True and wants device="cuda". Code that then prints “CUDA” is naming a vendor that is not there; code that prints “no NVIDIA GPU” is denying one that is.

  • Apple’s Metal answers it False and wants device="mps". MPS is Metal, not an Apple-Silicon feature. On the Intel iMac this was written on it drives an AMD Radeon Pro 5300 – measured 16x on conv2d inference and 2.7x on a training step against the same machine’s CPU. ROCm has no macOS build at all, so Metal is the ONLY route to an AMD card on a Mac.

So this module answers both questions separately, and a third that matters more than either: what is this device unable to do? Re-pointing a tensor at MPS is not enough, because three things fail rather than degrade, and each was measured rather than assumed:

  • float64 raises TypeError outright – not a slow path, a hard stop. Anything wanting double precision must stay on the CPU.

  • autocast(device_type="mps") raises RuntimeError. Mixed precision has to be branched off, not merely re-pointed.

  • Some operators are simply missing (aten::linalg_qr.out among them). PYTORCH_ENABLE_MPS_FALLBACK=1 turns those from a crash into a quiet CPU detour, so this module sets it when it selects MPS.

DETECTED AND USABLE ARE REPORTED SEPARATELY. A setup screen announcing an accelerator spaCR will never dispatch to is worse than one that says nothing, because the user then blames their hardware for CPU speed. Neural engines are the whole reason that distinction exists here: the Apple Neural Engine, Intel AI Boost and Qualcomm Hexagon are all real silicon with no portable torch device, so they are named as FOUND and never selected.

NOTHING HERE RAISES. A half-installed ROCm, a broken driver, a torch built without a backend it advertises – all resolve to the CPU with a note. The machines most likely to have a strange accelerator are the least able to afford a traceback at import time.

Classes

Accelerator

What spaCR will compute on, and what it can be told about it.

Functions

autocast_device_type(→ Optional[str])

The string for torch.autocast(device_type=...), or None.

capabilities(→ Tuple[Tuple[str, bool, str], ...])

(task, accelerated, detail) for what this machine can actually do.

cellpose_gpu(→ bool)

What to pass cellpose as gpu=.

cellpose_kwargs(→ dict)

Consistent device and weight precision for CellposeModel inference.

describe(→ str)

One line for a log or a console: what was found and whether it runs.

device_string(→ str)

The resolved device as a string, e.g. "cuda:0", "mps".

empty_cache(→ str)

Hand the driver back whatever this backend caches. Never raises.

inspect_torch(→ Accelerator)

Resolve against a SPECIFIC torch module, without touching the cache.

is_cuda(→ bool)

Is the resolved accelerator genuinely NVIDIA CUDA.

is_gpu(→ bool)

Is there a usable accelerator of any vendor.

neural_engines(→ Tuple[str, ...])

Inference silicon that is present and that spaCR will NOT use.

resolve(→ Accelerator)

The accelerator spaCR will use. Cached; never raises.

supports_autocast(→ bool)

Whether torch.autocast accepts this device type.

supports_bfloat16(→ bool)

Whether torch.bfloat16 tensors may be put on the device.

supports_float64(→ bool)

Whether double precision may be sent to the device.

torch_device()

torch.device for the resolved accelerator.

Module Contents

class spacr.accelerator.Accelerator[source]

What spaCR will compute on, and what it can be told about it.

Parameters:
  • kind – one of KINDS. rocm is reported distinctly from cuda even though both dispatch to device="cuda", because the only thing a user can act on is the true vendor.

  • device – the string to hand torch.device.

  • label – human text for the setup slide and the doctor.

  • name – undecorated hardware marketing name shown to users, without the backend or version suffix carried by label; empty when no device name was discovered.

  • detected – the hardware is present.

  • usable – spaCR will actually dispatch to it. Never true unless detected; false for accelerators with no torch device.

  • note – why it is not usable, when it is not. Empty otherwise.

  • float64 – the device accepts double precision.

  • autocast – torch.autocast accepts this device type.

  • fallback – missing operators silently run on the CPU instead of raising, because this module set the backend’s fallback flag.

  • bfloat16 – the device accepts torch.bfloat16. PROBED rather than assumed from the backend name, because it is the one capability that moves with the torch version: Metal gained it after 2.2, and 2.2.2 is the last x86-64 macOS wheel, so the same backend answers differently on an Intel Mac and an Apple Silicon one. Cellpose’s cpsam loads its weights in bfloat16 by default, so a wrong answer here is a TypeError at model construction.

property is_cuda: bool[source]

Is this genuinely NVIDIA CUDA.

Narrower than is_gpu and deliberately so: anything reading nvidia-smi, a CUDA version, or NVIDIA-specific memory interfaces wants this one and would be wrong on ROCm.

property is_gpu: bool[source]

Is there a usable non-CPU device.

THE QUESTION MOST CALL SITES ACTUALLY MEANT when they wrote torch.cuda.is_available().

property torch_device[source]

A torch.device for device, or the CPU if torch is missing entirely.

spacr.accelerator.autocast_device_type() → str | None[source]

The string for torch.autocast(device_type=...), or None.

None means “do not use mixed precision here”, which is a different answer from “use it on the CPU” and has to stay distinguishable.

spacr.accelerator.capabilities(found: Accelerator | None = None) → Tuple[Tuple[str, bool, str], ...][source]

(task, accelerated, detail) for what this machine can actually do.

WHAT THE SETUP SCREEN IS FOR. “Compatible GPU” on its own answers a question nobody asked: users want to know whether THE SLOW STEP will be slow. So this reports per task, and it reports the truth per backend rather than one verdict for all of them – on Metal the segmentation and the classifier are accelerated while the cuML reductions are not, and a single green tick would be a lie about the second.

Ordered by how much the acceleration is worth: segmentation on the CPU took 444 s for one 256x256 image on the machine this was written on, and 3.2 s on its Radeon.

Parameters:

found – optional existing accelerator snapshot. Providing it avoids another resolution and any dtype-probe allocations, as required by the doctor’s metadata-only check.

Returns:

task, accelerated flag and explanation for each capability.

spacr.accelerator.cellpose_gpu() → bool[source]

What to pass cellpose as gpu=.

Cellpose does its own device resolution and already knows about MPS – assign_device(gpu=True) answers mps on a Metal machine. What it cannot do is guess, so it must be TOLD there is a GPU. Passing torch.cuda.is_available() here pins a Mac to the CPU when no explicit device is supplied. An explicit Cellpose device takes precedence.

spacr.accelerator.cellpose_kwargs() → dict[source]

Consistent device and weight precision for CellposeModel inference.

THREE ARGUMENTS THAT HAVE TO AGREE, which is why they are produced together rather than spelled out at six call sites:

  • gpu – controls Cellpose’s selection when no explicit device is supplied, including callers that deliberately drop device below.

  • device – from the one resolver, so cellpose and spaCR cannot disagree about the same machine; an explicit device takes precedence.

  • use_bfloat16 – cpsam loads its weights in bfloat16 by default and Metal on torch 2.2 has no bfloat16, so the default is a TypeError: BFloat16 is not supported on MPS at construction. Measured on the reporting iMac; float32 weights work there and cost VRAM, which is the right trade for a card that otherwise sits idle. CPU inference also uses float32: bfloat16 support does not imply native arithmetic, and emulation can be substantially slower. This matches Make Masks; float32 arithmetic need not produce identical predictions to bfloat16.

Callers that pass device=None on purpose – letting cellpose resolve it – should take gpu and use_bfloat16 from here and drop device.

spacr.accelerator.describe() → str[source]

One line for a log or a console: what was found and whether it runs.

spacr.accelerator.device_string() → str[source]

The resolved device as a string, e.g. "cuda:0", "mps".

spacr.accelerator.empty_cache(torch_module=None) → str[source]

Hand the driver back whatever this backend caches. Never raises.

Parameters:

torch_module – ask about THIS torch instead of the cached answer for the machine. spacr.qt.resource_cleanup holds its own handle – it deliberately does not import torch just to free memory – and a cached global would send the call to the wrong backend when that handle is a stand-in.

spacr.accelerator.inspect_torch(torch, *, device_names=True, include_cuda=True) → Accelerator[source]

Resolve against a SPECIFIC torch module, without touching the cache.

For callers that already hold a torch handle and must be answered about that one – spacr.doctor.check_gpu() is the case: it imports torch through its own indirection so the diagnosis can be exercised against a stand-in, and a cached answer about the real machine would defeat that entirely. Same probes and same order as resolve(), so the two cannot drift.

Parameters:
  • torch – torch module whose availability metadata should be inspected.

  • device_names – False avoids CUDA/XPU property queries that can initialize their runtime. Availability queries still run.

  • include_cuda – False skips CUDA/ROCm entirely, for a process whose environment deliberately hides those devices. Other backends remain eligible. This function does not allocate dtype-probe tensors.

Returns:

accelerator described by the supplied torch module.

spacr.accelerator.is_cuda() → bool[source]

Is the resolved accelerator genuinely NVIDIA CUDA.

For the places that legitimately need CUDA specifically – memory interrogation, nvidia-smi advice, CUDA version reporting.

spacr.accelerator.is_gpu() → bool[source]

Is there a usable accelerator of any vendor.

What a call site means when it writes torch.cuda.is_available() to decide whether to use a GPU at all.

spacr.accelerator.neural_engines() → Tuple[str, ...][source]

Inference silicon that is present and that spaCR will NOT use.

Reported so the setup slide can say “found, not used” rather than leaving a user to wonder why their Neural Engine is idle. There is no portable torch device for any of these – the ANE is reachable only through CoreML, Intel’s AI Boost only through OpenVINO – so naming one as a compute device would be a promise nothing behind it can keep.

spacr.accelerator.resolve(refresh: bool = False) → Accelerator[source]

The accelerator spaCR will use. Cached; never raises.

Parameters:

refresh – probe again instead of answering from the cache. For tests, which fake the probes.

spacr.accelerator.supports_autocast() → bool[source]

Whether torch.autocast accepts this device type.

spacr.accelerator.supports_bfloat16() → bool[source]

Whether torch.bfloat16 tensors may be put on the device.

spacr.accelerator.supports_float64() → bool[source]

Whether double precision may be sent to the device.

spacr.accelerator.torch_device()[source]

torch.device for the resolved accelerator.

The direct replacement for torch.device("cuda:0" if torch.cuda.is_available() else "cpu").

Nested helpers

_keep_cellpose_flows_off_metal.on_the_cpu(masks, flows, threshold=0.4, device=None)

Run Cellpose flow-error filtering on the CPU for Metal.

Parameters:
  • masks – label masks forwarded unchanged to Cellpose.

  • flows – flow field forwarded unchanged to Cellpose.

  • threshold – flow-error threshold forwarded to Cellpose.

  • device – ignored compatibility argument; the wrapped call always receives torch.device("cpu").

Returns:

the original remove_bad_flow_masks return value.

spacr/accelerator.py:736

_measure_dtypes.accepts(name: str) → bool | None

Probe whether the resolved device accepts one named torch dtype.

Parameters:

name – torch dtype attribute name, such as "float64" or "bfloat16".

Returns:

True when a two-element tensor allocates on the device, False for the errors used by unsupported dtypes, or None when torch lacks the dtype or the probe fails unexpectedly.

spacr/accelerator.py:466