spacr.model_zoo

Workflow inputs and outputs

Model Zoo

Inspect model provenance and download/install compatible checkpoints or backends. A listed backend is not itself a checkpoint file.

Open: Make Masks → Model Zoo.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Segmentation checkpoint — Saved Cellpose-compatible checkpoint or a compatible installed backend selected with its own configuration.

Outputs

  • Segmentation checkpoint — Saved Cellpose-compatible checkpoint or a compatible installed backend selected with its own configuration.

API reference.

Module tutorial.

Browse, verify, fetch and benchmark spaCR segmentation and classification models.

Why this exists

spaCR can run a stock Cellpose model, a Cellpose model somebody on the team fine-tuned last year, a classifier checkpoint from a run three folders up, or a checkpoint downloaded from Hugging Face. Today the only way to find out which of those exist on a machine is find / -name '*.pth', the only way to know what one of them was trained on is to remember, and the only way to know whether the copy on disk is the file the author published is to hope.

This module answers those three questions and nothing else:

  • what is here — discover_local() walks folders and returns the checkpoints, classified as Cellpose or classifier, with whatever provenance is recoverable from the settings snapshots spaCR already writes;

  • is it the right bytes — sha256_file() / verify() / fetch(), which downloads atomically, checksums what arrived, and refuses to install a mismatch;

  • what does it do on my data — benchmark(), which is spacr.model_compare’s “three fields” harness pointed at one model instead of two.

Nothing here imports torch or cellpose at module import time. Browsing the zoo, reading provenance, checksumming a file and rendering the table all work on a machine with neither installed; only benchmark() (through spacr.model_compare.segment_with_cellpose()) needs them, and only when it is called. tests/test_model_zoo.py asserts that.

The four things that make this trustworthy rather than merely convenient

Every download is checksummed, and verification happens before use.

A checkpoint truncated by a dropped connection, or swapped for a different one at the same URL, still loads. It does not raise; it produces silently different masks, and the run that used it looks exactly like the run that did not. So fetch() hashes what arrived and compares it to the hash the catalogue published: a mismatch deletes the file and raises ChecksumMismatch — the entry is never registered. The hash that was actually computed is stored on the returned entry, so “verified” means “these bytes”, not “this filename”.

A catalogue entry with no published hash cannot be verified at all, and that is refused by default (require_checksum=True) rather than quietly treated as fine. Callers that knowingly accept an unverifiable source pass require_checksum=False; the resulting entry carries verified=False and says so in format_zoo().

Every download is atomic, and never overwrites.

Bytes stream into a temporary file in the destination directory (so the final os.replace is a same-filesystem rename, which is atomic) and the rename happens only after the checksum passes. An interrupted or cancelled download therefore leaves nothing behind that looks like a model — the failure mode where half a checkpoint sits at the real filename, loads, and segments badly, cannot happen.

The destination is versioned rather than overwritten (versioned_path(): foo.CP_model, foo_v2.CP_model, …). Two models with the same filename are a normal thing to have; losing the first one to the second is not.

Provenance is recorded, and “unknown” is written out in full.

A Cellpose model fine-tuned on 60x confluent HeLa is not interchangeable with one trained on 20x sparse fibroblasts, and no amount of benchmark score makes it so. What a model was trained on is the single most useful thing the zoo can show, so ModelEntry carries trained_on and trained_by, both recovered from the settings snapshots spaCR already writes beside its models (<file>_settings.csv for a Cellpose model, <dst>/settings.csv for a classifier — read through spacr.train_compare.load_run(), which already knows where to look).

Where it could not be recovered the field reads 'unknown', never ''. A blank cell in a provenance table reads as “no constraints”; that is the opposite of what it means.

Benchmarks are only comparable inside one field set.

A model’s score on your three fields says nothing whatsoever about its score on somebody else’s, and a table that sorts the two together invents a ranking out of two unrelated numbers. So every BenchmarkResult records a fieldset_id() — a hash of the actual pixels, not the folder name — and rank() raises IncomparableBenchmarks when handed results from more than one field set. rank_groups() and format_benchmarks() are the supported alternative: they group by field set, rank within each group, and label the groups.

And a fifth, smaller one: a model file that is missing, empty, or not a torch checkpoint at all fails in inspect_checkpoint() with a message naming the file, before anything tries to load it. The default failure — a KeyError on a state-dict key from deep inside torch — names nothing the user chose.

What the benchmark can and cannot say

There is no ground truth here. benchmark() runs one model over N fields and reports what came out: object counts, timings and the spacr.seg_qc verdict per field (fused? shattered? empty? all on the border?). That is a quality-control score, not an accuracy — it can tell you a model collapsed on your data, it cannot tell you which of two plausible segmentations is right. RANK_KEYS therefore offers exactly two keys, 'qc' and 'seconds', and no key that would read as accuracy. To compare two models against each other, use compare_entries(), which hands both to spacr.model_compare.compare_models() — the A/B harness that is explicit about neither side being the truth.

Cellpose 4 accepts and ignores model_type, diam_mean, nchan, channels and rescale; only diameter at eval still changes the masks. BenchmarkResult carries the resolved honoured and ignored parameter dicts straight from spacr.model_compare.ModelConfig so a benchmark cannot silently be a benchmark of settings nothing read.

Example:

from spacr import model_zoo as zoo

entries = zoo.catalogue() + zoo.discover_local('/data/screen1')
print(zoo.format_zoo(entries))

entry = zoo.resolve('cpsam_plaque_r3', entries)
result = zoo.benchmark(entry, source='/data/screen1/plate1/1', n_fields=3)
print(zoo.format_benchmarks([result]))

See also

spacr.model_compare — the A/B harness this module reuses for segmentation and for the two-model comparison. spacr.train_compare — run discovery and settings recovery, reused wholesale for the classifier half of the zoo. spacr.utils.download_models() — the legacy bulk pull of the bundled Hugging Face model pack, wrapped by download_bundled_models().

Exceptions

ChecksumMismatch

What arrived is not what the catalogue published. Nothing was installed.

DownloadCancelled

The caller cancelled a fetch. Nothing was left at the destination.

IncomparableBenchmarks

Benchmarks from different field sets cannot be ranked against each other.

ModelUnreadable

A model file is missing, empty, or not a checkpoint. Names the file.

ModelZooError

Base class for every refusal in this module.

Classes

BenchmarkResult

One model over one field set. Only comparable to results on the same set.

FieldBenchmark

One field's result for one model.

ModelEntry

One model the zoo knows about, wherever it lives.

Functions

benchmark(→ BenchmarkResult)

Run one model over N fields and report what came out. "Test on 3 fields".

bioimageio_entries(→ List[ModelEntry])

Every Cellpose model bioimage.io publishes, as zoo rows.

catalogue(→ List[ModelEntry])

Everything the zoo knows about without scanning the user's disks.

classify_kind(→ Optional[str])

Say whether a file is a Cellpose model, a classifier, or not a model.

community_entries(→ List[ModelEntry])

Unvetted models uploaded by spaCR users, or an empty list.

compare_entries(entry_a, entry_b[, images, source, ...])

Put two zoo entries head to head on the same fields.

config_for(entry[, overrides])

The spacr.model_compare.ModelConfig that runs this entry.

default_local_roots(→ List[pathlib.Path])

Folders worth scanning when the caller has not named one.

discover_local(→ List[ModelEntry])

Find the model checkpoints already on this machine.

download_bundled_models(→ str)

Pull the bundled Hugging Face model pack via the existing downloader.

entries_from_sources(→ List[Any])

The rows belonging to the headings that are on, in the given order.

entry_from_file() → ModelEntry)

Build a ModelEntry for a checkpoint on this machine.

fetch(→ pathlib.Path)

Download a model, verify it, and only then put it where it belongs.

fieldset_id(→ str)

A stable id for a set of fields, taken from the pixels.

format_benchmarks(→ str)

Render benchmarks grouped by field set, ranked only within a group.

format_zoo(→ str)

Render a listing a human reads before choosing a model.

group_by_fieldset(→ Dict[str, List[BenchmarkResult]])

Bucket benchmarks by the field set they ran on, first-seen order.

group_by_source(→ Dict[str, List[Any]])

Split a listing into ZOO_SOURCES, keeping each source's order.

hf_uri(→ str)

The download URL for a file in a Hugging Face repo.

inspect_checkpoint(→ Dict[str, Any])

Check a file is a loadable checkpoint, failing with the filename in it.

install(→ ModelEntry)

fetch() the model and return the registered local entry.

installable_backend_entries(→ List[ModelEntry])

Every optional segmentation backend, in whatever state it is here.

load_catalogue_file(→ List[ModelEntry])

Read a JSON catalogue of remote models.

open_uri(→ Tuple[Iterable[bytes], int])

Open a model URI for streaming. (chunks, total_bytes).

package_model_root(→ pathlib.Path)

<spacr>/resources/models — where the bundled pack lives.

publish_model() → Dict[str, Any])

Upload a model to Hugging Face and return its catalogue row.

rank(→ List[BenchmarkResult])

Order benchmarks best-first — within one field set only.

rank_groups(→ Dict[str, List[BenchmarkResult]])

Rank inside each field set, keeping the sets apart. The safe alternative.

resolve(→ ModelEntry)

Turn a key, a name or a path into a ModelEntry.

scorecard_html(→ str)

The model's scorecard as an HTML table, for a tooltip.

sha256_file(→ str)

Hex SHA-256 of a file, read in chunks so a 2 GB checkpoint is not RAM.

shared_catalogue(→ Tuple[ModelEntry, ...])

The community catalogue, fetched from REMOTE_CATALOGUE_URI.

shared_catalogue_is_stale(→ bool)

Whether shared_catalogue() would go to the network to answer.

source_of(→ str)

Which of ZOO_SOURCES this row belongs under.

stock_cellpose_entries(→ List[ModelEntry])

Every model the installed Cellpose can fetch for itself.

verify(→ bool)

Hash the file this entry points at and compare it to a known digest.

versioned_path(→ pathlib.Path)

The first free destination for filename in dest.

Module Contents

exception spacr.model_zoo.ChecksumMismatch[source]

Bases: ModelZooError

What arrived is not what the catalogue published. Nothing was installed.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.model_zoo.DownloadCancelled[source]

Bases: ModelZooError

The caller cancelled a fetch. Nothing was left at the destination.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.model_zoo.IncomparableBenchmarks[source]

Bases: ModelZooError

Benchmarks from different field sets cannot be ranked against each other.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.model_zoo.ModelUnreadable[source]

Bases: ModelZooError

A model file is missing, empty, or not a checkpoint. Names the file.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.model_zoo.ModelZooError[source]

Bases: Exception

Base class for every refusal in this module.

Initialize self. See help(type(self)) for accurate signature.

class spacr.model_zoo.BenchmarkResult[source]

One model over one field set. Only comparable to results on the same set.

Parameters:
  • entry – the model that ran.

  • fieldset – fieldset_id() of the images — a hash of the pixels, so two runs over the same three fields share it and two runs over different fields never do, whatever the folders were called.

  • fieldset_label – the same thing for a human.

  • rows – one FieldBenchmark per field, in field order.

  • seconds – wall-clock seconds the model spent segmenting.

  • honoured – the parameters that reached the model.

  • ignored – what was set and Cellpose 4 dropped (diam_mean and friends) — carried so a benchmark cannot silently be a benchmark of settings nothing read.

  • notes – warnings the reader must see before the numbers.

  • masks – the label images, when kept, so a GUI can draw them.

  • images – the source fields, likewise.

  • object_type – what was segmented.

property fields: List[str][source]

Every field this model was benchmarked on.

Returns:

the field names.

property mean_objects: float[source]

Objects per field.

NaN rather than zero when nothing was benchmarked: no fields is a different statement from a model that found nothing.

Returns:

the mean, or NaN.

property n_failed: int[source]

Fields spacr.seg_qc scored 'fail'.

property n_fields: int[source]

How many fields were benchmarked.

Returns:

the field count.

property n_ok: int[source]

How many fields came back without a quality complaint.

Returns:

the count of fields at severity ok.

property qc_score: float[source]

Fraction of fields seg_qc scored 'ok'; nan without QC.

A quality-control verdict on this model’s own masks — it says the masks are not obviously broken, not that they are right. There is no ground truth in a benchmark (see the module docstring), so this is as close to a score as the zoo will produce.

property summary: str[source]

The model, what it found, over how many fields, and its QC score.

Returns:

a one-line summary.

property total_objects: int[source]

Every object the model found, across all fields.

Returns:

the object count.

class spacr.model_zoo.FieldBenchmark[source]

One field’s result for one model.

Parameters:
  • field – the field name.

  • n_objects – labels in the mask this model produced.

  • severity – spacr.seg_qc’s verdict — 'ok', 'warn' or 'fail', or '-' when QC was off.

  • flags – the named defects seg_qc raised.

  • note – seg_qc’s verdict in prose, with its numbers in it.

class spacr.model_zoo.ModelEntry[source]

One model the zoo knows about, wherever it lives.

Frozen because an entry is a record of a file at a moment — the hash, the size and the provenance describe those bytes. Changing one in place would silently invalidate the other two; dataclasses.replace() makes the new record explicit.

Parameters:
  • key – stable id, unique within a listing. For a local file this is derived from the filename; for a catalogue entry it is whatever the catalogue declared.

  • name – the filename (or the published name) — what a human reads.

  • kind – 'cellpose' or 'classifier'; see KINDS.

  • source – 'bundled' (ships with spaCR), 'local' (found on this machine) or 'remote' (declared in a catalogue, not yet fetched).

  • path – absolute path on this machine, or '' for a remote entry.

  • uri – where a remote entry is fetched from, or ''.

  • version – the zoo’s own version number for a filename. '1' for a plain name, '2' for foo_v2.CP_model (see versioned_path()), or whatever a catalogue declared.

  • sha256 – hex digest. For a downloaded model this is the digest of the bytes that were actually written; for a catalogue entry it is the published digest to check against; '' means “no checksum known”, which fetch() treats as a refusal rather than a pass.

  • size_bytes – file size, 0 when unknown.

  • trained_on – what data produced this model, in prose, or UNKNOWN. Never ''.

  • trained_by – who produced it, or UNKNOWN. Never ''.

  • metrics – whatever numbers came with it — for a classifier, the best/last epoch metrics spacr.train_compare.load_run() recovered. Excluded from equality: two records of the same bytes are the same model whether or not somebody attached numbers to one of them.

  • notes – everything the reader needs to know that is not a field: missing provenance, an unverified download, a file that does not look like a checkpoint.

  • verified – True only when sha256 was checked against a published digest. A downloaded file whose hash was merely recorded is not verified, and says so.

  • settings_path – where the provenance came from, for the reader who wants to go and look at it.

  • licence – the licence the model or package is published under, as its publisher states it (an SPDX identifier where there is one), or '' when none is recorded.

__post_init__()[source]

Fill in the provenance fields and validate the kind.

A blank trained_on or trained_by reads as “no constraints”, so it is replaced with an explicit unknown – the field has to say so out loud rather than by omission.

Raises:

ValueError – if kind is not one of the known model kinds.

describe() → str[source]

The multi-line provenance card shown next to a selected model.

scorecard_lines() → List[str][source]

The scorecard as display lines, or the sentence saying there is none.

ONE SOURCE, FOUR RENDERINGS. The tooltip, the API page, the Zoo screen and the Hugging Face table all render THIS, so they cannot disagree – 366 found six README tiles pointing at three different API pages, and that is what happens when a number is written down in more than one place.

summary_line() → str[source]

One line for a list widget.

property checksum_state: str[source]

What the checksum column says, in one word.

'none'

no hash at all — nothing can be checked, and fetch() refuses such an entry unless the caller overrides it.

'published'

a hash came with the entry but the bytes are not here yet, so it is a promise about what will arrive.

'recorded'

the honest middle: the hash of the file on disk is known, but nobody published one to compare it against. It proves the file has not changed since we looked, and nothing more.

'verified'

the bytes on disk were compared with a published digest and match.

property exists: bool[source]

True when path names a file that is here now.

property model_card_url: str[source]

The Hugging Face page for this model, derived from its download uri.

A checksum and a metrics table are not enough on their own: the reader wants the page that says what the model was trained on and shows its training curves. Derived rather than declared, so every Hugging Face entry has one without a per-entry field to forget.

property provenance_known: bool[source]

True when this model says what it was trained on.

property scorecard_holdout: str[source]

name @ version of the hold-out set, or "".

A SCORECARD WITHOUT ITS SET IS A NUMBER WITHOUT A UNIT. Two people quoting an F1 for the same model have said nothing to each other unless they scored the same masks, so the set travels with the numbers into every surface that shows them.

property scorecard_known: bool[source]

True when this model says how accurate it is.

The accuracy twin of provenance_known, which 370 asks for by name. A model with no numbers did not score zero, and a table of empty cells reads as the second – so the absence is a state to report rather than a gap to render.

spacr.model_zoo.benchmark(entry: ModelEntry, images: Sequence[Any] | None = None, source: Any = None, n_fields: int = DEFAULT_N_FIELDS, field_names: Sequence[str] | None = None, segment_fn: Callable | None = None, settings: Mapping[str, Any] | None = None, object_type: str = 'cell', qc: bool = True, keep_images: bool = True, channel: int | None = None, progress: Callable[[str, int, int], None] | None = None) → BenchmarkResult[source]

Run one model over N fields and report what came out. “Test on 3 fields”.

This is spacr.model_compare’s harness with one model instead of two: the same load_fields() reader, the same ModelConfig (so the same arguments are honoured and the same ones reported as ignored), the same segment_with_cellpose() backend, and the same spacr.seg_qc scorecards. To put two models side by side use compare_entries(), which calls compare_models() proper.

The checkpoint is checked before it is loaded, so a missing or corrupt file fails with its own name in the message rather than a torch KeyError.

Parameters:
  • entry – the model to run.

  • images – fields already in memory; None loads them from source.

  • source – a folder of fields (.tif / .png / .npy / .npz), read by spacr.model_compare.load_fields().

  • n_fields – how many fields to take from source.

  • field_names – names for the rows.

  • segment_fn – fn(images, config) -> masks; defaults to spacr.model_compare.segment_with_cellpose(). This is the seam the GUI and the tests use, and the reason no test here loads Cellpose.

  • settings – eval overrides (diameter, flow_threshold, …).

  • object_type – what is being segmented, for the seg_qc scorecards.

  • qc – score the masks with spacr.seg_qc.

  • keep_images – keep images and masks on the result for a GUI to draw.

  • channel – index into the last axis for multi-channel fields.

  • progress – fn(message, done, total).

Returns:

a BenchmarkResult.

Raises:
  • ModelUnreadable – when the checkpoint is missing or not a checkpoint.

  • ValueError – when there is no field, or the model returned the wrong number of masks.

spacr.model_zoo.bioimageio_entries(timeout: float = 5.0, url: str | None = None, allow_network: bool = False) → List[ModelEntry][source]

Every Cellpose model bioimage.io publishes, as zoo rows.

Read from bioimage.io’s own collection, so the category cannot go stale the way a typed list would. Three kinds of row:

  • a Cellpose-SAM or Cellpose-DINO model, which spaCR’s own Cellpose 4 loads: a cellpose row that downloads its weights;

  • a Cellpose 3-format checkpoint: a cellpose3 row that downloads its weights for the Cellpose 3 backend, which runs it as it runs cyto3;

  • one of either that spaCR cannot run, which says why in its first note (_bioimageio_cannot_run()) and offers nothing to download.

Each weights file is checked against the SHA-256 its manifest publishes, and carries the licence the uploader chose and what it was trained on.

Best effort and never raises: no network, a slow mirror or a changed schema all mean fewer rows, never a zoo that fails to open. With allow_network the collection is refreshed once it is a day old, and the weights files are measured, so a stand-in file is refused before anyone downloads it.

spacr.model_zoo.catalogue(include_bundled: bool = True, remote: bool = True, catalogue_path: Any = None, include_plugins: bool = True, block: bool | None = None) → List[ModelEntry][source]

Everything the zoo knows about without scanning the user’s disks.

That is: the models bundled with the installed package (whatever spacr.utils.download_models() has put in resources/models), plus the declared remote entries — BUNDLED_REMOTE_MODELS and, if one is configured, the JSON catalogue named by catalogue_path or the CATALOGUE_ENV_VAR environment variable.

Local apart from one thing, and the exception used to be undocumented: with remote on, this also asks shared_catalogue() for the community rows, which is a network call. It works offline either way – that fetch never raises – and it never blocks Qt’s GUI thread, which shared_catalogue() enforces for itself.

Parameters:
  • include_bundled – list the models in the package resources folder.

  • remote – list declared remote entries.

  • catalogue_path – a JSON catalogue to add; defaults to $SPACR_MODEL_CATALOGUE when that names a file.

  • include_plugins – include entries returned by installed spaCR model providers. Provider failures are recorded in plugin diagnostics and do not hide built-in entries.

  • block – passed to shared_catalogue(). False takes the community rows from its cache rather than waiting for the network; None lets that function decide from the thread it is on.

Returns:

bundled entries first, then remote ones already present locally are dropped (a downloaded model is listed once, as the local file).

spacr.model_zoo.classify_kind(path: Any) → str | None[source]

Say whether a file is a Cellpose model, a classifier, or not a model.

The rules, in order:

  1. *.CP_model is a Cellpose checkpoint — that is what spacr.submodules.train_cellpose() names its output.

  2. *.pth / *.pt is a classifier checkpoint (spacr.io._save_model() writes <model_type>_epoch_<n>_channels_<ch>.pth) unless it sits in a Cellpose folder or has cellpose/cp_model in its name.

  3. An extensionless file inside a Cellpose folder is a Cellpose checkpoint only if its first bytes are a torch save. cellpose.train writes <save_path>/models/<name> with no suffix, and that folder also holds READMEs and logs — the magic-byte check is what keeps a README out of the zoo.

  4. Anything else is not a model. CSVs, PNGs, .npy masks and settings snapshots all land here and are ignored.

Parameters:

path – a file path.

Returns:

'cellpose', 'classifier' or None.

spacr.model_zoo.community_entries(allow_network: bool = False, repo: str = COMMUNITY_REPO) → List[ModelEntry][source]

Unvetted models uploaded by spaCR users, or an empty list.

These are shown only when the user asks for them, because an unreviewed checkpoint sitting beside a measured one invites the reader to treat them alike. Every row carries COMMUNITY_WARNING, and the checksum comes from the uploader’s own submission record – it proves the file has not changed since it was uploaded, NOT that it is any good.

Cache-first for the same reason as the bioimage.io listing: catalogue() must not reach the network.

spacr.model_zoo.compare_entries(entry_a: ModelEntry, entry_b: ModelEntry, images: Sequence[Any] | None = None, source: Any = None, n_fields: int = DEFAULT_N_FIELDS, field_names: Sequence[str] | None = None, settings_a: Mapping[str, Any] | None = None, settings_b: Mapping[str, Any] | None = None, **kwargs: Any)[source]

Put two zoo entries head to head on the same fields.

Straight delegation to spacr.model_compare.compare_models() — the metrics, the split/merge attribution and the “neither model is ground truth” wording all come from there, unchanged. This function’s only job is turning two ModelEntry objects into two ModelConfig objects.

Parameters:
  • entry_a – the A side.

  • entry_b – the B side.

  • images – fields already in memory; None loads them from source.

  • source – a folder of fields.

  • n_fields – how many fields to take from source.

  • field_names – names for the rows.

  • settings_a – eval overrides for A.

  • settings_b – eval overrides for B.

  • kwargs – forwarded to spacr.model_compare.compare_models().

Returns:

a spacr.model_compare.ComparisonReport.

spacr.model_zoo.config_for(entry: ModelEntry, overrides: Mapping[str, Any] | None = None)[source]

The spacr.model_compare.ModelConfig that runs this entry.

A local checkpoint is passed by path, which spacr.utils._choose_model() and spacr.model_compare.segment_with_cellpose() both load as pretrained_model; anything else goes through by name and is subject to Cellpose 4’s legacy-name remapping, which the config reports.

Parameters:
  • entry – the model.

  • overrides – eval settings (diameter, flow_threshold, …).

Returns:

the config.

spacr.model_zoo.default_local_roots() → List[pathlib.Path][source]

Folders worth scanning when the caller has not named one.

The bundled pack, the Cellpose user folder, and spaCR’s own model cache. Only the ones that exist come back.

spacr.model_zoo.discover_local(roots: Any = None, max_depth: int = DEFAULT_SCAN_DEPTH, compute_hashes: bool = False, limit: int = DEFAULT_SCAN_LIMIT) → List[ModelEntry][source]

Find the model checkpoints already on this machine.

Cellpose models and classifier checkpoints are told apart by classify_kind(); everything else in the folders — settings CSVs, mask .npy files, montage PNGs, logs — is ignored.

Nothing is downloaded and nothing is hashed unless compute_hashes is set: this is the function behind a list widget, and it has to be fast enough to run on a folder the user just typed.

Parameters:
  • roots – a folder, a file, or an iterable of them; None uses default_local_roots().

  • max_depth – how deep below each root to look.

  • compute_hashes – hash every file found (minutes on a big folder).

  • limit – stop after examining this many files per root.

Returns:

entries, Cellpose first, then by name.

spacr.model_zoo.download_bundled_models(**kwargs: Any) → str[source]

Pull the bundled Hugging Face model pack via the existing downloader.

Thin, deliberate wrapper over spacr.utils.download_models() — the downloader spaCR already ships and the one spacr.submodules.analyze_plaques() depends on. It is not reimplemented here, so there is one code path that fills resources/models and one place to fix when the repo moves.

It is also the unverified path: that function has no checksum, writes straight to the destination filename, and skips the whole pull when the folder is non-empty. Prefer a catalogue entry with a hash and install(); this exists so the zoo can offer the legacy pack rather than pretend it does not exist.

spacr.utils imports torch, so it is imported here and not at module level. Nothing else in this module reaches for it.

Parameters:

kwargs – forwarded to spacr.utils.download_models().

Returns:

the local directory the pack landed in.

spacr.model_zoo.entries_from_sources(entries: Iterable[Any], sources: Iterable[str]) → List[Any][source]

The rows belonging to the headings that are on, in the given order.

Parameters:
  • entries – the rows to filter.

  • sources – the headings currently on.

spacr.model_zoo.entry_from_file(path: Any, kind: str | None = None, source: str = 'local', key: str | None = None, runs: Mapping[str, Any] | None = None, compute_hash: bool = False, sha256: str = '', verified: bool = False, extra_notes: Sequence[str] = ()) → ModelEntry[source]

Build a ModelEntry for a checkpoint on this machine.

Provenance is recovered from whatever spaCR already wrote beside the model: a *_settings.csv or <src>/settings/<name>.csv for a Cellpose model, and — for a classifier — the training run the checkpoint sits in, loaded through spacr.train_compare.load_run() so the zoo and the training-run comparison agree about where settings live and what they say.

Parameters:
  • path – the checkpoint.

  • kind – override classify_kind().

  • source – 'local' or 'bundled'.

  • key – override the generated key.

  • runs – {folder: TrainingRun} from _runs_under(), so a scan of 40 checkpoints in one run folder reads that folder once.

  • compute_hash – hash the file now. Off by default: hashing every checkpoint on a machine to populate a list widget is minutes.

  • sha256 – a digest already known for these bytes.

  • verified – whether sha256 was checked against a published digest.

  • extra_notes – notes to carry onto the entry.

Returns:

the entry.

Raises:

ModelUnreadable – when the path is not a file.

spacr.model_zoo.fetch(entry: ModelEntry, dest: Any, expected_sha256: str | None = None, require_checksum: bool = True, opener: Callable[[str], Any] | None = None, progress: Callable[[int, int], None] | None = None, cancel: Callable[[], bool] | None = None, chunk_size: int = DEFAULT_CHUNK, timeout: int = DEFAULT_TIMEOUT) → pathlib.Path[source]

Download a model, verify it, and only then put it where it belongs.

The order is the whole point:

  1. bytes stream into a temporary file inside dest, so the rename in step 4 is a same-filesystem os.replace and therefore atomic;

  2. the checksum of what actually arrived is computed;

  3. if it does not match the published digest the temporary file is deleted and ChecksumMismatch is raised — nothing is installed, and the destination still holds whatever it held before;

  4. only now is the temporary file renamed, to a versioned_path() that does not exist yet.

Every failure — a dead server, a cancel, a bad hash, a full disk — leaves the destination directory exactly as it was. There is no window in which a half-written file sits at a name that looks like a model.

Parameters:
  • entry – what to fetch. ModelEntry.uri is the source.

  • dest – destination directory; created if missing.

  • expected_sha256 – digest to require, overriding ModelEntry.sha256.

  • require_checksum – refuse to install when no digest is known. True by default — a download nobody can check is exactly the thing this module exists to stop being routine. Pass False to accept one knowingly; the entry install() returns then reports verified=False.

  • opener – fn(uri) -> chunks or fn(uri) -> (chunks, total); defaults to open_uri().

  • progress – fn(done_bytes, total_bytes); total is 0 when the server did not say.

  • cancel – fn() -> bool, polled between chunks. Returning True deletes the partial file and raises DownloadCancelled.

  • chunk_size – bytes per read.

  • timeout – seconds, HTTP only.

Returns:

the path the model was written to.

Raises:
spacr.model_zoo.fieldset_id(names: Sequence[str], images: Sequence[Any]) → str[source]

A stable id for a set of fields, taken from the pixels.

Folder names are not identity: plate1/1 on two machines is two different sets of images, and the same three images copied to a new folder are the same benchmark input. So the id hashes each array’s bytes, shape and dtype together with its name.

This is what makes rank() able to refuse. Without it, two benchmarks run on different data are two numbers, and two numbers always sort.

Parameters:
  • names – field names, in order.

  • images – the arrays, in the same order.

Returns:

a 16-character hex id.

spacr.model_zoo.format_benchmarks(results: Sequence[BenchmarkResult], key: str = DEFAULT_RANK_KEY) → str[source]

Render benchmarks grouped by field set, ranked only within a group.

Two models benchmarked on different fields appear under two headers with a line saying the two blocks cannot be compared. That is the alternative to rank()’s refusal, and it is the only way this module will ever put incomparable numbers on the same page.

Parameters:
  • results – benchmarks, from any number of field sets.

  • key – one of RANK_KEYS.

Returns:

a multi-line string.

spacr.model_zoo.format_zoo(entries: Sequence[ModelEntry]) → str[source]

Render a listing a human reads before choosing a model.

Provenance is a column, not a footnote: “trained on” is the field that decides whether a model is applicable to your images at all, and it is printed for every row — reading unknown where it is unknown, because a blank there would read as “no constraints”.

Parameters:

entries – what to list.

Returns:

a multi-line string.

spacr.model_zoo.group_by_fieldset(results: Sequence[BenchmarkResult]) → Dict[str, List[BenchmarkResult]][source]

Bucket benchmarks by the field set they ran on, first-seen order.

Parameters:

results – benchmarks.

Returns:

{fieldset_id: [results]}.

spacr.model_zoo.group_by_source(entries: Iterable[Any]) → Dict[str, List[Any]][source]

Split a listing into ZOO_SOURCES, keeping each source’s order.

Every heading is present even when it has no rows, so a caller drawing the strip does not have to know which of the five happened to be empty this time.

Parameters:

entries – the rows to split.

Returns:

heading -> rows, in ZOO_SOURCES order.

spacr.model_zoo.hf_uri(repo_id: str, filename: str, repo_type: str = 'dataset') → str[source]

The download URL for a file in a Hugging Face repo.

Parameters:
  • repo_id – Hugging Face repository identifier.

  • filename – repository-relative name of the file to download.

  • repo_type – "dataset" (the default, and what spaCR shipped first) or "model".

Exactly the URL spacr.utils.download_models() and spacr.qt.hf_download._download_one() build, kept in one place so the zoo cannot drift away from the downloader spaCR already ships.

THE TWO REPO KINDS HAVE DIFFERENT URLS, which is not cosmetic: a dataset file lives under /datasets/<repo>/resolve/... and a model file under /<repo>/resolve/.... Asking for one at the other’s URL returns a 404 page, and a downloader that does not check the content type writes that HTML into the destination and leaves a “checkpoint” that fails to load with a torch error naming neither the URL nor the repo.

dataset remains the default because HF_MODELS_REPO is a DATASET repo – einarolafsson/models – and every entry written before this parameter existed assumes it. New model repos pass "model".

spacr.model_zoo.inspect_checkpoint(path: Any, loader: Callable[[str], Any] | None = None, deep: bool = False) → Dict[str, Any][source]

Check a file is a loadable checkpoint, failing with the filename in it.

The default failure for a wrong or corrupt checkpoint is a KeyError on a state-dict key raised somewhere inside torch, which names nothing the user chose and reads like a spaCR bug. This turns all of it — missing, empty, truncated, a PNG somebody renamed, a Cellpose model handed to the classifier path — into one ModelUnreadable naming the file.

The shallow check needs no torch at all: it is a stat and four bytes.

Parameters:
  • path – the checkpoint.

  • loader – fn(path) -> object used for the deep check; defaults to torch.load(..., map_location='cpu'), imported only if used.

  • deep – actually load the file. Off by default because loading a 2 GB checkpoint to populate a list widget is not acceptable.

Returns:

{'path', 'size_bytes', 'format', 'loaded'}.

Raises:

ModelUnreadable – naming the file, always.

spacr.model_zoo.install(entry: ModelEntry, dest: Any, **kwargs: Any) → ModelEntry[source]

fetch() the model and return the registered local entry.

The returned entry carries the digest of the bytes that were actually written — not the one the catalogue claimed — and ModelEntry.verified is True only when the two were compared and matched. Provenance from the catalogue entry is carried over, because that is the whole reason for having had a catalogue.

Parameters:
  • entry – the remote entry.

  • dest – destination directory.

  • kwargs – passed to fetch().

Returns:

a source='local' entry pointing at the new file.

spacr.model_zoo.installable_backend_entries() → List[ModelEntry][source]

Every optional segmentation backend, in whatever state it is here.

A backend absent from the zoo teaches nobody that it exists, so each one is listed, and its source says where it stands – installed, installable, installing or not installable here – with the reason as its first note and its licence on the row. Installing one builds it an environment of its own under ~/.spacr/backends and leaves spaCR’s own environment alone.

spacr.model_zoo.load_catalogue_file(path: Any) → List[ModelEntry][source]

Read a JSON catalogue of remote models.

Format — a list, or an object with a models list:

{"models": [
  {"key": "hela_60x",
   "name": "hela_60x_confluent.CP_model",
   "kind": "cellpose",
   "uri": "https://…/hela_60x_confluent.CP_model",
   "sha256": "9f86d0…",
   "size_bytes": 26566572,
   "trained_on": "HeLa, 60x, confluent monolayer, 512px crops",
   "trained_by": "A. Researcher, 2026-02",
   "metrics": {"note": "benchmarked on plate3 fields 1-3"}}
]}

sha256 is the field that decides whether the entry is usable without an explicit override, so a catalogue is worth exactly as much as its hashes.

Parameters:

path – the JSON file.

Returns:

the entries.

Raises:

ModelZooError – when the file cannot be read or is not a catalogue, naming the file.

spacr.model_zoo.open_uri(uri: str, timeout: int = DEFAULT_TIMEOUT, chunk_size: int = DEFAULT_CHUNK) → Tuple[Iterable[bytes], int][source]

Open a model URI for streaming. (chunks, total_bytes).

http:// and https:// stream over requests — the same call spacr.utils.download_models() and spacr.qt.hf_download._download_one() make, imported here so this module has no hard dependency on it. file:// and a plain existing path are read from disk, which is what a lab mirror on a NAS looks like and what the tests use, so the whole fetch path is exercised without a network.

Parameters:
  • uri – where the model lives.

  • timeout – seconds, HTTP only.

  • chunk_size – bytes per chunk.

Returns:

(iterable of byte chunks, total size or 0 when unknown).

Raises:

ModelZooError – for a scheme this does not speak.

spacr.model_zoo.package_model_root() → pathlib.Path[source]

<spacr>/resources/models — where the bundled pack lives.

The same folder spacr.utils.download_models() fills and spacr.submodules.analyze_plaques() reads from.

spacr.model_zoo.publish_model(local_path: Any, repo_id: str, *, key: str, kind: str = 'cellpose', trained_on: str = UNKNOWN, trained_by: str = UNKNOWN, private: bool = False, notes: Sequence[str] = ()) → Dict[str, Any][source]

Upload a model to Hugging Face and return its catalogue row.

Parameters:
  • local_path – the checkpoint to upload.

  • repo_id – <user>/<repo> – YOUR OWN account.

  • key – the short name spaCR will offer the model under.

  • kind – one of KINDS.

  • trained_on – what the model was trained on. Say it properly: this is the only thing another lab has to decide whether it applies to them.

  • trained_by – who trained it, and roughly when.

  • private – keep the repo private. A private model cannot be fetched by other spaCR users, so it is off by default.

  • notes – caveats worth carrying next to the model.

Returns:

the catalogue row, with the sha256 filled in.

Raises:

ImportError – when huggingface_hub is not installed.

THE CHECKSUM IS COMPUTED HERE, from the file that was actually uploaded, which is the whole reason this exists as a function rather than as instructions in a README. fetch() refuses an entry it cannot verify, so a row written by hand without a hash produces a model nobody can install without disabling the check – which is what the one pre-existing bundled entry does, and it is a hole rather than a precedent.

Publishing does NOT distribute the model on its own: add the returned row to the shared catalogue (REMOTE_CATALOGUE_URI) and every spaCR user sees it within CATALOGUE_CACHE_SECONDS.

spacr.model_zoo.rank(results: Sequence[BenchmarkResult], key: str = DEFAULT_RANK_KEY) → List[BenchmarkResult][source]

Order benchmarks best-first — within one field set only.

A model’s numbers on your three fields say nothing about its numbers on somebody else’s: different cell density, different exposure, different magnification. Sorting results from two field sets into one list produces a ranking that looks exactly like a real one and means nothing, which is the failure this function exists to prevent. So it refuses.

Parameters:
  • results – benchmarks, all from the same field set.

  • key – one of RANK_KEYS.

Returns:

the results, best first.

Raises:
spacr.model_zoo.rank_groups(results: Sequence[BenchmarkResult], key: str = DEFAULT_RANK_KEY) → Dict[str, List[BenchmarkResult]][source]

Rank inside each field set, keeping the sets apart. The safe alternative.

Parameters:
  • results – benchmarks from any number of field sets.

  • key – one of RANK_KEYS.

Returns:

{fieldset_id: [results, best first]}.

spacr.model_zoo.resolve(key_or_path: Any, entries: Sequence[ModelEntry] | None = None) → ModelEntry[source]

Turn a key, a name or a path into a ModelEntry.

A path that exists wins over a key: pointing the zoo at a file you just trained has to work without registering it anywhere first.

Parameters:
  • key_or_path – an entry key, a model filename, or a path to a file.

  • entries – the listing to search; defaults to catalogue().

Returns:

the entry.

Raises:
spacr.model_zoo.scorecard_html(entry) → str[source]

The model’s scorecard as an HTML table, for a tooltip.

A paragraph of prose is what a tooltip used to show, and a reader comparing two models had to parse two paragraphs to find two numbers. The same table the model card prints answers that at a glance. Falls back to the prose when an entry publishes no metrics, because an empty table is worse than a sentence.

Metrics that hold none of the scorecard’s keys – a free-form note, a training loss under a name of its own – are not a scorecard either, and return nothing too: a table of eleven “not recorded” rows would replace a two-line note that said something.

Parameters:

entry – a catalogue entry, or anything with metrics and a name.

Returns:

the HTML table, or "" when there is no scorecard to show.

spacr.model_zoo.sha256_file(path: Any, chunk_size: int = 1 << 20) → str[source]

Hex SHA-256 of a file, read in chunks so a 2 GB checkpoint is not RAM.

Parameters:
  • path – the file.

  • chunk_size – bytes per read.

Returns:

the lowercase hex digest.

Raises:

ModelUnreadable – when the file is missing or cannot be read, with the path in the message.

spacr.model_zoo.shared_catalogue(uri: str | None = None, *, timeout: float = DEFAULT_TIMEOUT, force: bool = False, block: bool | None = None) → Tuple[ModelEntry, ...][source]

The community catalogue, fetched from REMOTE_CATALOGUE_URI.

Parameters:
  • uri – override the catalogue location.

  • timeout – seconds to wait for the request.

  • force – ignore the cache and re-fetch.

  • block – whether to wait for the network. None – the default, and what an unthinking caller gets – waits everywhere EXCEPT Qt’s GUI thread, where it answers from the cache and refreshes on a daemon thread. True waits wherever it is called, which only a caller that knows it is on a worker or in a CLI may ask for. False never waits.

Returns:

the entries, or () when the catalogue cannot be read.

NEVER FETCHES ON THE GUI THREAD, and that is not an optimisation. This is reached from spacr.settings.downloaded_zoo_models while a settings panel is being built, so it ran inside MainWindow._on_nav_selected with nothing able to paint or answer the compositor. Measured with the catalogue host non-routable (10.255.255.1, the shape of a down VPN or a captive portal – the connect neither completes nor is refused): opening the Mask module took 32.2 s, all of it a GUI thread stuck in urlopen. GNOME asks a window whether it is alive after five, so what the user sees is spaCR’s “force quit” dialog, which is how this was reported. With the fetch moved off the thread the same open is 2.4 s.

The cost of not waiting is a first module open whose Cellpose dropdown lists the bundled and local models but not the community ones; the background refresh means the second one has them.

NEVER RAISES, and that is deliberate. This runs when a user opens a module that offers a model list, and the list is useful without it: the bundled entries and any local models are still there. A laptop on a train, a lab behind a proxy and a Hugging Face outage all produce the same thing – a shorter list and a log line – rather than a module that will not open.

The failure that WOULD be silent and harmful is a corrupt or hostile catalogue, so entries that do not parse are dropped individually, and an entry without a sha256 still cannot be installed by fetch() without an explicit override. A catalogue row is a claim about where a file lives; the checksum is what makes it a claim about which file.

spacr.model_zoo.shared_catalogue_is_stale() → bool[source]

Whether shared_catalogue() would go to the network to answer.

For a caller that wants to do the waiting somewhere it is allowed to – a worker thread – rather than get the cached answer and not know it was one.

spacr.model_zoo.source_of(entry: Any) → str[source]

Which of ZOO_SOURCES this row belongs under.

Decided FROM THE ROW, never from a list of names kept somewhere else: a hand-written list goes stale the first time a model is added, and the failure it produces is a model that is in the catalogue and under no heading, which is a model nobody can see.

The order of the tests is the rule, and it matters in one place: a Cellpose 3 checkpoint published on bioimage.io is a bioimage.io row, not a Cellpose 3 one. cellpose3 means the backend’s OWN models – cyto, cyto2, cyto3, nuclei – and the backend package that runs them.

A row that matches nothing is filed under spaCR AND SAID OUT LOUD. It is the fallback rather than a sixth heading because a model under the wrong heading is a nuisance and a model under no heading is a bug the user experiences as a missing model.

Parameters:

entry – any zoo row – a ModelEntry, or anything with kind, source and uri.

Returns:

one of ZOO_SOURCES.

spacr.model_zoo.stock_cellpose_entries() → List[ModelEntry][source]

Every model the installed Cellpose can fetch for itself.

These are not spaCR’s files and carry no checksum of ours: Cellpose downloads and verifies them, and the name IS the path – passing “cpsam” to Cellpose resolves it. They are listed so that the zoo answers “what can I segment with” rather than “what has Einar trained”, which is the question a new user actually has.

spacr.model_zoo.verify(entry: ModelEntry, expected: str | None = None) → bool[source]

Hash the file this entry points at and compare it to a known digest.

Parameters:
  • entry – the entry to check.

  • expected – the digest to compare against; defaults to ModelEntry.sha256.

Returns:

True when the file’s digest matches.

Raises:
  • ModelUnreadable – when the entry has no local file, naming it.

  • ModelZooError – when there is no digest to compare against — that is a caller error, and returning False for it would read as “the file is wrong” when what happened is “nobody said what right looks like”.

spacr.model_zoo.versioned_path(dest: Any, filename: str) → pathlib.Path[source]

The first free destination for filename in dest.

foo.CP_model -> foo.CP_model, then foo_v2.CP_model, foo_v3.CP_model… An existing checkpoint is never overwritten: two models with the same filename are a normal thing to have (the same author retrained, or two people picked the same name), and the failure mode of overwriting — a run that used the old weights becoming unreproducible with no trace — is silent.

An input that already carries _vN counts from there rather than becoming foo_v2_v2.

Parameters:
  • dest – destination directory.

  • filename – the name to place there.

Returns:

a path that does not exist yet.

Nested helpers

_default_segmenter._segment(images, config)

Segment images with the backend the config’s model belongs to.

spacr/model_zoo.py:3923

_refresh_shared_catalogue_in_background.run() → None

Fetch the catalogue off the GUI thread, and always release the flag.

The finally is the whole point: _SHARED_CATALOGUE_FETCHING is what stops a second refresh being started while this one is in flight, so a fetch that raises must still clear it or no later refresh can ever begin.

spacr/model_zoo.py:1969

benchmark._tick(message: str, done: int) → None

Report one benchmark milestone through the captured callback.

Parameters:
  • message – stage description for the progress display.

  • done – completed-step index from zero through two.

Returns:

None. When a callback was supplied it receives the message, completed index, and captured total of two; otherwise this is a no-op.

spacr/model_zoo.py:3834

scorecard_html.cell(value)

A scorecard value as shown, or “not recorded” when blank.

spacr/model_zoo.py:4127