spacr.model_zoo¶
Workflow inputs and outputs¶
Model Zoo¶
Inspect model provenance and download/install compatible checkpoints or backends. A listed backend is not itself a checkpoint file.
Open: Make Masks → Model Zoo.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Segmentation checkpoint — Saved Cellpose-compatible checkpoint or a compatible installed backend selected with its own configuration.
Outputs
Segmentation checkpoint — Saved Cellpose-compatible checkpoint or a compatible installed backend selected with its own configuration.
Browse, verify, fetch and benchmark spaCR segmentation and classification models.
Why this exists¶
spaCR can run a stock Cellpose model, a Cellpose model somebody on the team
fine-tuned last year, a classifier checkpoint from a run three folders up, or a
checkpoint downloaded from Hugging Face. Today the only way to find out which
of those exist on a machine is find / -name '*.pth', the only way to know
what one of them was trained on is to remember, and the only way to know
whether the copy on disk is the file the author published is to hope.
This module answers those three questions and nothing else:
what is here —
discover_local()walks folders and returns the checkpoints, classified as Cellpose or classifier, with whatever provenance is recoverable from the settings snapshots spaCR already writes;is it the right bytes —
sha256_file()/verify()/fetch(), which downloads atomically, checksums what arrived, and refuses to install a mismatch;what does it do on my data —
benchmark(), which isspacr.model_compare’s “three fields” harness pointed at one model instead of two.
Nothing here imports torch or cellpose at module import time. Browsing the zoo,
reading provenance, checksumming a file and rendering the table all work on a
machine with neither installed; only benchmark() (through
spacr.model_compare.segment_with_cellpose()) needs them, and only when it
is called. tests/test_model_zoo.py asserts that.
The four things that make this trustworthy rather than merely convenient¶
- Every download is checksummed, and verification happens before use.
A checkpoint truncated by a dropped connection, or swapped for a different one at the same URL, still loads. It does not raise; it produces silently different masks, and the run that used it looks exactly like the run that did not. So
fetch()hashes what arrived and compares it to the hash the catalogue published: a mismatch deletes the file and raisesChecksumMismatch— the entry is never registered. The hash that was actually computed is stored on the returned entry, so “verified” means “these bytes”, not “this filename”.A catalogue entry with no published hash cannot be verified at all, and that is refused by default (
require_checksum=True) rather than quietly treated as fine. Callers that knowingly accept an unverifiable source passrequire_checksum=False; the resulting entry carriesverified=Falseand says so informat_zoo().- Every download is atomic, and never overwrites.
Bytes stream into a temporary file in the destination directory (so the final
os.replaceis a same-filesystem rename, which is atomic) and the rename happens only after the checksum passes. An interrupted or cancelled download therefore leaves nothing behind that looks like a model — the failure mode where half a checkpoint sits at the real filename, loads, and segments badly, cannot happen.The destination is versioned rather than overwritten (
versioned_path():foo.CP_model,foo_v2.CP_model, …). Two models with the same filename are a normal thing to have; losing the first one to the second is not.- Provenance is recorded, and “unknown” is written out in full.
A Cellpose model fine-tuned on 60x confluent HeLa is not interchangeable with one trained on 20x sparse fibroblasts, and no amount of benchmark score makes it so. What a model was trained on is the single most useful thing the zoo can show, so
ModelEntrycarriestrained_onandtrained_by, both recovered from the settings snapshots spaCR already writes beside its models (<file>_settings.csvfor a Cellpose model,<dst>/settings.csvfor a classifier — read throughspacr.train_compare.load_run(), which already knows where to look).Where it could not be recovered the field reads
'unknown', never''. A blank cell in a provenance table reads as “no constraints”; that is the opposite of what it means.- Benchmarks are only comparable inside one field set.
A model’s score on your three fields says nothing whatsoever about its score on somebody else’s, and a table that sorts the two together invents a ranking out of two unrelated numbers. So every
BenchmarkResultrecords afieldset_id()— a hash of the actual pixels, not the folder name — andrank()raisesIncomparableBenchmarkswhen handed results from more than one field set.rank_groups()andformat_benchmarks()are the supported alternative: they group by field set, rank within each group, and label the groups.
And a fifth, smaller one: a model file that is missing, empty, or not a torch
checkpoint at all fails in inspect_checkpoint() with a message naming the
file, before anything tries to load it. The default failure — a KeyError on
a state-dict key from deep inside torch — names nothing the user chose.
What the benchmark can and cannot say¶
There is no ground truth here. benchmark() runs one model over N fields
and reports what came out: object counts, timings and the
spacr.seg_qc verdict per field (fused? shattered? empty? all on the
border?). That is a quality-control score, not an accuracy — it can tell you
a model collapsed on your data, it cannot tell you which of two plausible
segmentations is right. RANK_KEYS therefore offers exactly two keys,
'qc' and 'seconds', and no key that would read as accuracy. To compare
two models against each other, use compare_entries(), which hands both to
spacr.model_compare.compare_models() — the A/B harness that is explicit
about neither side being the truth.
Cellpose 4 accepts and ignores model_type, diam_mean, nchan,
channels and rescale; only diameter at eval still changes the
masks. BenchmarkResult carries the resolved honoured and
ignored parameter dicts straight from
spacr.model_compare.ModelConfig so a benchmark cannot silently be a
benchmark of settings nothing read.
Example:
from spacr import model_zoo as zoo
entries = zoo.catalogue() + zoo.discover_local('/data/screen1')
print(zoo.format_zoo(entries))
entry = zoo.resolve('cpsam_plaque_r3', entries)
result = zoo.benchmark(entry, source='/data/screen1/plate1/1', n_fields=3)
print(zoo.format_benchmarks([result]))
See also
spacr.model_compare — the A/B harness this module reuses for
segmentation and for the two-model comparison.
spacr.train_compare — run discovery and settings recovery, reused
wholesale for the classifier half of the zoo.
spacr.utils.download_models() — the legacy bulk pull of the bundled
Hugging Face model pack, wrapped by download_bundled_models().
Exceptions¶
What arrived is not what the catalogue published. Nothing was installed. |
|
The caller cancelled a fetch. Nothing was left at the destination. |
|
Benchmarks from different field sets cannot be ranked against each other. |
|
A model file is missing, empty, or not a checkpoint. Names the file. |
|
Base class for every refusal in this module. |
Classes¶
One model over one field set. Only comparable to results on the same set. |
|
One field's result for one model. |
|
One model the zoo knows about, wherever it lives. |
Functions¶
|
Run one model over N fields and report what came out. "Test on 3 fields". |
|
Every Cellpose model bioimage.io publishes, as zoo rows. |
|
Everything the zoo knows about without scanning the user's disks. |
|
Say whether a file is a Cellpose model, a classifier, or not a model. |
|
Unvetted models uploaded by spaCR users, or an empty list. |
|
Put two zoo entries head to head on the same fields. |
|
The |
|
Folders worth scanning when the caller has not named one. |
|
Find the model checkpoints already on this machine. |
|
Pull the bundled Hugging Face model pack via the existing downloader. |
|
The rows belonging to the headings that are on, in the given order. |
|
Build a |
|
Download a model, verify it, and only then put it where it belongs. |
|
A stable id for a set of fields, taken from the pixels. |
|
Render benchmarks grouped by field set, ranked only within a group. |
|
Render a listing a human reads before choosing a model. |
|
Bucket benchmarks by the field set they ran on, first-seen order. |
|
Split a listing into |
|
The download URL for a file in a Hugging Face repo. |
|
Check a file is a loadable checkpoint, failing with the filename in it. |
|
|
|
Every optional segmentation backend, in whatever state it is here. |
|
Read a JSON catalogue of remote models. |
|
Open a model URI for streaming. |
|
|
|
Upload a model to Hugging Face and return its catalogue row. |
|
Order benchmarks best-first — within one field set only. |
|
Rank inside each field set, keeping the sets apart. The safe alternative. |
|
Turn a key, a name or a path into a |
|
The model's scorecard as an HTML table, for a tooltip. |
|
Hex SHA-256 of a file, read in chunks so a 2 GB checkpoint is not RAM. |
|
The community catalogue, fetched from |
|
Whether |
|
Which of |
|
Every model the installed Cellpose can fetch for itself. |
|
Hash the file this entry points at and compare it to a known digest. |
|
The first free destination for |
Module Contents¶
- exception spacr.model_zoo.ChecksumMismatch[source]¶
Bases:
ModelZooErrorWhat arrived is not what the catalogue published. Nothing was installed.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.model_zoo.DownloadCancelled[source]¶
Bases:
ModelZooErrorThe caller cancelled a fetch. Nothing was left at the destination.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.model_zoo.IncomparableBenchmarks[source]¶
Bases:
ModelZooErrorBenchmarks from different field sets cannot be ranked against each other.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.model_zoo.ModelUnreadable[source]¶
Bases:
ModelZooErrorA model file is missing, empty, or not a checkpoint. Names the file.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.model_zoo.ModelZooError[source]¶
Bases:
ExceptionBase class for every refusal in this module.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.model_zoo.BenchmarkResult[source]¶
One model over one field set. Only comparable to results on the same set.
- Parameters:
entry – the model that ran.
fieldset –
fieldset_id()of the images — a hash of the pixels, so two runs over the same three fields share it and two runs over different fields never do, whatever the folders were called.fieldset_label – the same thing for a human.
rows – one
FieldBenchmarkper field, in field order.seconds – wall-clock seconds the model spent segmenting.
honoured – the parameters that reached the model.
ignored – what was set and Cellpose 4 dropped (
diam_meanand friends) — carried so a benchmark cannot silently be a benchmark of settings nothing read.notes – warnings the reader must see before the numbers.
masks – the label images, when kept, so a GUI can draw them.
images – the source fields, likewise.
object_type – what was segmented.
- property fields: List[str][source]¶
Every field this model was benchmarked on.
- Returns:
the field names.
- property mean_objects: float[source]¶
Objects per field.
NaN rather than zero when nothing was benchmarked: no fields is a different statement from a model that found nothing.
- Returns:
the mean, or NaN.
- property n_failed: int[source]¶
Fields
spacr.seg_qcscored'fail'.
- property n_ok: int[source]¶
How many fields came back without a quality complaint.
- Returns:
the count of fields at severity
ok.
- property qc_score: float[source]¶
Fraction of fields seg_qc scored
'ok';nanwithout QC.A quality-control verdict on this model’s own masks — it says the masks are not obviously broken, not that they are right. There is no ground truth in a benchmark (see the module docstring), so this is as close to a score as the zoo will produce.
- class spacr.model_zoo.FieldBenchmark[source]¶
One field’s result for one model.
- Parameters:
field – the field name.
n_objects – labels in the mask this model produced.
severity –
spacr.seg_qc’s verdict —'ok','warn'or'fail', or'-'when QC was off.flags – the named defects seg_qc raised.
note – seg_qc’s verdict in prose, with its numbers in it.
- class spacr.model_zoo.ModelEntry[source]¶
One model the zoo knows about, wherever it lives.
Frozen because an entry is a record of a file at a moment — the hash, the size and the provenance describe those bytes. Changing one in place would silently invalidate the other two;
dataclasses.replace()makes the new record explicit.- Parameters:
key – stable id, unique within a listing. For a local file this is derived from the filename; for a catalogue entry it is whatever the catalogue declared.
name – the filename (or the published name) — what a human reads.
kind –
'cellpose'or'classifier'; seeKINDS.source –
'bundled'(ships with spaCR),'local'(found on this machine) or'remote'(declared in a catalogue, not yet fetched).path – absolute path on this machine, or
''for a remote entry.uri – where a remote entry is fetched from, or
''.version – the zoo’s own version number for a filename.
'1'for a plain name,'2'forfoo_v2.CP_model(seeversioned_path()), or whatever a catalogue declared.sha256 – hex digest. For a downloaded model this is the digest of the bytes that were actually written; for a catalogue entry it is the published digest to check against;
''means “no checksum known”, whichfetch()treats as a refusal rather than a pass.size_bytes – file size,
0when unknown.trained_on – what data produced this model, in prose, or
UNKNOWN. Never''.trained_by – who produced it, or
UNKNOWN. Never''.metrics – whatever numbers came with it — for a classifier, the best/last epoch metrics
spacr.train_compare.load_run()recovered. Excluded from equality: two records of the same bytes are the same model whether or not somebody attached numbers to one of them.notes – everything the reader needs to know that is not a field: missing provenance, an unverified download, a file that does not look like a checkpoint.
verified – True only when
sha256was checked against a published digest. A downloaded file whose hash was merely recorded is not verified, and says so.settings_path – where the provenance came from, for the reader who wants to go and look at it.
licence – the licence the model or package is published under, as its publisher states it (an SPDX identifier where there is one), or
''when none is recorded.
- __post_init__()[source]¶
Fill in the provenance fields and validate the kind.
A blank
trained_onortrained_byreads as “no constraints”, so it is replaced with an explicit unknown – the field has to say so out loud rather than by omission.- Raises:
ValueError – if
kindis not one of the known model kinds.
- scorecard_lines() List[str][source]¶
The scorecard as display lines, or the sentence saying there is none.
ONE SOURCE, FOUR RENDERINGS. The tooltip, the API page, the Zoo screen and the Hugging Face table all render THIS, so they cannot disagree – 366 found six README tiles pointing at three different API pages, and that is what happens when a number is written down in more than one place.
- property checksum_state: str[source]¶
What the checksum column says, in one word.
'none'no hash at all — nothing can be checked, and
fetch()refuses such an entry unless the caller overrides it.'published'a hash came with the entry but the bytes are not here yet, so it is a promise about what will arrive.
'recorded'the honest middle: the hash of the file on disk is known, but nobody published one to compare it against. It proves the file has not changed since we looked, and nothing more.
'verified'the bytes on disk were compared with a published digest and match.
- property model_card_url: str[source]¶
The Hugging Face page for this model, derived from its download uri.
A checksum and a metrics table are not enough on their own: the reader wants the page that says what the model was trained on and shows its training curves. Derived rather than declared, so every Hugging Face entry has one without a per-entry field to forget.
- property scorecard_holdout: str[source]¶
name @ versionof the hold-out set, or"".A SCORECARD WITHOUT ITS SET IS A NUMBER WITHOUT A UNIT. Two people quoting an F1 for the same model have said nothing to each other unless they scored the same masks, so the set travels with the numbers into every surface that shows them.
- property scorecard_known: bool[source]¶
True when this model says how accurate it is.
The accuracy twin of
provenance_known, which 370 asks for by name. A model with no numbers did not score zero, and a table of empty cells reads as the second – so the absence is a state to report rather than a gap to render.
- spacr.model_zoo.benchmark(entry: ModelEntry, images: Sequence[Any] | None = None, source: Any = None, n_fields: int = DEFAULT_N_FIELDS, field_names: Sequence[str] | None = None, segment_fn: Callable | None = None, settings: Mapping[str, Any] | None = None, object_type: str = 'cell', qc: bool = True, keep_images: bool = True, channel: int | None = None, progress: Callable[[str, int, int], None] | None = None) BenchmarkResult[source]¶
Run one model over N fields and report what came out. “Test on 3 fields”.
This is
spacr.model_compare’s harness with one model instead of two: the sameload_fields()reader, the sameModelConfig(so the same arguments are honoured and the same ones reported as ignored), the samesegment_with_cellpose()backend, and the samespacr.seg_qcscorecards. To put two models side by side usecompare_entries(), which callscompare_models()proper.The checkpoint is checked before it is loaded, so a missing or corrupt file fails with its own name in the message rather than a torch
KeyError.- Parameters:
entry – the model to run.
images – fields already in memory; None loads them from
source.source – a folder of fields (
.tif/.png/.npy/.npz), read byspacr.model_compare.load_fields().n_fields – how many fields to take from
source.field_names – names for the rows.
segment_fn –
fn(images, config) -> masks; defaults tospacr.model_compare.segment_with_cellpose(). This is the seam the GUI and the tests use, and the reason no test here loads Cellpose.settings – eval overrides (
diameter,flow_threshold, …).object_type – what is being segmented, for the seg_qc scorecards.
qc – score the masks with
spacr.seg_qc.keep_images – keep images and masks on the result for a GUI to draw.
channel – index into the last axis for multi-channel fields.
progress –
fn(message, done, total).
- Returns:
- Raises:
ModelUnreadable – when the checkpoint is missing or not a checkpoint.
ValueError – when there is no field, or the model returned the wrong number of masks.
- spacr.model_zoo.bioimageio_entries(timeout: float = 5.0, url: str | None = None, allow_network: bool = False) List[ModelEntry][source]¶
Every Cellpose model bioimage.io publishes, as zoo rows.
Read from bioimage.io’s own collection, so the category cannot go stale the way a typed list would. Three kinds of row:
a Cellpose-SAM or Cellpose-DINO model, which spaCR’s own Cellpose 4 loads: a
cellposerow that downloads its weights;a Cellpose 3-format checkpoint: a
cellpose3row that downloads its weights for the Cellpose 3 backend, which runs it as it runs cyto3;one of either that spaCR cannot run, which says why in its first note (
_bioimageio_cannot_run()) and offers nothing to download.
Each weights file is checked against the SHA-256 its manifest publishes, and carries the licence the uploader chose and what it was trained on.
Best effort and never raises: no network, a slow mirror or a changed schema all mean fewer rows, never a zoo that fails to open. With
allow_networkthe collection is refreshed once it is a day old, and the weights files are measured, so a stand-in file is refused before anyone downloads it.
- spacr.model_zoo.catalogue(include_bundled: bool = True, remote: bool = True, catalogue_path: Any = None, include_plugins: bool = True, block: bool | None = None) List[ModelEntry][source]¶
Everything the zoo knows about without scanning the user’s disks.
That is: the models bundled with the installed package (whatever
spacr.utils.download_models()has put inresources/models), plus the declared remote entries —BUNDLED_REMOTE_MODELSand, if one is configured, the JSON catalogue named bycatalogue_pathor theCATALOGUE_ENV_VARenvironment variable.Local apart from one thing, and the exception used to be undocumented: with
remoteon, this also asksshared_catalogue()for the community rows, which is a network call. It works offline either way – that fetch never raises – and it never blocks Qt’s GUI thread, whichshared_catalogue()enforces for itself.- Parameters:
include_bundled – list the models in the package resources folder.
remote – list declared remote entries.
catalogue_path – a JSON catalogue to add; defaults to
$SPACR_MODEL_CATALOGUEwhen that names a file.include_plugins – include entries returned by installed spaCR model providers. Provider failures are recorded in plugin diagnostics and do not hide built-in entries.
block – passed to
shared_catalogue().Falsetakes the community rows from its cache rather than waiting for the network;Nonelets that function decide from the thread it is on.
- Returns:
bundled entries first, then remote ones already present locally are dropped (a downloaded model is listed once, as the local file).
- spacr.model_zoo.classify_kind(path: Any) str | None[source]¶
Say whether a file is a Cellpose model, a classifier, or not a model.
The rules, in order:
*.CP_modelis a Cellpose checkpoint — that is whatspacr.submodules.train_cellpose()names its output.*.pth/*.ptis a classifier checkpoint (spacr.io._save_model()writes<model_type>_epoch_<n>_channels_<ch>.pth) unless it sits in a Cellpose folder or hascellpose/cp_modelin its name.An extensionless file inside a Cellpose folder is a Cellpose checkpoint only if its first bytes are a torch save.
cellpose.trainwrites<save_path>/models/<name>with no suffix, and that folder also holds READMEs and logs — the magic-byte check is what keeps aREADMEout of the zoo.Anything else is not a model. CSVs, PNGs,
.npymasks and settings snapshots all land here and are ignored.
- Parameters:
path – a file path.
- Returns:
'cellpose','classifier'or None.
- spacr.model_zoo.community_entries(allow_network: bool = False, repo: str = COMMUNITY_REPO) List[ModelEntry][source]¶
Unvetted models uploaded by spaCR users, or an empty list.
These are shown only when the user asks for them, because an unreviewed checkpoint sitting beside a measured one invites the reader to treat them alike. Every row carries
COMMUNITY_WARNING, and the checksum comes from the uploader’s own submission record – it proves the file has not changed since it was uploaded, NOT that it is any good.Cache-first for the same reason as the bioimage.io listing: catalogue() must not reach the network.
- spacr.model_zoo.compare_entries(entry_a: ModelEntry, entry_b: ModelEntry, images: Sequence[Any] | None = None, source: Any = None, n_fields: int = DEFAULT_N_FIELDS, field_names: Sequence[str] | None = None, settings_a: Mapping[str, Any] | None = None, settings_b: Mapping[str, Any] | None = None, **kwargs: Any)[source]¶
Put two zoo entries head to head on the same fields.
Straight delegation to
spacr.model_compare.compare_models()— the metrics, the split/merge attribution and the “neither model is ground truth” wording all come from there, unchanged. This function’s only job is turning twoModelEntryobjects into twoModelConfigobjects.- Parameters:
entry_a – the A side.
entry_b – the B side.
images – fields already in memory; None loads them from
source.source – a folder of fields.
n_fields – how many fields to take from
source.field_names – names for the rows.
settings_a – eval overrides for A.
settings_b – eval overrides for B.
kwargs – forwarded to
spacr.model_compare.compare_models().
- Returns:
- spacr.model_zoo.config_for(entry: ModelEntry, overrides: Mapping[str, Any] | None = None)[source]¶
The
spacr.model_compare.ModelConfigthat runs this entry.A local checkpoint is passed by path, which
spacr.utils._choose_model()andspacr.model_compare.segment_with_cellpose()both load aspretrained_model; anything else goes through by name and is subject to Cellpose 4’s legacy-name remapping, which the config reports.- Parameters:
entry – the model.
overrides – eval settings (
diameter,flow_threshold, …).
- Returns:
the config.
- spacr.model_zoo.default_local_roots() List[pathlib.Path][source]¶
Folders worth scanning when the caller has not named one.
The bundled pack, the Cellpose user folder, and spaCR’s own model cache. Only the ones that exist come back.
- spacr.model_zoo.discover_local(roots: Any = None, max_depth: int = DEFAULT_SCAN_DEPTH, compute_hashes: bool = False, limit: int = DEFAULT_SCAN_LIMIT) List[ModelEntry][source]¶
Find the model checkpoints already on this machine.
Cellpose models and classifier checkpoints are told apart by
classify_kind(); everything else in the folders — settings CSVs, mask.npyfiles, montage PNGs, logs — is ignored.Nothing is downloaded and nothing is hashed unless
compute_hashesis set: this is the function behind a list widget, and it has to be fast enough to run on a folder the user just typed.- Parameters:
roots – a folder, a file, or an iterable of them; None uses
default_local_roots().max_depth – how deep below each root to look.
compute_hashes – hash every file found (minutes on a big folder).
limit – stop after examining this many files per root.
- Returns:
entries, Cellpose first, then by name.
- spacr.model_zoo.download_bundled_models(**kwargs: Any) str[source]¶
Pull the bundled Hugging Face model pack via the existing downloader.
Thin, deliberate wrapper over
spacr.utils.download_models()— the downloader spaCR already ships and the onespacr.submodules.analyze_plaques()depends on. It is not reimplemented here, so there is one code path that fillsresources/modelsand one place to fix when the repo moves.It is also the unverified path: that function has no checksum, writes straight to the destination filename, and skips the whole pull when the folder is non-empty. Prefer a catalogue entry with a hash and
install(); this exists so the zoo can offer the legacy pack rather than pretend it does not exist.spacr.utilsimports torch, so it is imported here and not at module level. Nothing else in this module reaches for it.- Parameters:
kwargs – forwarded to
spacr.utils.download_models().- Returns:
the local directory the pack landed in.
- spacr.model_zoo.entries_from_sources(entries: Iterable[Any], sources: Iterable[str]) List[Any][source]¶
The rows belonging to the headings that are on, in the given order.
- Parameters:
entries – the rows to filter.
sources – the headings currently on.
- spacr.model_zoo.entry_from_file(path: Any, kind: str | None = None, source: str = 'local', key: str | None = None, runs: Mapping[str, Any] | None = None, compute_hash: bool = False, sha256: str = '', verified: bool = False, extra_notes: Sequence[str] = ()) ModelEntry[source]¶
Build a
ModelEntryfor a checkpoint on this machine.Provenance is recovered from whatever spaCR already wrote beside the model: a
*_settings.csvor<src>/settings/<name>.csvfor a Cellpose model, and — for a classifier — the training run the checkpoint sits in, loaded throughspacr.train_compare.load_run()so the zoo and the training-run comparison agree about where settings live and what they say.- Parameters:
path – the checkpoint.
kind – override
classify_kind().source –
'local'or'bundled'.key – override the generated key.
runs –
{folder: TrainingRun}from_runs_under(), so a scan of 40 checkpoints in one run folder reads that folder once.compute_hash – hash the file now. Off by default: hashing every checkpoint on a machine to populate a list widget is minutes.
sha256 – a digest already known for these bytes.
verified – whether
sha256was checked against a published digest.extra_notes – notes to carry onto the entry.
- Returns:
the entry.
- Raises:
ModelUnreadable – when the path is not a file.
- spacr.model_zoo.fetch(entry: ModelEntry, dest: Any, expected_sha256: str | None = None, require_checksum: bool = True, opener: Callable[[str], Any] | None = None, progress: Callable[[int, int], None] | None = None, cancel: Callable[[], bool] | None = None, chunk_size: int = DEFAULT_CHUNK, timeout: int = DEFAULT_TIMEOUT) pathlib.Path[source]¶
Download a model, verify it, and only then put it where it belongs.
The order is the whole point:
bytes stream into a temporary file inside
dest, so the rename in step 4 is a same-filesystemos.replaceand therefore atomic;the checksum of what actually arrived is computed;
if it does not match the published digest the temporary file is deleted and
ChecksumMismatchis raised — nothing is installed, and the destination still holds whatever it held before;only now is the temporary file renamed, to a
versioned_path()that does not exist yet.
Every failure — a dead server, a cancel, a bad hash, a full disk — leaves the destination directory exactly as it was. There is no window in which a half-written file sits at a name that looks like a model.
- Parameters:
entry – what to fetch.
ModelEntry.uriis the source.dest – destination directory; created if missing.
expected_sha256 – digest to require, overriding
ModelEntry.sha256.require_checksum – refuse to install when no digest is known. True by default — a download nobody can check is exactly the thing this module exists to stop being routine. Pass False to accept one knowingly; the entry
install()returns then reportsverified=False.opener –
fn(uri) -> chunksorfn(uri) -> (chunks, total); defaults toopen_uri().progress –
fn(done_bytes, total_bytes);totalis 0 when the server did not say.cancel –
fn() -> bool, polled between chunks. Returning True deletes the partial file and raisesDownloadCancelled.chunk_size – bytes per read.
timeout – seconds, HTTP only.
- Returns:
the path the model was written to.
- Raises:
ChecksumMismatch – the bytes are not the published bytes.
DownloadCancelled –
cancel()returned True.ModelZooError – no URI, or no checksum with
require_checksum.
- spacr.model_zoo.fieldset_id(names: Sequence[str], images: Sequence[Any]) str[source]¶
A stable id for a set of fields, taken from the pixels.
Folder names are not identity:
plate1/1on two machines is two different sets of images, and the same three images copied to a new folder are the same benchmark input. So the id hashes each array’s bytes, shape and dtype together with its name.This is what makes
rank()able to refuse. Without it, two benchmarks run on different data are two numbers, and two numbers always sort.- Parameters:
names – field names, in order.
images – the arrays, in the same order.
- Returns:
a 16-character hex id.
- spacr.model_zoo.format_benchmarks(results: Sequence[BenchmarkResult], key: str = DEFAULT_RANK_KEY) str[source]¶
Render benchmarks grouped by field set, ranked only within a group.
Two models benchmarked on different fields appear under two headers with a line saying the two blocks cannot be compared. That is the alternative to
rank()’s refusal, and it is the only way this module will ever put incomparable numbers on the same page.- Parameters:
results – benchmarks, from any number of field sets.
key – one of
RANK_KEYS.
- Returns:
a multi-line string.
- spacr.model_zoo.format_zoo(entries: Sequence[ModelEntry]) str[source]¶
Render a listing a human reads before choosing a model.
Provenance is a column, not a footnote: “trained on” is the field that decides whether a model is applicable to your images at all, and it is printed for every row — reading
unknownwhere it is unknown, because a blank there would read as “no constraints”.- Parameters:
entries – what to list.
- Returns:
a multi-line string.
- spacr.model_zoo.group_by_fieldset(results: Sequence[BenchmarkResult]) Dict[str, List[BenchmarkResult]][source]¶
Bucket benchmarks by the field set they ran on, first-seen order.
- Parameters:
results – benchmarks.
- Returns:
{fieldset_id: [results]}.
- spacr.model_zoo.group_by_source(entries: Iterable[Any]) Dict[str, List[Any]][source]¶
Split a listing into
ZOO_SOURCES, keeping each source’s order.Every heading is present even when it has no rows, so a caller drawing the strip does not have to know which of the five happened to be empty this time.
- Parameters:
entries – the rows to split.
- Returns:
heading -> rows, in
ZOO_SOURCESorder.
- spacr.model_zoo.hf_uri(repo_id: str, filename: str, repo_type: str = 'dataset') str[source]¶
The download URL for a file in a Hugging Face repo.
- Parameters:
repo_id – Hugging Face repository identifier.
filename – repository-relative name of the file to download.
repo_type –
"dataset"(the default, and what spaCR shipped first) or"model".
Exactly the URL
spacr.utils.download_models()andspacr.qt.hf_download._download_one()build, kept in one place so the zoo cannot drift away from the downloader spaCR already ships.THE TWO REPO KINDS HAVE DIFFERENT URLS, which is not cosmetic: a dataset file lives under
/datasets/<repo>/resolve/...and a model file under/<repo>/resolve/.... Asking for one at the other’s URL returns a 404 page, and a downloader that does not check the content type writes that HTML into the destination and leaves a “checkpoint” that fails to load with a torch error naming neither the URL nor the repo.datasetremains the default becauseHF_MODELS_REPOis a DATASET repo –einarolafsson/models– and every entry written before this parameter existed assumes it. New model repos pass"model".
- spacr.model_zoo.inspect_checkpoint(path: Any, loader: Callable[[str], Any] | None = None, deep: bool = False) Dict[str, Any][source]¶
Check a file is a loadable checkpoint, failing with the filename in it.
The default failure for a wrong or corrupt checkpoint is a
KeyErroron a state-dict key raised somewhere inside torch, which names nothing the user chose and reads like a spaCR bug. This turns all of it — missing, empty, truncated, a PNG somebody renamed, a Cellpose model handed to the classifier path — into oneModelUnreadablenaming the file.The shallow check needs no torch at all: it is a stat and four bytes.
- Parameters:
path – the checkpoint.
loader –
fn(path) -> objectused for the deep check; defaults totorch.load(..., map_location='cpu'), imported only if used.deep – actually load the file. Off by default because loading a 2 GB checkpoint to populate a list widget is not acceptable.
- Returns:
{'path', 'size_bytes', 'format', 'loaded'}.- Raises:
ModelUnreadable – naming the file, always.
- spacr.model_zoo.install(entry: ModelEntry, dest: Any, **kwargs: Any) ModelEntry[source]¶
fetch()the model and return the registered local entry.The returned entry carries the digest of the bytes that were actually written — not the one the catalogue claimed — and
ModelEntry.verifiedis True only when the two were compared and matched. Provenance from the catalogue entry is carried over, because that is the whole reason for having had a catalogue.- Parameters:
entry – the remote entry.
dest – destination directory.
kwargs – passed to
fetch().
- Returns:
a
source='local'entry pointing at the new file.
- spacr.model_zoo.installable_backend_entries() List[ModelEntry][source]¶
Every optional segmentation backend, in whatever state it is here.
A backend absent from the zoo teaches nobody that it exists, so each one is listed, and its
sourcesays where it stands –installed,installable,installingornot installable here– with the reason as its first note and its licence on the row. Installing one builds it an environment of its own under~/.spacr/backendsand leaves spaCR’s own environment alone.
- spacr.model_zoo.load_catalogue_file(path: Any) List[ModelEntry][source]¶
Read a JSON catalogue of remote models.
Format — a list, or an object with a
modelslist:{"models": [ {"key": "hela_60x", "name": "hela_60x_confluent.CP_model", "kind": "cellpose", "uri": "https://…/hela_60x_confluent.CP_model", "sha256": "9f86d0…", "size_bytes": 26566572, "trained_on": "HeLa, 60x, confluent monolayer, 512px crops", "trained_by": "A. Researcher, 2026-02", "metrics": {"note": "benchmarked on plate3 fields 1-3"}} ]}
sha256is the field that decides whether the entry is usable without an explicit override, so a catalogue is worth exactly as much as its hashes.- Parameters:
path – the JSON file.
- Returns:
the entries.
- Raises:
ModelZooError – when the file cannot be read or is not a catalogue, naming the file.
- spacr.model_zoo.open_uri(uri: str, timeout: int = DEFAULT_TIMEOUT, chunk_size: int = DEFAULT_CHUNK) Tuple[Iterable[bytes], int][source]¶
Open a model URI for streaming.
(chunks, total_bytes).http://andhttps://stream overrequests— the same callspacr.utils.download_models()andspacr.qt.hf_download._download_one()make, imported here so this module has no hard dependency on it.file://and a plain existing path are read from disk, which is what a lab mirror on a NAS looks like and what the tests use, so the whole fetch path is exercised without a network.- Parameters:
uri – where the model lives.
timeout – seconds, HTTP only.
chunk_size – bytes per chunk.
- Returns:
(iterable of byte chunks, total size or 0 when unknown).- Raises:
ModelZooError – for a scheme this does not speak.
- spacr.model_zoo.package_model_root() pathlib.Path[source]¶
<spacr>/resources/models— where the bundled pack lives.The same folder
spacr.utils.download_models()fills andspacr.submodules.analyze_plaques()reads from.
- spacr.model_zoo.publish_model(local_path: Any, repo_id: str, *, key: str, kind: str = 'cellpose', trained_on: str = UNKNOWN, trained_by: str = UNKNOWN, private: bool = False, notes: Sequence[str] = ()) Dict[str, Any][source]¶
Upload a model to Hugging Face and return its catalogue row.
- Parameters:
local_path – the checkpoint to upload.
repo_id –
<user>/<repo>– YOUR OWN account.key – the short name spaCR will offer the model under.
kind – one of
KINDS.trained_on – what the model was trained on. Say it properly: this is the only thing another lab has to decide whether it applies to them.
trained_by – who trained it, and roughly when.
private – keep the repo private. A private model cannot be fetched by other spaCR users, so it is off by default.
notes – caveats worth carrying next to the model.
- Returns:
the catalogue row, with the sha256 filled in.
- Raises:
ImportError – when
huggingface_hubis not installed.
THE CHECKSUM IS COMPUTED HERE, from the file that was actually uploaded, which is the whole reason this exists as a function rather than as instructions in a README.
fetch()refuses an entry it cannot verify, so a row written by hand without a hash produces a model nobody can install without disabling the check – which is what the one pre-existing bundled entry does, and it is a hole rather than a precedent.Publishing does NOT distribute the model on its own: add the returned row to the shared catalogue (
REMOTE_CATALOGUE_URI) and every spaCR user sees it withinCATALOGUE_CACHE_SECONDS.
- spacr.model_zoo.rank(results: Sequence[BenchmarkResult], key: str = DEFAULT_RANK_KEY) List[BenchmarkResult][source]¶
Order benchmarks best-first — within one field set only.
A model’s numbers on your three fields say nothing about its numbers on somebody else’s: different cell density, different exposure, different magnification. Sorting results from two field sets into one list produces a ranking that looks exactly like a real one and means nothing, which is the failure this function exists to prevent. So it refuses.
- Parameters:
results – benchmarks, all from the same field set.
key – one of
RANK_KEYS.
- Returns:
the results, best first.
- Raises:
IncomparableBenchmarks – when the results span more than one field set. Use
rank_groups()orformat_benchmarks(), which group and label instead.ValueError – on an unknown
key.
- spacr.model_zoo.rank_groups(results: Sequence[BenchmarkResult], key: str = DEFAULT_RANK_KEY) Dict[str, List[BenchmarkResult]][source]¶
Rank inside each field set, keeping the sets apart. The safe alternative.
- Parameters:
results – benchmarks from any number of field sets.
key – one of
RANK_KEYS.
- Returns:
{fieldset_id: [results, best first]}.
- spacr.model_zoo.resolve(key_or_path: Any, entries: Sequence[ModelEntry] | None = None) ModelEntry[source]¶
Turn a key, a name or a path into a
ModelEntry.A path that exists wins over a key: pointing the zoo at a file you just trained has to work without registering it anywhere first.
- Parameters:
key_or_path – an entry key, a model filename, or a path to a file.
entries – the listing to search; defaults to
catalogue().
- Returns:
the entry.
- Raises:
ModelUnreadable – when it looks like a path and no file is there.
ModelZooError – when no entry matches, listing the near misses.
- spacr.model_zoo.scorecard_html(entry) str[source]¶
The model’s scorecard as an HTML table, for a tooltip.
A paragraph of prose is what a tooltip used to show, and a reader comparing two models had to parse two paragraphs to find two numbers. The same table the model card prints answers that at a glance. Falls back to the prose when an entry publishes no metrics, because an empty table is worse than a sentence.
Metrics that hold none of the scorecard’s keys – a free-form note, a training loss under a name of its own – are not a scorecard either, and return nothing too: a table of eleven “not recorded” rows would replace a two-line note that said something.
- Parameters:
entry – a catalogue entry, or anything with
metricsand a name.- Returns:
the HTML table, or
""when there is no scorecard to show.
- spacr.model_zoo.sha256_file(path: Any, chunk_size: int = 1 << 20) str[source]¶
Hex SHA-256 of a file, read in chunks so a 2 GB checkpoint is not RAM.
- Parameters:
path – the file.
chunk_size – bytes per read.
- Returns:
the lowercase hex digest.
- Raises:
ModelUnreadable – when the file is missing or cannot be read, with the path in the message.
The community catalogue, fetched from
REMOTE_CATALOGUE_URI.- Parameters:
uri – override the catalogue location.
timeout – seconds to wait for the request.
force – ignore the cache and re-fetch.
block – whether to wait for the network.
None– the default, and what an unthinking caller gets – waits everywhere EXCEPT Qt’s GUI thread, where it answers from the cache and refreshes on a daemon thread.Truewaits wherever it is called, which only a caller that knows it is on a worker or in a CLI may ask for.Falsenever waits.
- Returns:
the entries, or
()when the catalogue cannot be read.
NEVER FETCHES ON THE GUI THREAD, and that is not an optimisation. This is reached from
spacr.settings.downloaded_zoo_modelswhile a settings panel is being built, so it ran insideMainWindow._on_nav_selectedwith nothing able to paint or answer the compositor. Measured with the catalogue host non-routable (10.255.255.1, the shape of a down VPN or a captive portal – the connect neither completes nor is refused): opening the Mask module took 32.2 s, all of it a GUI thread stuck inurlopen. GNOME asks a window whether it is alive after five, so what the user sees is spaCR’s “force quit” dialog, which is how this was reported. With the fetch moved off the thread the same open is 2.4 s.The cost of not waiting is a first module open whose Cellpose dropdown lists the bundled and local models but not the community ones; the background refresh means the second one has them.
NEVER RAISES, and that is deliberate. This runs when a user opens a module that offers a model list, and the list is useful without it: the bundled entries and any local models are still there. A laptop on a train, a lab behind a proxy and a Hugging Face outage all produce the same thing – a shorter list and a log line – rather than a module that will not open.
The failure that WOULD be silent and harmful is a corrupt or hostile catalogue, so entries that do not parse are dropped individually, and an entry without a
sha256still cannot be installed byfetch()without an explicit override. A catalogue row is a claim about where a file lives; the checksum is what makes it a claim about which file.
Whether
shared_catalogue()would go to the network to answer.For a caller that wants to do the waiting somewhere it is allowed to – a worker thread – rather than get the cached answer and not know it was one.
- spacr.model_zoo.source_of(entry: Any) str[source]¶
Which of
ZOO_SOURCESthis row belongs under.Decided FROM THE ROW, never from a list of names kept somewhere else: a hand-written list goes stale the first time a model is added, and the failure it produces is a model that is in the catalogue and under no heading, which is a model nobody can see.
The order of the tests is the rule, and it matters in one place: a Cellpose 3 checkpoint published on bioimage.io is a bioimage.io row, not a Cellpose 3 one.
cellpose3means the backend’s OWN models – cyto, cyto2, cyto3, nuclei – and the backend package that runs them.A row that matches nothing is filed under
spaCRAND SAID OUT LOUD. It is the fallback rather than a sixth heading because a model under the wrong heading is a nuisance and a model under no heading is a bug the user experiences as a missing model.- Parameters:
entry – any zoo row – a
ModelEntry, or anything withkind,sourceanduri.- Returns:
one of
ZOO_SOURCES.
- spacr.model_zoo.stock_cellpose_entries() List[ModelEntry][source]¶
Every model the installed Cellpose can fetch for itself.
These are not spaCR’s files and carry no checksum of ours: Cellpose downloads and verifies them, and the name IS the path – passing “cpsam” to Cellpose resolves it. They are listed so that the zoo answers “what can I segment with” rather than “what has Einar trained”, which is the question a new user actually has.
- spacr.model_zoo.verify(entry: ModelEntry, expected: str | None = None) bool[source]¶
Hash the file this entry points at and compare it to a known digest.
- Parameters:
entry – the entry to check.
expected – the digest to compare against; defaults to
ModelEntry.sha256.
- Returns:
True when the file’s digest matches.
- Raises:
ModelUnreadable – when the entry has no local file, naming it.
ModelZooError – when there is no digest to compare against — that is a caller error, and returning False for it would read as “the file is wrong” when what happened is “nobody said what right looks like”.
- spacr.model_zoo.versioned_path(dest: Any, filename: str) pathlib.Path[source]¶
The first free destination for
filenameindest.foo.CP_model->foo.CP_model, thenfoo_v2.CP_model,foo_v3.CP_model… An existing checkpoint is never overwritten: two models with the same filename are a normal thing to have (the same author retrained, or two people picked the same name), and the failure mode of overwriting — a run that used the old weights becoming unreproducible with no trace — is silent.An input that already carries
_vNcounts from there rather than becomingfoo_v2_v2.- Parameters:
dest – destination directory.
filename – the name to place there.
- Returns:
a path that does not exist yet.
Nested helpers¶
- _default_segmenter._segment(images, config)¶
Segment
imageswith the backend the config’s model belongs to.spacr/model_zoo.py:3923
Fetch the catalogue off the GUI thread, and always release the flag.
The
finallyis the whole point:_SHARED_CATALOGUE_FETCHINGis what stops a second refresh being started while this one is in flight, so a fetch that raises must still clear it or no later refresh can ever begin.spacr/model_zoo.py:1969
- benchmark._tick(message: str, done: int) None¶
Report one benchmark milestone through the captured callback.
- Parameters:
message – stage description for the progress display.
done – completed-step index from zero through two.
- Returns:
None. When a callback was supplied it receives the message, completed index, and captured total of two; otherwise this is a no-op.
spacr/model_zoo.py:3834
- scorecard_html.cell(value)¶
A scorecard value as shown, or “not recorded” when blank.
spacr/model_zoo.py:4127