spacr.example_archives¶
The published example datasets: what they are, and how to fetch them.
WHY THIS IS NOT IN spacr.qt.hf_download. That module owns the same
downloads with a Qt progress dialog wrapped round them, and it imports PySide6
at module scope. spacr-download has to work on a cluster login node with no
display and no Qt installed at all, so everything that is only about the DATA –
which repositories publish it, what each archive is called, how a stream is
verified, how an archive is safely unpacked, and what has to be repaired on
arrival – lives here instead. hf_download imports these names back out, so
every existing GUI caller (and every test that patches one of them on that
module) keeps working unchanged.
NOT TO BE CONFUSED WITH spacr.example_data, which fetches the Regression
example screen’s count and score CSVs from a GitHub release. That is a
different set of files, a different host and a different transport. This module
is about the .tar archives of imaging data: the Mask demo plate, the
Measure example, the Annotate/Classify example, the Replication and
Recruitment assay examples, and – through spacr.screen_data – the
pieces of the published TSG101 screen.
SIZES ARE STATED BEFORE ANYTHING IS FETCHED. A user choosing between example
sets is choosing how many gigabytes to spend, and a picker (or a CLI) that
lists names without sizes makes that choice blind. EXAMPLE_SETS carries
one size per set for the same reason spacr.screen_data.ScreenAsset
does.
Classes¶
One published example dataset, described before it is downloaded. |
Functions¶
|
Stream one file from the HF repo to |
|
The ONE folder every example dataset unpacks into. |
|
The set called |
|
Where the set called |
|
Write each |
|
Turn a download exception into something a user can act on. |
|
Unpack a downloaded example archive under |
Point a downloaded example at where it actually landed. |
Module Contents¶
- class spacr.example_archives.ExampleSet[source]¶
One published example dataset, described before it is downloaded.
- Parameters:
key – what a user types to ask for this one.
repo – the Hugging Face dataset repository it is published in.
summary – one line saying what the set is and what it is for.
bytes – the archive’s size. APPROXIMATE, and deliberately so: these are the figures the GUI buttons already state (“about 280 MB”), taken from the archives at publication. They exist to let someone decide whether to spend the download, not to be checked against. The exact length is verified during the transfer, against what the server declares.
markers – what the unpacked set LEAVES BEHIND, as globs relative to the folder it unpacks into. Presence is tested on these rather than on the folder, because all three sets share one plate directory: the directory existing says nothing about which of them is in it.
expands_npz – whether
.npzarrays have to be written back out as the.npyMeasure reads. Seeexpand_measure_arrays().in_default – whether
spacr-downloadwith no arguments fetches it. False for the OPS and Align & Stitch samples: together they are 0.6 GB of one screen’s tiles, and adding them would push the default run pastCONFIRM_ABOVE_BYTES, so the command would start asking to confirm the thing it does with no arguments. They are fetched when named, and byall.folder – the folder under the example-data root the set unpacks into.
plate1– the shared plate – for the three sets that are stages of one plate; a folder of its own for a set that would collide with them.
- is_present(folder) bool[source]¶
Whether this set is already unpacked under
folder.Every marker has to match. A half-unpacked set – a transfer that died between the images and the database – is therefore reported as absent, which is the answer that gets it repaired.
- Parameters:
folder – the plate directory the set unpacks into.
- spacr.example_archives.download_archive(repo_id: str, file_name: str, dest_dir: pathlib.Path, *, progress: Callable[[int, int | None], None] | None = None, chunk_size: int = 1 << 15) pathlib.Path[source]¶
Stream one file from the HF repo to
dest_dir/basename.Uses plain HTTP + streaming so we don’t need the full
hf_hubdownload machinery (and its cache dir) for a one-shot demo pull.The body lands in a sibling
.partfile and is only moved onto the final path once every advertised byte has arrived. Writing straight to the destination meant a dropped connection left a truncated image behind that was indistinguishable from a good download — the next pipeline run then failed deep inside the mask stage instead of at the download.- Parameters:
repo_id – the Hugging Face dataset repository.
file_name – the path within it; only the basename is kept on disk.
dest_dir – the folder the file lands in.
progress – called as
progress(written, expected)after each chunk, whereexpectedisNonewhen the server declared no length. A CLI draws a percentage from it; the GUI does not use it, because its workers emit Qt signals from inside their own loops.chunk_size – 32 KB by default, which is the size the per-file demo pull has always used. An archive of gigabytes passes a bigger one.
- spacr.example_archives.example_plate_folder() pathlib.Path[source]¶
The ONE folder every example dataset unpacks into.
~/.cache/spacr/example_data/plate1– a real spaCR plate directory, holding whichever ofmerged/,data/,measurements/andsettings/have been downloaded.ONE FOLDER BECAUSE THE SETS ARE USED TOGETHER.
data/is the crops andmeasurements/measurements.dbis what indexes them; downloading them into separate trees meant the two halves of one plate could not be opened at once, and the user had to know which download had put what where. Each archive’s members are relative to this folder, so the three unpack into it side by side and compose into a plate that Measure, Annotate and Classify can all be pointed at.The assay examples are not stages of this plate and each has a folder of its own beside it; see
example_set_folder().
- spacr.example_archives.example_set(key: str) ExampleSet[source]¶
The set called
key.- Parameters:
key –
import,mask,measure,annotate,replication,recruitmentorinvasion.- Raises:
KeyError – naming the keys that do exist. A typo that returned
Nonewould download nothing and report success, which is the one outcome a download command must never produce.
- spacr.example_archives.example_set_folder(key: str) pathlib.Path[source]¶
Where the set called
keyunpacks: its folder beside the plate.- Parameters:
key – an
EXAMPLE_SETSkey.- Returns:
~/.cache/spacr/example_data/<folder>. Not created here.- Raises:
KeyError – as
example_set()does.
- spacr.example_archives.expand_measure_arrays(merged: pathlib.Path) None[source]¶
Write each
.npzback out as the.npyMeasure reads.The compression is a TRANSPORT detail – it halves a 700 MB download – and Measure loads
npy. Converting on arrival keeps that entirely inside the downloader rather than teaching every reader about a second format.The
.npzis removed afterwards: keeping both doubles the disk cost of an example dataset for a file nothing will open again.A MODULE FUNCTION, NOT A METHOD, and that is the point.
after_extractruns on the download thread, and reaching this code through_MeasureExampleWorker(dest)._expand_arrays(...)CONSTRUCTED a QObject there purely to borrow a helper.thread_guardreported it exactly as it should have: the object then lived on ‘Dummy-6’ and every later touch from the GUI thread was illegal. Nothing in here ever readself, so there was never an object to need.- Parameters:
merged – the folder the
.npzarrays were unpacked into.
- spacr.example_archives.explain_download_failure(exc: BaseException) str[source]¶
Turn a download exception into something a user can act on.
Downloading test data requires a network connection. What the user saw before was
str(exc), which for the ordinary offline case is a nested urllib3 dump:(MaxRetryError("HTTPSConnectionPool(host='huggingface.co', port=443): Max retries exceeded with url: /api/datasets/... (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7e8d...>: Failed to establish a new connection: [Errno 101] Network is unreachable'))"), '(Request ID: 73ac20ed-...)')
— 300 characters that never say “you are offline” and never say what to do instead. The three conditions this actually fails on are: no network, the
huggingface_hubextra not installed, and a truncated transfer. Each gets a sentence naming the cause and the way out; anything else keeps its own message with the same closing advice attached.- Parameters:
exc – the exception raised inside the download worker.
- Returns:
a multi-line message for the failure dialog.
- spacr.example_archives.extract_example_archive(archive, dest) int[source]¶
Unpack a downloaded example archive under
dest.EXTRACTED WITH
filter="data", which is the whole reason this is a function rather than two lines at the call site. A tar can name../../etc/somethingor an absolute path, and a plainextractallwill happily write there – so unpacking downloaded content without a filter hands whoever can publish to the repo a write anywhere the user can write. The filter rejects those members, along with device nodes, setuid bits and symlinks pointing outside the tree.Python 3.12 and later have it built in. Older interpreters get an explicit check instead of a silent unfiltered unpack.
- Parameters:
archive – the
.taron disk.dest – the folder to unpack into.
- Returns:
how many members were written.
- spacr.example_archives.make_the_example_paths_absolute(root) int[source]¶
Point a downloaded example at where it actually landed.
A measurements database stores ABSOLUTE paths to its crops, which name the machine that made it and resolve nowhere else. The published copy stores them relative to the dataset root instead, so it is portable and carries no account name – and this is what turns them back into paths that open.
The settings files are rewritten the same way: they carry
DATASET_PLACEHOLDERwhere the unpack location goes, so a user can press Run without first editing a path.Idempotent. A path that is already absolute is left alone, so running this twice – a re-download over an existing copy – does not produce
/home/me/data//home/me/data/....- Parameters:
root – the folder the dataset was unpacked into.
- Returns:
how many values were rewritten.