spacr.example_archives

The published example datasets: what they are, and how to fetch them.

WHY THIS IS NOT IN spacr.qt.hf_download. That module owns the same downloads with a Qt progress dialog wrapped round them, and it imports PySide6 at module scope. spacr-download has to work on a cluster login node with no display and no Qt installed at all, so everything that is only about the DATA – which repositories publish it, what each archive is called, how a stream is verified, how an archive is safely unpacked, and what has to be repaired on arrival – lives here instead. hf_download imports these names back out, so every existing GUI caller (and every test that patches one of them on that module) keeps working unchanged.

NOT TO BE CONFUSED WITH spacr.example_data, which fetches the Regression example screen’s count and score CSVs from a GitHub release. That is a different set of files, a different host and a different transport. This module is about the .tar archives of imaging data: the Mask demo plate, the Measure example, the Annotate/Classify example, the Replication and Recruitment assay examples, and – through spacr.screen_data – the pieces of the published TSG101 screen.

SIZES ARE STATED BEFORE ANYTHING IS FETCHED. A user choosing between example sets is choosing how many gigabytes to spend, and a picker (or a CLI) that lists names without sizes makes that choice blind. EXAMPLE_SETS carries one size per set for the same reason spacr.screen_data.ScreenAsset does.

Classes

ExampleSet

One published example dataset, described before it is downloaded.

Functions

download_archive(→ pathlib.Path)

Stream one file from the HF repo to dest_dir/basename.

example_plate_folder(→ pathlib.Path)

The ONE folder every example dataset unpacks into.

example_set(→ ExampleSet)

The set called key.

example_set_folder(→ pathlib.Path)

Where the set called key unpacks: its folder beside the plate.

expand_measure_arrays(→ None)

Write each .npz back out as the .npy Measure reads.

explain_download_failure(→ str)

Turn a download exception into something a user can act on.

extract_example_archive(→ int)

Unpack a downloaded example archive under dest.

make_the_example_paths_absolute(→ int)

Point a downloaded example at where it actually landed.

Module Contents

class spacr.example_archives.ExampleSet[source]

One published example dataset, described before it is downloaded.

Parameters:
  • key – what a user types to ask for this one.

  • repo – the Hugging Face dataset repository it is published in.

  • summary – one line saying what the set is and what it is for.

  • bytes – the archive’s size. APPROXIMATE, and deliberately so: these are the figures the GUI buttons already state (“about 280 MB”), taken from the archives at publication. They exist to let someone decide whether to spend the download, not to be checked against. The exact length is verified during the transfer, against what the server declares.

  • markers – what the unpacked set LEAVES BEHIND, as globs relative to the folder it unpacks into. Presence is tested on these rather than on the folder, because all three sets share one plate directory: the directory existing says nothing about which of them is in it.

  • expands_npz – whether .npz arrays have to be written back out as the .npy Measure reads. See expand_measure_arrays().

  • in_default – whether spacr-download with no arguments fetches it. False for the OPS and Align & Stitch samples: together they are 0.6 GB of one screen’s tiles, and adding them would push the default run past CONFIRM_ABOVE_BYTES, so the command would start asking to confirm the thing it does with no arguments. They are fetched when named, and by all.

  • folder – the folder under the example-data root the set unpacks into. plate1 – the shared plate – for the three sets that are stages of one plate; a folder of its own for a set that would collide with them.

is_present(folder) → bool[source]

Whether this set is already unpacked under folder.

Every marker has to match. A half-unpacked set – a transfer that died between the images and the database – is therefore reported as absent, which is the answer that gets it repaired.

Parameters:

folder – the plate directory the set unpacks into.

property archive: str[source]

The .tar this set ships as.

spacr.example_archives.download_archive(repo_id: str, file_name: str, dest_dir: pathlib.Path, *, progress: Callable[[int, int | None], None] | None = None, chunk_size: int = 1 << 15) → pathlib.Path[source]

Stream one file from the HF repo to dest_dir/basename.

Uses plain HTTP + streaming so we don’t need the full hf_hub download machinery (and its cache dir) for a one-shot demo pull.

The body lands in a sibling .part file and is only moved onto the final path once every advertised byte has arrived. Writing straight to the destination meant a dropped connection left a truncated image behind that was indistinguishable from a good download — the next pipeline run then failed deep inside the mask stage instead of at the download.

Parameters:
  • repo_id – the Hugging Face dataset repository.

  • file_name – the path within it; only the basename is kept on disk.

  • dest_dir – the folder the file lands in.

  • progress – called as progress(written, expected) after each chunk, where expected is None when the server declared no length. A CLI draws a percentage from it; the GUI does not use it, because its workers emit Qt signals from inside their own loops.

  • chunk_size – 32 KB by default, which is the size the per-file demo pull has always used. An archive of gigabytes passes a bigger one.

spacr.example_archives.example_plate_folder() → pathlib.Path[source]

The ONE folder every example dataset unpacks into.

~/.cache/spacr/example_data/plate1 – a real spaCR plate directory, holding whichever of merged/, data/, measurements/ and settings/ have been downloaded.

ONE FOLDER BECAUSE THE SETS ARE USED TOGETHER. data/ is the crops and measurements/measurements.db is what indexes them; downloading them into separate trees meant the two halves of one plate could not be opened at once, and the user had to know which download had put what where. Each archive’s members are relative to this folder, so the three unpack into it side by side and compose into a plate that Measure, Annotate and Classify can all be pointed at.

The assay examples are not stages of this plate and each has a folder of its own beside it; see example_set_folder().

spacr.example_archives.example_set(key: str) → ExampleSet[source]

The set called key.

Parameters:

key – import, mask, measure, annotate, replication, recruitment or invasion.

Raises:

KeyError – naming the keys that do exist. A typo that returned None would download nothing and report success, which is the one outcome a download command must never produce.

spacr.example_archives.example_set_folder(key: str) → pathlib.Path[source]

Where the set called key unpacks: its folder beside the plate.

Parameters:

key – an EXAMPLE_SETS key.

Returns:

~/.cache/spacr/example_data/<folder>. Not created here.

Raises:

KeyError – as example_set() does.

spacr.example_archives.expand_measure_arrays(merged: pathlib.Path) → None[source]

Write each .npz back out as the .npy Measure reads.

The compression is a TRANSPORT detail – it halves a 700 MB download – and Measure loads npy. Converting on arrival keeps that entirely inside the downloader rather than teaching every reader about a second format.

The .npz is removed afterwards: keeping both doubles the disk cost of an example dataset for a file nothing will open again.

A MODULE FUNCTION, NOT A METHOD, and that is the point. after_extract runs on the download thread, and reaching this code through _MeasureExampleWorker(dest)._expand_arrays(...) CONSTRUCTED a QObject there purely to borrow a helper. thread_guard reported it exactly as it should have: the object then lived on ‘Dummy-6’ and every later touch from the GUI thread was illegal. Nothing in here ever read self, so there was never an object to need.

Parameters:

merged – the folder the .npz arrays were unpacked into.

spacr.example_archives.explain_download_failure(exc: BaseException) → str[source]

Turn a download exception into something a user can act on.

Downloading test data requires a network connection. What the user saw before was str(exc), which for the ordinary offline case is a nested urllib3 dump:

(MaxRetryError("HTTPSConnectionPool(host='huggingface.co', port=443):
Max retries exceeded with url: /api/datasets/... (Caused by
NewConnectionError('<urllib3.connection.HTTPSConnection object at
0x7e8d...>: Failed to establish a new connection: [Errno 101] Network
is unreachable'))"), '(Request ID: 73ac20ed-...)')

— 300 characters that never say “you are offline” and never say what to do instead. The three conditions this actually fails on are: no network, the huggingface_hub extra not installed, and a truncated transfer. Each gets a sentence naming the cause and the way out; anything else keeps its own message with the same closing advice attached.

Parameters:

exc – the exception raised inside the download worker.

Returns:

a multi-line message for the failure dialog.

spacr.example_archives.extract_example_archive(archive, dest) → int[source]

Unpack a downloaded example archive under dest.

EXTRACTED WITH filter="data", which is the whole reason this is a function rather than two lines at the call site. A tar can name ../../etc/something or an absolute path, and a plain extractall will happily write there – so unpacking downloaded content without a filter hands whoever can publish to the repo a write anywhere the user can write. The filter rejects those members, along with device nodes, setuid bits and symlinks pointing outside the tree.

Python 3.12 and later have it built in. Older interpreters get an explicit check instead of a silent unfiltered unpack.

Parameters:
  • archive – the .tar on disk.

  • dest – the folder to unpack into.

Returns:

how many members were written.

spacr.example_archives.make_the_example_paths_absolute(root) → int[source]

Point a downloaded example at where it actually landed.

A measurements database stores ABSOLUTE paths to its crops, which name the machine that made it and resolve nowhere else. The published copy stores them relative to the dataset root instead, so it is portable and carries no account name – and this is what turns them back into paths that open.

The settings files are rewritten the same way: they carry DATASET_PLACEHOLDER where the unpack location goes, so a user can press Run without first editing a path.

Idempotent. A path that is already absolute is left alone, so running this twice – a re-download over an existing copy – does not produce /home/me/data//home/me/data/....

Parameters:

root – the folder the dataset was unpacked into.

Returns:

how many values were rewritten.