spacr.qt.make_masks_datasets

A sample of every dataset spaCR has a model for, one dialog away in Make Masks.

Make Masks offers several datasets to choose from: a small sample of each dataset a published model was trained on, fetched from Hugging Face.

The “Load test data…” example set gives Make Masks ONE dataset – ten Toxoplasma vacuole fields. This is that idea for the rest of the zoo: every published model was trained on a dataset that is already on Hugging Face, and a user who wants to see what a model was taught should be able to open ten of its fields without knowing a repository name.

NO NEW REPOSITORIES, AND NO ARCHIVES. The example set is a purpose-built repo with a single tar. The training datasets are not: they are 556 to 6,062 files of images/ and masks/, and nobody is downloading 3,029 fields to look at ten. Each sample is fetched FILE BY FILE with hf_hub_download, which is why MASK_DATASETS carries the two folder names per repo rather than assuming them – cross-channel-toxoplasma-from-cellmask calls its masks masks_pv, and assuming masks would have given ten images and no labels with no error to explain it.

WHICH FIELDS. A random draw in which every field of the dataset has the same chance – the first N by name are the first plate of the first domain, which is not the dataset. SEEDED by the dataset key, so the draw is the same on every machine: a sample that changes between two people’s machines is a sample nobody can talk about – “the third field looks wrong” has to mean the same field for both of them. A dataset with a quota draws that way within each domain.

THE LAYOUT IS MAKE MASKS’ OWN: images at the top of the folder, masks in masks/ beneath them, which is what curation_queue.detect_layout calls nested and the layout spaCR uses everywhere. So a sample opens for EDITING, with the published masks as the drafts – deliberately unlike the example set, which hides its truth in ground_truth_masks/ so the fields open raw. Those are two different jobs: the example set is “try segmenting this”, and a training-dataset sample is “look at what the model was taught”.

THE WELL-DETECTOR DATASET IS NOT HERE. toxoplasma-plaque-well-detector-dataset has images/ and labels/, and those labels are YOLO bounding boxes. A box is not a mask and Make Masks would open every field blank.

Classes

DatasetPicker

The dataset list: one row per dataset, with what it is.

MaskDataset

One published dataset, and how to take a sample of it.

Functions

choose_sample(→ List[Tuple[str, str]])

Pick which files a sample holds, as (image path, mask path) in the repo.

datasets_for(→ Tuple[MaskDataset, ...])

The sets a given module offers.

examples_root(→ pathlib.Path)

Where spaCR keeps downloaded example data.

foreground_stems(→ Optional[set])

The stems of the fields that have at least one object, or None.

install_dataset_button(screen[, app_key, use])

Build Make Masks' "Training datasets…" button, wired to screen.

is_present(→ bool)

Whether a complete sample is already unpacked in folder.

open_a_training_dataset(→ bool)

Ask which dataset, fetch a sample if it is not cached, and open it.

pool_stems(→ Optional[set])

The stems a sample may be drawn from, when the dataset names a pool.

sample_folder(→ pathlib.Path)

Where dataset's sample is cached.

Module Contents

class spacr.qt.make_masks_datasets.DatasetPicker(parent=None, datasets=MASK_DATASETS)[source]

Bases: PySide6.QtWidgets.QDialog

The dataset list: one row per dataset, with what it is.

A dialog rather than a dropdown on the toolbar, because each row needs two lines – a title a user recognises and the model it trained – and a dropdown gives one.

Build the dialog with one row per dataset, the first selected.

Parameters:
  • parent – the Qt parent.

  • datasets – the datasets to list.

chosen() → MaskDataset | None[source]

The dataset that is selected, or None.

class spacr.qt.make_masks_datasets.MaskDataset[source]

One published dataset, and how to take a sample of it.

Variables:
  • key – the stable name this is stored and tested under.

  • title – what the picker shows.

  • repo – the Hugging Face dataset repository.

  • images – the folder inside it holding the fields.

  • masks – the folder holding their labels. NOT assumed to be “masks”.

  • model – the model this dataset trained, for the line under the title.

  • note – what a reader should know before opening it.

  • apps – which modules offer this set. A dataset belongs to the module whose job it illustrates, and the plaque sets belong to two – Make Masks, where a curator edits the masks, and Plaque Analysis, where the pipeline runs on them.

  • quota – how many fields to take from each domain, as (domain, count), a domain being the <domain>__ prefix of a file name. Empty means SAMPLE_SIZE fields from the whole set.

  • counts – where the repository says which fields have objects: a CSV with an n_objects column keyed by stem or name, or labels/ for a YOLO set, whose empty label files are the images with nothing in them. Read by foreground_stems().

  • foreground_share – the smallest share of the sample that must show objects; FOREGROUND_SHARE unless the set says otherwise.

  • pool – a CSV in spacr/resources/data whose filename column lists the only fields the sample may be drawn from; empty draws from the whole folder. The plaque figures use it for what the repository does not record – which figures are high resolution.

  • revision – added to the cache folder’s name when this set’s draw changes, so a sample drawn the old way is not reopened as the new one.

property size: int[source]

How many fields a complete sample of this dataset holds.

spacr.qt.make_masks_datasets.choose_sample(dataset: MaskDataset, listing: List[str], size: int = SAMPLE_SIZE, *, foreground: set | None = None) → List[Tuple[str, str]][source]

Pick which files a sample holds, as (image path, mask path) in the repo.

A field is only taken when BOTH its image and its mask are in the listing, by stem. Pairing on the stem rather than the full name is what lets a repo store images/x.tif beside masks/x.tif or masks_pv/x.tif and still pair.

A DATASET WITH NO MASKS FOLDER – masks="" – yields images alone, with an empty mask path. That is the plaque FIGURES set: the Plaque Analysis pipeline takes a figure and finds the wells itself, so there is nothing to pair and demanding a pair would return an empty sample and no reason why.

Parameters:
  • dataset – which dataset, for its two folder names.

  • listing – every path in the repository.

  • size – how many pairs to take. Ignored when the dataset has a MaskDataset.quota, which says how many per domain instead.

  • foreground – from foreground_stems(): at least 80% of what is drawn (per domain, with a quota) has objects, at most 20% is empty.

Returns:

up to size pairs drawn by _random_pick(), sorted, the same on every machine.

spacr.qt.make_masks_datasets.datasets_for(app_key: str) → Tuple[MaskDataset, ...][source]

The sets a given module offers.

Parameters:

app_key – the module, e.g. mask or analyze_plaques.

Returns:

its datasets, in registry order.

spacr.qt.make_masks_datasets.examples_root() → pathlib.Path[source]

Where spaCR keeps downloaded example data.

The same ~/.cache/spacr/example_data that the example set unpacks beside, so a user who clears one clears both and there is one place to look.

Returns:

the folder. It is not created here.

spacr.qt.make_masks_datasets.foreground_stems(dataset: MaskDataset, *, download: Callable | None = None, tree: Callable | None = None) → set | None[source]

The stems of the fields that have at least one object, or None.

Parameters:
  • dataset – the dataset, whose MaskDataset.counts says where to look.

  • download – fn(repo, filename) -> local path; defaults to huggingface_hub.hf_hub_download().

  • tree – fn(repo, folder) -> [(path, size)]; defaults to huggingface_hub.HfApi.list_repo_tree().

Returns:

the stems, or None when the dataset says nothing, in which case every field is treated alike.

spacr.qt.make_masks_datasets.install_dataset_button(screen, app_key: str = 'mask', use=None)[source]

Build Make Masks’ “Training datasets…” button, wired to screen.

It sits beside the example set’s “Load test data…” rather than replacing it. The two answer different questions: that one gives raw fields to segment, this one gives fields WITH the masks a published model was trained on.

Parameters:

screen – the Make Masks screen the sample opens in.

Returns:

the button, for the caller to place.

spacr.qt.make_masks_datasets.is_present(folder, expected: int = SAMPLE_SIZE) → bool[source]

Whether a complete sample is already unpacked in folder.

A half-downloaded sample reads as ABSENT, which is the answer that fetches the rest of it rather than opening a folder with four fields in it and saying nothing.

Parameters:
  • folder – the sample folder.

  • expected – how many pairs a complete sample has.

Returns:

whether that many image/mask pairs are there.

spacr.qt.make_masks_datasets.open_a_training_dataset(screen, *, pick=None, fetch=None, root=None, app_key: str = 'mask', use=None) → bool[source]

Ask which dataset, fetch a sample if it is not cached, and open it.

A cached sample opens at once with no request. Otherwise the fetch runs on a worker thread behind the shared progress dialog and the folder opens when it lands.

Parameters:
  • screen – the Make Masks screen to open in.

  • pick – replaces the dialog, for tests. Called as pick(screen) and returns a MaskDataset or None.

  • fetch – replaces the download, for tests. Called as fetch(screen, dataset, folder, on_done).

  • root – where samples are cached; defaults to the example-data folder.

  • app_key – which module is asking, so the picker offers its sets.

  • use – what to DO with the folder, as use(folder) -> bool. Make Masks opens it in the editor; Plaque Analysis points its src at it instead. The default calls screen._open_folder, which is Make Masks’ own.

Returns:

whether a folder was opened synchronously. A download that has to run returns False and opens later, which is what a caller can check.

spacr.qt.make_masks_datasets.pool_stems(dataset: MaskDataset) → set | None[source]

The stems a sample may be drawn from, when the dataset names a pool.

Written by tools/build_plaque_figure_sample_pool.py, which measures what the repository does not record.

Parameters:

dataset – the dataset.

Returns:

the stems, or None when every field may be drawn.

spacr.qt.make_masks_datasets.sample_folder(root, dataset: MaskDataset) → pathlib.Path[source]

Where dataset’s sample is cached.

Parameters:
  • root – the folder example data is kept in.

  • dataset – the dataset.

Returns:

<root>/mask_datasets/<key>_<rule>, the rule naming how the sample was drawn, so a copy drawn under an older rule is never opened as if it were this one.

Nested helpers

foreground_stems.download(repo, filename)

Fetch filename from dataset repo and return its local path.

spacr/qt/make_masks_datasets.py:307

foreground_stems.tree(repo, folder)

List folder of repo as [(path, size)], files only.

spacr/qt/make_masks_datasets.py:295

open_a_training_dataset.done(result, error) → None

Re-enable the button, then open the fetched folder or report why not.

spacr/qt/make_masks_datasets.py:617