spacr.qt.make_masks_datasets¶
A sample of every dataset spaCR has a model for, one dialog away in Make Masks.
Make Masks offers several datasets to choose from: a small sample of each dataset a published model was trained on, fetched from Hugging Face.
The “Load test data…” example set gives Make Masks ONE dataset – ten Toxoplasma vacuole fields. This is that idea for the rest of the zoo: every published model was trained on a dataset that is already on Hugging Face, and a user who wants to see what a model was taught should be able to open ten of its fields without knowing a repository name.
NO NEW REPOSITORIES, AND NO ARCHIVES. The example set is a purpose-built repo with a
single tar. The training datasets are not: they are 556 to 6,062 files of images/ and
masks/, and nobody is downloading 3,029 fields to look at ten. Each sample is fetched
FILE BY FILE with hf_hub_download, which is why MASK_DATASETS carries the two
folder names per repo rather than assuming them – cross-channel-toxoplasma-from-cellmask
calls its masks masks_pv, and assuming masks would have given ten images and no
labels with no error to explain it.
WHICH FIELDS. A random draw in which every field of the dataset has the same chance – the first N by name are the first plate of the first domain, which is not the dataset. SEEDED by the dataset key, so the draw is the same on every machine: a sample that changes between two people’s machines is a sample nobody can talk about – “the third field looks wrong” has to mean the same field for both of them. A dataset with a quota draws that way within each domain.
THE LAYOUT IS MAKE MASKS’ OWN: images at the top of the folder, masks in masks/
beneath them, which is what curation_queue.detect_layout calls nested and the layout
spaCR uses everywhere. So a sample opens for EDITING, with the published masks as the
drafts – deliberately unlike the example set, which hides its truth in
ground_truth_masks/ so the fields open raw. Those are two different jobs: the example
set is “try segmenting this”, and a training-dataset sample is “look at what the model
was taught”.
THE WELL-DETECTOR DATASET IS NOT HERE. toxoplasma-plaque-well-detector-dataset has
images/ and labels/, and those labels are YOLO bounding boxes. A box is not a mask
and Make Masks would open every field blank.
Classes¶
The dataset list: one row per dataset, with what it is. |
|
One published dataset, and how to take a sample of it. |
Functions¶
|
Pick which files a sample holds, as |
|
The sets a given module offers. |
|
Where spaCR keeps downloaded example data. |
|
The stems of the fields that have at least one object, or |
|
Build Make Masks' "Training datasets…" button, wired to |
|
Whether a complete sample is already unpacked in |
|
Ask which dataset, fetch a sample if it is not cached, and open it. |
|
The stems a sample may be drawn from, when the dataset names a pool. |
|
Where |
Module Contents¶
- class spacr.qt.make_masks_datasets.DatasetPicker(parent=None, datasets=MASK_DATASETS)[source]¶
Bases:
PySide6.QtWidgets.QDialogThe dataset list: one row per dataset, with what it is.
A dialog rather than a dropdown on the toolbar, because each row needs two lines – a title a user recognises and the model it trained – and a dropdown gives one.
Build the dialog with one row per dataset, the first selected.
- Parameters:
parent – the Qt parent.
datasets – the datasets to list.
- chosen() MaskDataset | None[source]¶
The dataset that is selected, or None.
- class spacr.qt.make_masks_datasets.MaskDataset[source]¶
One published dataset, and how to take a sample of it.
- Variables:
key – the stable name this is stored and tested under.
title – what the picker shows.
repo – the Hugging Face dataset repository.
images – the folder inside it holding the fields.
masks – the folder holding their labels. NOT assumed to be “masks”.
model – the model this dataset trained, for the line under the title.
note – what a reader should know before opening it.
apps – which modules offer this set. A dataset belongs to the module whose job it illustrates, and the plaque sets belong to two – Make Masks, where a curator edits the masks, and Plaque Analysis, where the pipeline runs on them.
quota – how many fields to take from each domain, as
(domain, count), a domain being the<domain>__prefix of a file name. Empty meansSAMPLE_SIZEfields from the whole set.counts – where the repository says which fields have objects: a CSV with an
n_objectscolumn keyed bystemorname, orlabels/for a YOLO set, whose empty label files are the images with nothing in them. Read byforeground_stems().foreground_share – the smallest share of the sample that must show objects;
FOREGROUND_SHAREunless the set says otherwise.pool – a CSV in
spacr/resources/datawhosefilenamecolumn lists the only fields the sample may be drawn from; empty draws from the whole folder. The plaque figures use it for what the repository does not record – which figures are high resolution.revision – added to the cache folder’s name when this set’s draw changes, so a sample drawn the old way is not reopened as the new one.
- spacr.qt.make_masks_datasets.choose_sample(dataset: MaskDataset, listing: List[str], size: int = SAMPLE_SIZE, *, foreground: set | None = None) List[Tuple[str, str]][source]¶
Pick which files a sample holds, as
(image path, mask path)in the repo.A field is only taken when BOTH its image and its mask are in the listing, by stem. Pairing on the stem rather than the full name is what lets a repo store
images/x.tifbesidemasks/x.tiformasks_pv/x.tifand still pair.A DATASET WITH NO MASKS FOLDER –
masks=""– yields images alone, with an empty mask path. That is the plaque FIGURES set: the Plaque Analysis pipeline takes a figure and finds the wells itself, so there is nothing to pair and demanding a pair would return an empty sample and no reason why.- Parameters:
dataset – which dataset, for its two folder names.
listing – every path in the repository.
size – how many pairs to take. Ignored when the dataset has a
MaskDataset.quota, which says how many per domain instead.foreground – from
foreground_stems(): at least 80% of what is drawn (per domain, with a quota) has objects, at most 20% is empty.
- Returns:
up to
sizepairs drawn by_random_pick(), sorted, the same on every machine.
- spacr.qt.make_masks_datasets.datasets_for(app_key: str) Tuple[MaskDataset, ...][source]¶
The sets a given module offers.
- Parameters:
app_key – the module, e.g.
maskoranalyze_plaques.- Returns:
its datasets, in registry order.
- spacr.qt.make_masks_datasets.examples_root() pathlib.Path[source]¶
Where spaCR keeps downloaded example data.
The same
~/.cache/spacr/example_datathat the example set unpacks beside, so a user who clears one clears both and there is one place to look.- Returns:
the folder. It is not created here.
- spacr.qt.make_masks_datasets.foreground_stems(dataset: MaskDataset, *, download: Callable | None = None, tree: Callable | None = None) set | None[source]¶
The stems of the fields that have at least one object, or
None.- Parameters:
dataset – the dataset, whose
MaskDataset.countssays where to look.download –
fn(repo, filename) -> local path; defaults tohuggingface_hub.hf_hub_download().tree –
fn(repo, folder) -> [(path, size)]; defaults tohuggingface_hub.HfApi.list_repo_tree().
- Returns:
the stems, or
Nonewhen the dataset says nothing, in which case every field is treated alike.
- spacr.qt.make_masks_datasets.install_dataset_button(screen, app_key: str = 'mask', use=None)[source]¶
Build Make Masks’ “Training datasets…” button, wired to
screen.It sits beside the example set’s “Load test data…” rather than replacing it. The two answer different questions: that one gives raw fields to segment, this one gives fields WITH the masks a published model was trained on.
- Parameters:
screen – the Make Masks screen the sample opens in.
- Returns:
the button, for the caller to place.
- spacr.qt.make_masks_datasets.is_present(folder, expected: int = SAMPLE_SIZE) bool[source]¶
Whether a complete sample is already unpacked in
folder.A half-downloaded sample reads as ABSENT, which is the answer that fetches the rest of it rather than opening a folder with four fields in it and saying nothing.
- Parameters:
folder – the sample folder.
expected – how many pairs a complete sample has.
- Returns:
whether that many image/mask pairs are there.
- spacr.qt.make_masks_datasets.open_a_training_dataset(screen, *, pick=None, fetch=None, root=None, app_key: str = 'mask', use=None) bool[source]¶
Ask which dataset, fetch a sample if it is not cached, and open it.
A cached sample opens at once with no request. Otherwise the fetch runs on a worker thread behind the shared progress dialog and the folder opens when it lands.
- Parameters:
screen – the Make Masks screen to open in.
pick – replaces the dialog, for tests. Called as
pick(screen)and returns aMaskDatasetor None.fetch – replaces the download, for tests. Called as
fetch(screen, dataset, folder, on_done).root – where samples are cached; defaults to the example-data folder.
app_key – which module is asking, so the picker offers its sets.
use – what to DO with the folder, as
use(folder) -> bool. Make Masks opens it in the editor; Plaque Analysis points itssrcat it instead. The default callsscreen._open_folder, which is Make Masks’ own.
- Returns:
whether a folder was opened synchronously. A download that has to run returns False and opens later, which is what a caller can check.
- spacr.qt.make_masks_datasets.pool_stems(dataset: MaskDataset) set | None[source]¶
The stems a sample may be drawn from, when the dataset names a pool.
Written by
tools/build_plaque_figure_sample_pool.py, which measures what the repository does not record.- Parameters:
dataset – the dataset.
- Returns:
the stems, or None when every field may be drawn.
- spacr.qt.make_masks_datasets.sample_folder(root, dataset: MaskDataset) pathlib.Path[source]¶
Where
dataset’s sample is cached.- Parameters:
root – the folder example data is kept in.
dataset – the dataset.
- Returns:
<root>/mask_datasets/<key>_<rule>, the rule naming how the sample was drawn, so a copy drawn under an older rule is never opened as if it were this one.
Nested helpers¶
- foreground_stems.download(repo, filename)¶
Fetch
filenamefrom datasetrepoand return its local path.spacr/qt/make_masks_datasets.py:307
- foreground_stems.tree(repo, folder)¶
List
folderofrepoas[(path, size)], files only.spacr/qt/make_masks_datasets.py:295