spacr.screen_data

The published TSG101 screen, as separately downloadable pieces.

WHY PIECES. The screen is four plates of about 8 GB of crops each plus a half-gigabyte database each – 33 GB in total. Almost nobody wants all of it: the regression measurement and cell functions read the DATABASES, and the crops are only needed when something has to display an image. Shipping it as one download would make trying one function cost 33 GB.

So each plate’s database and each plate’s crop folder is its own archive, and ScreenAsset says how big each one is BEFORE it is fetched. The picker in the Regression screen lists them with their sizes and downloads only what is selected.

The merged/ folders are deliberately absent. They are about 300 GB per plate – 1.2 TB for the screen – which is past what a public dataset host will take without arrangement, and nothing in Regression reads them.

Classes

ScreenAsset

One separately downloadable piece of the screen.

Functions

assets_for(→ List[ScreenAsset])

The pieces matching kind and plate.

human_size(→ str)

1234567 as 1.2 MB.

published_archives([repo, timeout])

The archive names actually present in repo, or None.

total_size(→ int)

Return how many bytes a selection will cost.

Module Contents

class spacr.screen_data.ScreenAsset[source]

One separately downloadable piece of the screen.

Parameters:
  • archive – the .tar in SCREEN_REPO.

  • plate – which plate it belongs to.

  • kind – "measurements" or "crops".

  • bytes – the archive’s size, so the picker can say what a tick costs before it is paid.

  • unpacks_to – the path inside the plate folder the archive fills, used to tell an already-downloaded piece from one that is missing.

is_present(folder) → bool[source]

Whether this piece is already unpacked under folder.

Checks what the archive WRITES, not the folder it writes into: every piece shares one plate directory, so the directory existing says nothing about which pieces are in it.

Parameters:

folder – plate directory under which this asset would unpack.

Returns:

whether the database file exists, or the crop directory contains at least one PNG, according to this asset’s kind.

property label: str[source]

Return the plate-and-content label shown by the data picker.

Returns:

human-readable plate and archive-content label.

spacr.screen_data.assets_for(kind: str | None = None, plate: int | None = None) → List[ScreenAsset][source]

The pieces matching kind and plate.

Parameters:
  • kind – "measurements", "crops", or None for both.

  • plate – a plate number, or None for all of them.

Returns:

assets matching both supplied filters, in picker order.

spacr.screen_data.human_size(count: int) → str[source]

1234567 as 1.2 MB.

Decimal units, matching what a download manager and a disk vendor both report – a user comparing this figure with either should not have to know which of two conventions each of us picked.

Parameters:

count – byte count to format.

Returns:

decimal-unit size from bytes through terabytes, with one decimal above the byte unit.

spacr.screen_data.published_archives(repo: str = SCREEN_REPO, *, timeout: float = 8.0)[source]

The archive names actually present in repo, or None.

None means “could not tell” – offline, or the hub did not answer – and is DELIBERATELY different from an empty set. A caller that treated a failed lookup as “nothing is published” would grey out every row and leave the user with a picker that offers nothing and explains nothing.

Parameters:
  • repo – Hugging Face dataset repository to inspect.

  • timeout – give up rather than hold a dialog open on a slow network.

Returns:

names of published .tar archives, or None when the repository could not be inspected.

spacr.screen_data.total_size(assets) → int[source]

Return how many bytes a selection will cost.

Parameters:

assets – iterable of ScreenAsset objects.

Returns:

sum of the assets’ published archive sizes in bytes.