spacr.curation_queue

A curation session is a queue with a memory, not one field.

Make Masks is the better editor and the bespoke tool outside the repository is the better workflow, and neither is both. This module is the workflow half, pulled inside spaCR: which fields are waiting, which order to offer them in, how many to offer before the session ends, and what was already decided about each one.

It holds no pixels and draws nothing. That is deliberate and it is the whole argument for bringing the work in here: the external tool’s defects are found in the middle of a curation session, because there is no way to test a 2,000-line Qt application, whereas everything below is ordinary Python over a folder and a CSV, so a defect is found by tests/test_a_curation_session_resumes_where_it_stopped.py instead. The module imports no Qt, needs no display, and touches numpy only inside the functions that actually read a draft mask off the disk.

Three layouts, all accepted

The layout mismatch is the thing to settle first: silently assuming one layout is how a folder full of work gets loaded as empty. So all three are accepted and the one in front of us is detected:

nested    <folder>/*.tif           + <folder>/masks/*.tif
sibling   <folder>/images/*.tif    + <folder>/masks/*.tif
seg       <folder>/*_seg.npy       (Cellpose pickles, the external tool)

Accepting all three is safer than picking a winner. An existing curation set opens unconverted, so nothing has to be moved before work can start – and the EDITOR takes all three too, reading and saving each where it lies (spacr.qt.mask_engine says how).

A guess is not accepted. A folder that matches none of the three, or that matches two of them, raises LayoutError. The alternative would be an empty queue, and an empty queue reads as “all done”.

Three reviewed states, not two

REVIEWED is done, skip and recropped. The ledger note says two; it is wrong. skip is not a lesser done — it records “this one cannot be curated”, and collapsing the two means the same unusable field is offered again every session. recropped is what the external tool writes when a bundle has been split into children: the parent is out of the queue and its pieces are in it.

The record lives at <folder>/curate_status.csv, inside the queue folder rather than beside it, so it syncs with the images and a session resumes on the other machine. A row about a stem that is not in this folder is kept on rewrite for exactly that reason: the curator may be resuming somewhere the rest of the set has not arrived yet, and dropping those rows would hand back work already done. A folder with no record of its own but with the external tool’s <parent>/<name>_status.csv beside it resumes from that one and writes back to it (status_path()), so a set curated half in each tool keeps one record.

Ordering

Ordering is the productivity feature: a curator’s judgement is worth more on the hard fields than on the first hundred easy ones. The four orders, and the constants inside them, are the external tool’s, chosen after a real 500-bundle curation set. order_items() says what each one means.

Where this knowingly differs from the external tool

Five places, each because this module must serve three layouts and a test suite rather than one folder of Cellpose pickles:

  • Objects are counted as distinct non-zero labels, not masks.max(). See count_draft_objects() for the two cases where those differ.

  • The median object diameter is computed directly rather than through skimage.measure.regionprops, so ordering a queue needs numpy alone. It is the same number, and a test pins the equality.

  • An unknown order raises instead of silently becoming value. A silent reorder is the failure this module is most careful about.

  • name order sorts by stem rather than by file name, so the order does not change when one field is a PNG and the next a TIFF.

  • Status rows are written sorted by stem, atomically. The file syncs between machines; a stable order keeps the diff to the rows that changed, and the rename keeps a half-written resume record off the disk.

The nested and sibling layouts also offer an image whose mask is missing, rather than pairing strictly as the external tool’s bundle builder does: a field with no draft is work, drawing from scratch, and easy already knows to offer it last.

Exceptions

CurationQueueError

A curation queue could not be built, read or written.

LayoutError

The folder does not unambiguously hold one of the three layouts.

StatusFileError

curate_status.csv exists but cannot be read as a resume record.

Classes

CurationQueue

A bounded session of curation work over one folder.

QueueItem

One field waiting to be curated.

QueueLayout

What was found in the folder, and where.

QueueSummary

What is left to do in a curation session.

StatusRow

One line of curate_status.csv.

Functions

build_queue(→ CurationQueue)

Open a folder as a bounded curation session.

count_draft_objects(→ int)

How many objects the draft holds.

detect_layout(→ QueueLayout)

Work out which of the three layouts folder holds, and read it.

discover_items(→ Tuple[QueueItem, ...])

Every field in folder, whatever layout it is in.

draft_stats(→ Tuple[int, float])

external_record(→ pathlib.Path)

Where the external curation tool keeps one of a queue's files.

is_reviewed(→ bool)

load_draft_counts(→ Dict[str, int])

Draft object counts for items, cached in curate_drafts.csv.

load_probabilities(→ Dict[str, float])

Read the queue's scores file, when there is one.

mark_state(→ Dict[str, StatusRow])

Record one decision and write the file, keeping every other row.

median_equivalent_diameter(→ float)

Median circle-equivalent diameter of the objects in a label image.

order_items(→ List[QueueItem])

Order a queue, with the external tool's semantics.

pending_items(→ List[QueueItem])

read_status(→ Dict[str, StatusRow])

Read the resume record.

record_state(→ Dict[str, StatusRow])

Return rows with one stem's decision recorded.

resolve_order(→ Tuple[str, Optional[str]])

Decide which ordering will actually be used, and what to say about it.

scores_path(→ Optional[pathlib.Path])

Where this queue's probabilities are read from, if anywhere.

status_path(→ pathlib.Path)

Where this queue's resume record is read from and written to.

summarize(→ QueueSummary)

Count the folder the way the ledger asks for it.

value_key(→ Tuple[int, float, int])

Sort key for value order: how much draft there is to correct.

write_status(→ pathlib.Path)

Write the resume record, atomically and in full.

Module Contents

exception spacr.curation_queue.CurationQueueError[source]

Bases: spacr.errors.SpacrError

A curation queue could not be built, read or written.

Subclass of spacr.errors.SpacrError, so a caller that catches spaCR’s own errors catches these too.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.curation_queue.LayoutError[source]

Bases: CurationQueueError

The folder does not unambiguously hold one of the three layouts.

Raised rather than returning an empty queue. An empty queue reads as “all done”, which is the one wrong answer a curation tool must never give about a folder full of work.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.curation_queue.StatusFileError[source]

Bases: CurationQueueError

curate_status.csv exists but cannot be read as a resume record.

Initialize self. See help(type(self)) for accurate signature.

class spacr.curation_queue.CurationQueue[source]

A bounded session of curation work over one folder.

Variables:
  • layout – what was found, and where.

  • items – the fields to offer THIS session — unreviewed, ordered, and cut to limit.

  • status – every status row read, including rows about stems that are not in this folder.

  • order – the ordering that was asked for.

  • effective_order – the ordering actually used, which differs from order when prob or easy fell back to value for want of probabilities.

  • limit – the cap that was applied, or None.

  • summary – counts over the WHOLE folder, not over this session.

  • notices – what the curator should know before the first field – a prob or easy order with no scores file, or with fields the scores file does not name, and a resume record that is the external tool’s file beside the folder, which every save then writes to. Printed when the queue is built and shown by the editor with it, because an ordering that quietly became another one, or a save that quietly lands outside the folder that was opened, is the failure this module is most careful about.

describe() → str[source]
Returns:

a printable line naming the layout, order and counts.

property all_items: Tuple[QueueItem, ...][source]

every field in the folder, reviewed ones included.

Type:

returns

property folder: pathlib.Path[source]

the queue folder.

Type:

returns

property order_phrase: str[source]

sorted by <order>, naming the order that was asked for too whenever it is not the one that was used.

Type:

returns

class spacr.curation_queue.QueueItem[source]

One field waiting to be curated.

Variables:
  • stem – the identity of the field, and the key of its status row.

  • layout – which of LAYOUTS this item came from.

  • image – the image file, when the layout has one on disk. None for a seg bundle that carries its own image, unless a display image of the same stem sits beside it.

  • mask – the draft mask file, or None when no draft exists yet.

  • bundle – the _seg.npy pickle, for the seg layout only.

property draft: pathlib.Path | None[source]

The file holding this item’s draft labels, or None.

Returns:

the bundle for a seg item, the mask file for the other two layouts, and None when there is no draft — which means drawing from scratch, not an error.

class spacr.curation_queue.QueueLayout[source]

What was found in the folder, and where.

Variables:
  • kind – one of LAYOUTS.

  • folder – the queue folder, which is also where the status file goes.

  • images_dir – where the images are, when they are on disk.

  • masks_dir – where masks are read from and written back to.

  • items – every field found, sorted by stem.

mask_destination(stem: str) → pathlib.Path[source]

Where a curated mask for stem is written.

The seg layout writes back into the bundle it was read from, as the external tool does; the other two write <masks_dir>/<stem>.tif, which is what spacr.qt.mask_engine.mask_save_path() already produces for the nested layout.

Parameters:

stem – the field’s stem.

Returns:

the path a curated mask belongs at.

Raises:

CurationQueueError – when stem is not in this queue.

property status_path: pathlib.Path[source]

where this queue’s resume record is read and written: <folder>/curate_status.csv, or the external curation tool’s <parent>/<name>_status.csv beside the folder when that is the record the folder has. status_path() decides.

Type:

returns

class spacr.curation_queue.QueueSummary[source]

What is left to do in a curation session.

Variables:
  • total – fields in the folder.

  • done – curated.

  • skip – refused — cannot be curated, and must not be offered again.

  • recropped – split into children, which are in the queue instead.

  • remaining – fields with no reviewed state.

  • unknown – status rows about stems that are not in this folder. Kept, never dropped, and reported here rather than silently: on a second machine they are usually work already done elsewhere.

describe() → str[source]

One line a caller can print.

The ledger’s measured state, 500 bundles, 366 done, 29 skip, 105 remaining, is exactly what this returns for it. The two rarer counts are appended only when they are not zero, so the common sentence stays the short one.

Returns:

a human-readable summary of the queue.

class spacr.curation_queue.StatusRow[source]

One line of curate_status.csv.

Variables:
  • stem – which field this is about.

  • state – done, skip, recropped, or anything else, which counts as pending. Normalised to lower case with the surrounding whitespace stripped, so a row typed by hand as Done still ends that stem’s turn.

  • n_objects – how many objects the field had when it was reviewed, or None when the file did not say.

  • updated – ISO-8601 timestamp, as written.

property reviewed: bool[source]

whether this row takes its stem out of the queue.

Type:

returns

spacr.curation_queue.build_queue(folder: PathLike, order: str = DEFAULT_ORDER, limit: int | None = None, probs=None, counts=None, announce: Callable[[str], None] | None = None, include_reviewed: bool = False, min_objects: int = MIN_OBJECTS_FOR_VALUE, min_diameter: float = MIN_DIAMETER_FOR_VALUE, cache_counts: bool = True) → CurationQueue[source]

Open a folder as a bounded curation session.

In order: detect the layout, read the resume record, drop everything already reviewed, order what is left, and only THEN apply the limit — so --limit 20 means the twenty most worthwhile fields still to do, not the first twenty found.

Probabilities are read only for the orders that use them, so name and value read no scores file at all. For prob and easy the absence of scores is never silent: no scores file falls back to value and says so, naming both places it looked; a scores file that leaves some waiting fields unscored says how many; and where the scores came from is printed. Those warnings, and the one that says progress is being written back to the external tool’s record beside the folder, are kept on the queue as CurationQueue.notices for the editor to show.

Parameters:
  • folder – the queue folder.

  • order – one of ORDERS; DEFAULT_ORDER by default.

  • limit – how many fields this session offers, or None for all.

  • probs – probabilities to use instead of reading curate_scores.csv.

  • counts – draft counts to use instead of reading the drafts.

  • announce – where notices go; defaults to print().

  • include_reviewed – offer reviewed fields too. Off by default; this is what “resume” means.

  • min_objects – passed to value_key().

  • min_diameter – passed to value_key().

  • cache_counts – whether easy may write its count cache.

Returns:

the CurationQueue.

Raises:
spacr.curation_queue.count_draft_objects(item: QueueItem) → int[source]

How many objects the draft holds.

Counted as DISTINCT NON-ZERO LABELS rather than as masks.max(), which is what the external tool uses. The two agree on a fresh Cellpose pickle, whose ids run 1..N, and disagree in the two places that matter here: a label TIFF that has had objects deleted has gaps in its ids, so max() overcounts; and a 0/255 binary mask, which the nested and sibling layouts can perfectly well hold, would report 255 objects.

Parameters:

item – the field to count.

Returns:

the number of objects, and 0 for a draft that is missing or unreadable — as the external tool’s count cache also does, since counts drive ordering only.

spacr.curation_queue.detect_layout(folder: PathLike) → QueueLayout[source]

Work out which of the three layouts folder holds, and read it.

Parameters:

folder – the queue folder.

Returns:

the QueueLayout, with every field found in it.

Raises:
  • LayoutError – when folder is not a directory, matches none of the three layouts, or matches more than one of them.

  • CurationQueueError – when two files in it share a stem.

Both refusals are loud on purpose. Guessing between two layouts loads half a folder; assuming one that is not there loads none of it, and reports that as a finished queue.

spacr.curation_queue.discover_items(folder: PathLike) → Tuple[QueueItem, ...][source]

Every field in folder, whatever layout it is in.

Parameters:

folder – the queue folder.

Returns:

the items, sorted by stem.

Raises:

LayoutError – as detect_layout() does.

spacr.curation_queue.draft_stats(item: QueueItem) → Tuple[int, float][source]
Parameters:

item – the field to measure.

Returns:

(n_objects, median_diameter) for its draft.

Raises:

Exception – whatever reading an unreadable draft raises. The value order catches that and sorts the field dead last.

spacr.curation_queue.external_record(folder: PathLike, suffix: str) → pathlib.Path[source]

Where the external curation tool keeps one of a queue’s files.

Parameters:
  • folder – the queue folder.

  • suffix – EXTERNAL_STATUS_SUFFIX or EXTERNAL_SCORES_SUFFIX.

Returns:

<parent>/<name><suffix>, beside the queue folder rather than in it. The folder is made absolute first, without touching the disk, so . has a name to put in front of the suffix.

spacr.curation_queue.is_reviewed(state: str | None) → bool[source]
Parameters:

state – a state string, or None.

Returns:

whether that state takes a stem out of the queue. done, skip and recropped all do, and they stay distinct: skip means “cannot be curated”, which is not a lesser done.

spacr.curation_queue.load_draft_counts(folder: PathLike, items: Sequence[QueueItem], write_cache: bool = True, announce: Callable[[str], None] | None = None) → Dict[str, int][source]

Draft object counts for items, cached in curate_drafts.csv.

Reading three hundred drafts at every launch is slow, so counts are cached by stem and only stems missing from the cache are read. Counts drive ORDERING only: a stale one mis-sorts a field, it never loses work.

SO AN UNREADABLE CACHE IS RECOUNTED, NEVER RAISED. A file truncated by a crash mid-write is not valid UTF-8, and letting that out would take the whole session down over an optimisation – the drafts are on disk and are the truth.

Parameters:
  • folder – the queue folder.

  • items – the fields to count.

  • write_cache – whether to write the cache back. A failure to write is announced, never raised — a read-only folder must still open.

  • announce – where progress goes; defaults to print().

Returns:

{stem: n_objects} for every item given.

spacr.curation_queue.load_probabilities(folder: PathLike) → Dict[str, float][source]

Read the queue’s scores file, when there is one.

Parameters:

folder – the queue folder.

Returns:

{stem: probability}, and {} when there is no scores file (see scores_path()) – which is what makes prob and easy fall back, loudly.

Raises:

CurationQueueError – when curate_scores.csv is present but has no stem column or no recognised probability column. A scores file that exists and cannot be used is a mistake worth stopping for; a missing one is an ordinary state.

spacr.curation_queue.mark_state(folder: PathLike, stem: str, state: str, n_objects: int | None = None, updated: str | None = None) → Dict[str, StatusRow][source]

Record one decision and write the file, keeping every other row.

This is the one call a curation session makes per field, and it is the call where rows about absent stems would otherwise be lost.

Parameters:
  • folder – the queue folder.

  • stem – the field being decided.

  • state – the state to record.

  • n_objects – how many objects it had, if known.

  • updated – the timestamp to record; defaults to now.

Returns:

the rows as they now stand on disk.

Raises:

StatusFileError – when the file cannot be read or written.

spacr.curation_queue.median_equivalent_diameter(labels) → float[source]

Median circle-equivalent diameter of the objects in a label image.

This is skimage.measure.regionprops’s equivalent_diameter_area for 2-D input, computed directly as sqrt(4 * area / pi) so that ordering a queue needs numpy and not scikit-image. The equality is pinned by a test rather than asserted here.

Parameters:

labels – a 2-D integer label image.

Returns:

the median diameter in pixels, or 0.0 when there are no objects.

spacr.curation_queue.order_items(items: Iterable[QueueItem], order: str = DEFAULT_ORDER, probs=None, counts=None, announce: Callable[[str], None] | None = None, min_objects: int = MIN_OBJECTS_FOR_VALUE, min_diameter: float = MIN_DIAMETER_FOR_VALUE, scores_hint: PathLike | None = None) → List[QueueItem][source]

Order a queue, with the external tool’s semantics.

name

By stem, and nothing else is read — not the probabilities, not the drafts. By stem rather than by file name so that the order does not change when one field is a PNG and the next a TIFF.

prob

Most likely first: the key is -probs.get(stem, -1.0). The default is the point. An unscored field yields +1.0, larger than any real negated probability, so it sorts LAST rather than pretending to be probability zero.

easy

(0 if n > 0 else 1, -probs.get(stem, -1.0), -n). A populated draft is accept-or-light-edit; an EMPTY draft means drawing from scratch, so empty ones go last REGARDLESS of probability. Within each group, most likely first, then most objects.

value

By value_key(): how much draft there is to correct.

uncertain

Most uncertain segmentation first, the key being -probs.get(stem, -1.0) with probs carrying each field’s uncertainty from curate_uncertainty.csv. An unscored field sorts last.

Ties are broken by stem, because the list is sorted by stem before any of this and Python’s sort is stable.

Parameters:
  • items – the fields to order.

  • order – one of ORDERS.

  • probs – {stem: probability}, or a callable returning one, or None for none available; for uncertain, {stem: uncertainty}.

  • counts – {stem: n_objects}, or a callable returning one. When None, easy reads the drafts itself.

  • announce – where the fallback notice goes; defaults to print(), so it lands on stdout as the external tool’s does.

  • min_objects – passed to value_key().

  • min_diameter – passed to value_key().

  • scores_hint – the scores file named in the fallback notice.

Returns:

a new list, ordered.

Raises:

ValueError – when order is not one of ORDERS.

spacr.curation_queue.pending_items(items: Iterable[QueueItem], status: Mapping[str, StatusRow]) → List[QueueItem][source]
Parameters:
  • items – fields in the folder.

  • status – the rows read from the status file.

Returns:

the items with no reviewed state, in the order given.

spacr.curation_queue.read_status(folder: PathLike) → Dict[str, StatusRow][source]

Read the resume record.

Parameters:

folder – the queue folder.

Returns:

{stem: StatusRow}, empty when no status file exists yet. Rows about stems that are not in the folder are included; they are the reason write_status() writes back everything it is given.

Raises:

StatusFileError – when the file exists but has no stem column, which means it is not a status file at all.

spacr.curation_queue.record_state(rows: Mapping[str, StatusRow], stem: str, state: str, n_objects: int | None = None, updated: str | None = None) → Dict[str, StatusRow][source]

Return rows with one stem’s decision recorded.

Pure: the mapping passed in is not modified, so a caller can decide whether the change reaches the disk.

Parameters:
  • rows – the rows read so far.

  • stem – the field being decided.

  • state – done, skip, recropped, or any other string, which leaves the stem pending.

  • n_objects – how many objects it had, if known.

  • updated – the timestamp to record; defaults to now, to the second.

Returns:

a new mapping including the recorded row.

spacr.curation_queue.resolve_order(order: str, probabilities: Mapping[str, float], scores_hint: PathLike | None = None) → Tuple[str, str | None][source]

Decide which ordering will actually be used, and what to say about it.

prob and easy both need probabilities. Without them they fall back to value — AND SAY SO. The announcement is half the behaviour: a silent reorder leaves the curator believing they are working the uncertain fields first when they are not.

Parameters:
  • order – the ordering asked for, one of ORDERS.

  • probabilities – the probabilities available, possibly empty.

  • scores_hint – the scores file that would have supplied them, named in the message so the fix is obvious.

Returns:

(effective_order, announcement), the announcement being None when nothing changed.

Raises:

ValueError – when order is not one of ORDERS. A typo must not quietly become a different ordering.

spacr.curation_queue.scores_path(folder: PathLike) → pathlib.Path | None[source]

Where this queue’s probabilities are read from, if anywhere.

<folder>/curate_scores.csv first; failing that, the external curation tool’s <parent>/<name>_scores.csv when its header has a stem column and a probability column, which is how the plaque project’s scores are written. A file beside the folder that does not look like scores is not a scores file and is passed over rather than allowed to stop the session.

Parameters:

folder – the queue folder.

Returns:

the scores file, or None when there is none.

spacr.curation_queue.status_path(folder: PathLike) → pathlib.Path[source]

Where this queue’s resume record is read from and written to.

<folder>/curate_status.csv, inside the queue folder so it travels with the images when the set is synced between machines – UNLESS the folder has none and the external curation tool’s record, <parent>/<name>_status.csv, sits beside it with stem and state columns. That record is then THE record, read and written back in place: the sets that tool curated keep their progress there and nothing inside the folder, and starting a second file would reopen every field already done or skipped as undone while the other tool went on reading the first.

Parameters:

folder – the queue folder.

Returns:

the path of the resume record, which need not exist yet.

spacr.curation_queue.summarize(items: Iterable[QueueItem], status: Mapping[str, StatusRow]) → QueueSummary[source]

Count the folder the way the ledger asks for it.

Parameters:
  • items – every field in the folder.

  • status – the rows read from the status file.

Returns:

the QueueSummary. Counts are over fields that are actually here, so done + skip + recropped + remaining == total; rows about stems that are not here are reported separately as unknown rather than subtracted from what is left to do.

spacr.curation_queue.value_key(item: QueueItem, min_objects: int = MIN_OBJECTS_FOR_VALUE, min_diameter: float = MIN_DIAMETER_FOR_VALUE) → Tuple[int, float, int][source]

Sort key for value order: how much draft there is to correct.

Exactly the external tool’s _value:

n == 0                     -> (2, 0, 0)   nothing to correct, last
unreadable                 -> (3, 0, 0)   dead last
otherwise                  -> (0 if rich else 1, -diameter, -n)
Parameters:
  • item – the field to rank.

  • min_objects – objects a “rich” draft needs. Plaque-specific.

  • min_diameter – median diameter a “rich” draft needs, in pixels. Plaque-specific.

Returns:

the sort key.

spacr.curation_queue.write_status(folder: PathLike, rows: Mapping[str, StatusRow]) → pathlib.Path[source]

Write the resume record, atomically and in full.

Every row given is written, including rows about stems that are not in this folder: the curator may be resuming on a machine where part of the set has not arrived, and dropping those rows offers work back that was already finished somewhere else.

Rows are written sorted by stem. The file syncs between machines, and a stable order keeps a sync diff to the rows that actually changed.

Parameters:
  • folder – the queue folder.

  • rows – {stem: StatusRow}, as read_status() returns.

Returns:

the path written.

Raises:

StatusFileError – when the file cannot be written.

Nested helpers

order_items.easy_key(item: QueueItem) → Tuple[int, float, int]
Parameters:

item – one field.

Returns:

its easy sort key, populated drafts first.

spacr/curation_queue.py:1349