spacr.annotation_dataset

Build an annotation set by streaming crops, and register it for annotating.

WHAT THIS IS FOR. Annotating means looking at single-object crops, and until now the only way to get a set of them was to run Measure over a whole plate – about twenty minutes for 52 fields on a 30-core machine. Deciding afterwards that the crops should have been cut differently meant running it again.

Every field’s objects are already described twice over once Measure has run: by the label masks inside the merged arrays, and by the coordinate columns in measurements.db. Either is enough to cut crops from, so a second set can be built in seconds without measuring anything again.

THE TWO ROUTES ARE NOT EQUIVALENT, and the difference is not guessable:

  • array reads the object masks out of the merged stacks, so it can cut to the object itself – masked, or to its bounding box;

  • database reads the coordinate columns, which are all the database stores, so it can only ever produce a BOUNDING BOX.

Comparing the two is therefore only meaningful with bounding_box=True. spacr.annotation_dataset() says so in the settings it accepts, and the GUI says it beside the picker.

Functions

crops_folder_for(→ str)

The crop folder that belongs to table.

filter_selection(→ pandas.DataFrame)

Drop objects a run would not have cropped.

generate_annotation_dataset(→ Dict[str, Any])

Stream a set of crops and register it for annotation.

next_png_table(→ str)

The name a new annotation set should be written under.

png_list_frame(→ pandas.DataFrame)

The rows the annotation viewer reads, in measure_crop's own schema.

read_objects_from_database(→ Optional[pandas.DataFrame])

The object table a streamed set can be built from.

reserve_png_table(→ str)

Claim the next free table name by creating it, empty.

write_png_list(→ str)

Write an annotation set into the measurements database.

Module Contents

spacr.annotation_dataset.crops_folder_for(table: str) → str[source]

The crop folder that belongs to table.

png_list -> data; png_list_2 -> data_2. The suffix is carried across rather than counted again, so the pair cannot drift.

spacr.annotation_dataset.filter_selection(selection: pandas.DataFrame, settings: Mapping[str, Any]) → pandas.DataFrame[source]

Drop objects a run would not have cropped.

THE SAME PREDICATES measure_crop APPLIES, so a streamed set and a measured one describe the same population. They are expressed against the selection table’s columns rather than against a measurement frame, because the array route has no measurements – only labels and geometry.

Recognised settings, each optional and each skipped when absent:

{object}_min_size / {object}_max_size

Object area in pixels. 0 and None both mean “no bound”, which is what the Measure panel writes for an unset field – treating 0 as a real minimum would drop nothing and look like it had worked.

wells / exclude_wells

Well ids to keep or drop, as rowID+columnID pairs or as plain well names.

max_objects

A cap, applied LAST and deterministically (by the sort the selection already has), so a capped set is reproducible.

Parameters:
Returns:

a new frame; the input is not modified.

spacr.annotation_dataset.generate_annotation_dataset(settings: Mapping[str, Any]) → Dict[str, Any][source]

Stream a set of crops and register it for annotation.

Parameters:

settings – needs src (the plate folder) and accepts stream_source (STREAM_SOURCES), object_array, channel_arrays, bounding_box, the filtration keys filter_selection() reads, and dst for where the crops go.

Returns:

the streaming report, plus table naming what was written.

spacr.annotation_dataset.next_png_table(connection: sqlite3.Connection) → str[source]

The name a new annotation set should be written under.

png_list when it is free, then png_list_2, png_list_3 and so on. An existing set is never overwritten: it may already carry annotations, and those are hand-made and unrecoverable.

THE CALLER MUST HOLD A WRITE TRANSACTION. Choosing a name and creating the table are two steps, and two runs started together would otherwise choose the same one. write_png_list() does this correctly; call it rather than this.

Parameters:

connection – an open connection to the measurements database.

Returns:

the free table name.

spacr.annotation_dataset.png_list_frame(selection: pandas.DataFrame, paths: Sequence[str]) → pandas.DataFrame[source]

The rows the annotation viewer reads, in measure_crop’s own schema.

Parameters:
  • selection – the filtered selection table.

  • paths – the written crop path for each of its rows, in order.

Returns:

a frame with exactly PNG_LIST_COLUMNS.

spacr.annotation_dataset.read_objects_from_database(db_path: str, object_type: str) → pandas.DataFrame | None[source]

The object table a streamed set can be built from.

Parameters:
  • db_path – the measurements database.

  • object_type – cell, nucleus, pathogen …

Returns:

the rows, or None when the table is not there.

spacr.annotation_dataset.reserve_png_table(db_path: str) → str[source]

Claim the next free table name by creating it, empty.

RESERVED BEFORE THE CROPS ARE CUT, not after, so the folder they are written into can be named to match: png_list_2 gets data_2. Deriving the folder from the table is what makes a set on disk traceable to the set in the database – with two independent counters they drift the first time either is deleted, and then nothing says which folder a table describes.

Creating the table is what reserves it: two runs started together would otherwise choose the same name, and the second would fail on insert after it had already written a folder full of crops.

Parameters:

db_path – the measurements database.

Returns:

the reserved table name.

spacr.annotation_dataset.write_png_list(db_path: str, frame: pandas.DataFrame, *, table: str | None = None) → str[source]

Write an annotation set into the measurements database.

The name is chosen and the table created inside ONE transaction, so two runs started together cannot pick the same one. A caller that already reserved a name with reserve_png_table() passes it as table.

Parameters:
  • db_path – the measurements database.

  • frame – rows as png_list_frame() builds them.

  • table – a name already reserved.

Returns:

the table actually written.

Nested helpers

generate_annotation_dataset._write(path, array)

Write and register one streamed crop through Measure’s writer.

Parameters:
  • path – proposed crop path; its extension is replaced with .png.

  • array – crop array to normalize, select, pad, resize, and save.

Returns:

None. The shared writer’s returned path is appended to the captured list using the captured channels and PNG size, preserving alignment with the subsequently registered selection rows.

spacr/annotation_dataset.py:341