spacr.annotation_dataset¶
Build an annotation set by streaming crops, and register it for annotating.
WHAT THIS IS FOR. Annotating means looking at single-object crops, and until now the only way to get a set of them was to run Measure over a whole plate – about twenty minutes for 52 fields on a 30-core machine. Deciding afterwards that the crops should have been cut differently meant running it again.
Every field’s objects are already described twice over once Measure has run:
by the label masks inside the merged arrays, and by the coordinate columns in
measurements.db. Either is enough to cut crops from, so a second set can be
built in seconds without measuring anything again.
THE TWO ROUTES ARE NOT EQUIVALENT, and the difference is not guessable:
arrayreads the object masks out of the merged stacks, so it can cut to the object itself – masked, or to its bounding box;databasereads the coordinate columns, which are all the database stores, so it can only ever produce a BOUNDING BOX.
Comparing the two is therefore only meaningful with bounding_box=True.
spacr.annotation_dataset() says so in the settings it accepts, and the
GUI says it beside the picker.
Functions¶
|
The crop folder that belongs to |
|
Drop objects a run would not have cropped. |
|
Stream a set of crops and register it for annotation. |
|
The name a new annotation set should be written under. |
|
The rows the annotation viewer reads, in |
|
The object table a streamed set can be built from. |
|
Claim the next free table name by creating it, empty. |
|
Write an annotation set into the measurements database. |
Module Contents¶
- spacr.annotation_dataset.crops_folder_for(table: str) str[source]¶
The crop folder that belongs to
table.png_list->data;png_list_2->data_2. The suffix is carried across rather than counted again, so the pair cannot drift.
- spacr.annotation_dataset.filter_selection(selection: pandas.DataFrame, settings: Mapping[str, Any]) pandas.DataFrame[source]¶
Drop objects a run would not have cropped.
THE SAME PREDICATES
measure_cropAPPLIES, so a streamed set and a measured one describe the same population. They are expressed against the selection table’s columns rather than against a measurement frame, because the array route has no measurements – only labels and geometry.Recognised settings, each optional and each skipped when absent:
{object}_min_size/{object}_max_sizeObject area in pixels.
0andNoneboth mean “no bound”, which is what the Measure panel writes for an unset field – treating 0 as a real minimum would drop nothing and look like it had worked.wells/exclude_wellsWell ids to keep or drop, as
rowID+columnIDpairs or as plain well names.max_objectsA cap, applied LAST and deterministically (by the sort the selection already has), so a capped set is reproducible.
- Parameters:
selection – a selection table from
spacr.stream_dataset.settings – the run settings.
- Returns:
a new frame; the input is not modified.
- spacr.annotation_dataset.generate_annotation_dataset(settings: Mapping[str, Any]) Dict[str, Any][source]¶
Stream a set of crops and register it for annotation.
- Parameters:
settings – needs
src(the plate folder) and acceptsstream_source(STREAM_SOURCES),object_array,channel_arrays,bounding_box, the filtration keysfilter_selection()reads, anddstfor where the crops go.- Returns:
the streaming report, plus
tablenaming what was written.
- spacr.annotation_dataset.next_png_table(connection: sqlite3.Connection) str[source]¶
The name a new annotation set should be written under.
png_listwhen it is free, thenpng_list_2,png_list_3and so on. An existing set is never overwritten: it may already carry annotations, and those are hand-made and unrecoverable.THE CALLER MUST HOLD A WRITE TRANSACTION. Choosing a name and creating the table are two steps, and two runs started together would otherwise choose the same one.
write_png_list()does this correctly; call it rather than this.- Parameters:
connection – an open connection to the measurements database.
- Returns:
the free table name.
- spacr.annotation_dataset.png_list_frame(selection: pandas.DataFrame, paths: Sequence[str]) pandas.DataFrame[source]¶
The rows the annotation viewer reads, in
measure_crop’s own schema.- Parameters:
selection – the filtered selection table.
paths – the written crop path for each of its rows, in order.
- Returns:
a frame with exactly
PNG_LIST_COLUMNS.
- spacr.annotation_dataset.read_objects_from_database(db_path: str, object_type: str) pandas.DataFrame | None[source]¶
The object table a streamed set can be built from.
- Parameters:
db_path – the measurements database.
object_type –
cell,nucleus,pathogen…
- Returns:
the rows, or
Nonewhen the table is not there.
- spacr.annotation_dataset.reserve_png_table(db_path: str) str[source]¶
Claim the next free table name by creating it, empty.
RESERVED BEFORE THE CROPS ARE CUT, not after, so the folder they are written into can be named to match:
png_list_2getsdata_2. Deriving the folder from the table is what makes a set on disk traceable to the set in the database – with two independent counters they drift the first time either is deleted, and then nothing says which folder a table describes.Creating the table is what reserves it: two runs started together would otherwise choose the same name, and the second would fail on insert after it had already written a folder full of crops.
- Parameters:
db_path – the measurements database.
- Returns:
the reserved table name.
- spacr.annotation_dataset.write_png_list(db_path: str, frame: pandas.DataFrame, *, table: str | None = None) str[source]¶
Write an annotation set into the measurements database.
The name is chosen and the table created inside ONE transaction, so two runs started together cannot pick the same one. A caller that already reserved a name with
reserve_png_table()passes it astable.- Parameters:
db_path – the measurements database.
frame – rows as
png_list_frame()builds them.table – a name already reserved.
- Returns:
the table actually written.
Nested helpers¶
- generate_annotation_dataset._write(path, array)¶
Write and register one streamed crop through Measure’s writer.
- Parameters:
path – proposed crop path; its extension is replaced with
.png.array – crop array to normalize, select, pad, resize, and save.
- Returns:
None. The shared writer’s returned path is appended to the captured list using the captured channels and PNG size, preserving alignment with the subsequently registered selection rows.
spacr/annotation_dataset.py:341