spacr.crop_loader

Choose a set of crops, then read them a page at a time.

WHAT IT IS FOR. Every screen that shows objects has had its own way of finding them, and the Embeddings screen had none at all – spacr.qt.screens.embeddings.EmbeddingsScreen.set_crops() existed and nothing in spacr/ called it, so the only route to a label-free vector was to open Python. This module is the missing half: it turns a question about a plate into a stack of pixels, and it does it in two steps so a GUI can stay answerable between them.

PLANNING IS NOT LOADING, and that separation is the whole design. plan_crops() asks the database or the folder which objects match and reads no pixels at all: a COUNT and a SELECT of identifiers, or one scandir. It is milliseconds on a plate of sixty thousand crops, so a screen can run it on a click and tell the user what they asked for before committing to reading any of it. load_crops() then reads the pixels in pages, reporting after each one. A caller that runs the second half on a worker thread gets a progress number per page and a stop that takes effect within a page rather than at the end.

WHAT IT NEEDS. Either a measurements.db – where the crops are chosen by object class, by plate, or by a predicate over png_list – or a folder of crop PNGs. Nothing else: the pixels themselves come from spacr.crops.resolve_crop_source(), so this module inherits the existing answer to “are there pre-generated crops here, or do we cut from merged/*.npy” rather than inventing a second one, and a legacy folder is corrected on load exactly as it is everywhere else.

WHAT IT PRODUCES. One (objects, height, width, channels) uint8 array, which is the layout spacr.embeddings.embed_array() documents and spacr.qt.screens.embeddings.EmbeddingsScreen.set_crops() requires, plus a JSON-serialisable record of what was actually read – how many matched, how many were taken, the crop shape, and how many crops had to be conformed to it.

WHAT TO DO NEXT. Hand the array to the Embeddings screen, or straight to spacr.embeddings.embed_array().

A CAP IS PART OF THE ANSWER, NOT A FAILURE TO LOAD EVERYTHING. Sixty thousand crops of 96x96x3 is 1.6 GB before a backbone has seen one of them, and a screen that tries it dies on a machine that could have embedded the first two thousand happily. So CropQuery.limit is a real default and the plan says what it left behind – “the first 2,000 of 61,433” – rather than quietly returning a subset that looks like the whole plate.

AND A CAP IS NOT THE ONLY THING THAT REMOVES A ROW, SO IT IS COUNTED ON ITS OWN. The merged route drops every row it cannot cut, and reporting that as a cap is the same sentence read backwards: a complete answer dressed as a subset, telling the user to raise a limit that never bit. So CropPlan.matched, CropPlan.selected and CropPlan.count are three numbers, CropPlan.capped is the gap between the first two and CropPlan.dropped the gap between the last two, and they reach the panel as separate clauses.

AN EMPTY RESULT IS A SENTENCE, NEVER AN EMPTY ARRAY. A plan that matched nothing carries CropPlan.empty_reason, which names the table, the class, the plate and the predicate that were asked for AND what the database does hold, because “no crops” on its own sends a user to look for a broken install when they have asked for plate 9 of an eight-plate screen.

THE PREDICATE IS SQL AND THE DATABASE IS OPENED READ-ONLY. where is a fragment of the user’s own query language against their own file; refusing it would mean inventing a worse one. It is spliced into a WHERE clause, so the connection is opened with readonly=True – SQLite’s query_only, which refuses every write at the engine – and a predicate carrying a statement separator is rejected before it is sent.

Exceptions

CropLoadError

A set of crops that cannot be loaded, and why.

Classes

CropPlan

What a query matched, before any pixel was read.

CropQuery

Which crops to load, and how much of the answer to take.

Functions

conform_crop(→ numpy.ndarray)

Centre a crop in a height x width frame, padding or trimming.

load_crops(→ numpy.ndarray)

Read every crop in plan, a page at a time.

load_page(→ Tuple[numpy.ndarray, int])

Read one page of plan's crops.

object_classes(→ Tuple[str, ...])

Which object classes this database actually has crops for.

plan_crops(→ CropPlan)

Answer query without reading a pixel.

plan_from_database(→ CropPlan)

Choose crops out of a measurements.db, reading no pixels.

plan_from_folder(→ CropPlan)

Choose crops out of a folder of *.png, reading no pixels.

plates(→ Tuple[str, ...])

Every plate the crop table names, sorted.

Module Contents

exception spacr.crop_loader.CropLoadError[source]

Bases: ValueError

A set of crops that cannot be loaded, and why.

Carries a sentence fit for a status bar: every raise in this module names the path, the class or the predicate that produced it, because the caller is a screen and its only way to explain a failure is to repeat this text.

Initialize self. See help(type(self)) for accurate signature.

class spacr.crop_loader.CropPlan[source]

What a query matched, before any pixel was read.

Parameters:
  • query – the CropQuery this answers.

  • rows – one opaque handle per crop, in the order they load – a row mapping for the database source, a path string for the folder one. Never longer than query.limit.

  • matched – how many crops the query matched in total, before the limit and before the merged route’s join. More than selected exactly when the cap bit.

  • selected – how many rows the selection actually took, which is min(matched, limit). Left at 0 it defaults to len(rows), which is right for every plan that has no join to lose rows to.

  • source_label – which pixel route was chosen and why, from spacr.crops.CropSource.describe().

  • empty_reason – the sentence to show when nothing matched, naming both what was asked and what is there. Empty when something did.

THREE NUMBERS, NOT TWO, BECAUSE TWO THINGS REMOVE ROWS AND THEY ASK THE READER FOR DIFFERENT THINGS. matched is what the query found, selected is what the limit left of it, and count is what survived the merged route’s join – which drops every row with no single object label ('omulti' / 'onone') or no merged array recorded. Only the first gap is a cap, and only a cap is repaired by raising ‘At most’; telling a user to raise a limit that never bit hands them an action that returns the same rows and reprints the same sentence.

__post_init__() → None[source]

Default selected to the rows, for a plan built without a join.

A plan whose rows came straight out of the selection – the folder route, and anything a test or a script builds by hand – selected exactly what it carries, and should not have to say so.

describe() → str[source]

One line a status bar can show the moment planning returns.

Says what was taken out of what, so a capped load reads as a deliberate subset rather than as the whole plate – and names a dropped row as a dropped row, in its own clause, so the two never arrive as one number.

pages(page_size: int = 0) → Tuple[Tuple[int, int], ...][source]

(start, stop) for each page, covering every row exactly once.

Parameters:

page_size – crops per page; 0 takes the query’s own.

property capped: bool[source]

Whether the limit kept crops out of this plan.

Against selected, never against count: a merged-route load that lost rows to the join has matched > count with no limit anywhere near it, and reporting that as a cap is this module’s own promise – that a cap is part of the answer – running backwards.

property count: int[source]

How many crops this plan will load.

property dropped: int[source]

How many selected rows could not be turned into a crop at all.

Nonzero only on the merged route, and a complete answer rather than a capped one: the rows are gone because they cannot be cut, so there is nothing a larger limit would add.

property is_empty: bool[source]

Whether the query matched nothing at all.

class spacr.crop_loader.CropQuery[source]

Which crops to load, and how much of the answer to take.

Parameters:
  • source – CROP_SOURCE_DATABASE or CROP_SOURCE_FOLDER.

  • path – the measurements.db, or the folder of crop PNGs.

  • object_type – which object class the crops are of – 'cell', 'nucleus', 'pathogen', 'cytoplasm' or an organelle role. It picks the object-id column in png_list and the mask plane a streamed crop is cut by, so it is not cosmetic: asking for nuclei out of a cell png_list yields nothing rather than yielding cells. Ignored by the folder source, which has no classes.

  • plate – one plate, or '' for every plate.

  • where – an SQL predicate over png_list, or ''. Spliced into the WHERE clause of a read-only connection; see the module docstring for why it is accepted at all.

  • limit – the most crops to take, 0 for no cap. The plan records what the cap left behind.

  • page_size – crops read per page by load_crops().

  • prefer – 'png' to read pre-generated crops, 'merged' to cut them from merged/*.npy, '' to let spacr.crops.resolve_crop_source() decide. Database source only.

describe() → str[source]

One line naming everything this query asked for.

Used in the status bar and in every refusal, so a user reading “no crops” can see the question that produced it without reopening the panel.

spacr.crop_loader.conform_crop(crop: numpy.ndarray, height: int, width: int) → numpy.ndarray[source]

Centre a crop in a height x width frame, padding or trimming.

Padding rather than resizing, and that is the whole of the decision: a stack has to be rectangular before a backbone sees it, and rescaling a crop changes the apparent size of the object in it, which is a phenotype. Zero-padding changes the background, which is not. It is also what spacr.crop_source.crop_at() already does to a box that runs off the edge of a field, so two crops of the same object arrive the same way whichever route cut them.

Parameters:
  • crop – (height, width, channels).

  • height – the target first axis.

  • width – the target second axis.

Returns:

crop itself when it already fits, so the common case copies nothing.

spacr.crop_loader.load_crops(plan: CropPlan, *, progress: Callable[[int, int], None] | None = None, cancelled: Callable[[], bool] | None = None, record: MutableMapping[str, Any] | None = None) → numpy.ndarray[source]

Read every crop in plan, a page at a time.

Nothing here touches a widget and nothing here is Qt: progress and cancelled are plain callables, so a screen passes a signal emit and a flag and the same function drives a script or a test.

Parameters:
  • plan – from plan_crops().

  • progress – called (done, total) after each page, on whatever thread this runs on. A caller on a worker thread must not touch a widget from it – emit a signal.

  • cancelled – called before each page; a true answer stops the load and returns the pages already read, which is why the result can be shorter than plan.count. Checked between pages, so a stop takes at most one page to take effect.

  • record – filled in with what was read – matched, selected, dropped, loaded, crop_shape, conformed, capped, source, stopped. matched minus selected is the cap; selected minus loaded is rows that could not be cut, plus whatever a stop left unread. JSON-serialisable, so a caller can keep it beside the matrix.

Returns:

(objects, height, width, channels), the layout spacr.embeddings.embed_array() documents. Every crop is padded or trimmed to the first one’s frame; see conform_crop().

Raises:

CropLoadError – the plan is empty, or a page could not be read.

spacr.crop_loader.load_page(plan: CropPlan, start: int, stop: int, shape: Tuple[int, int, int] | None = None) → Tuple[numpy.ndarray, int][source]

Read one page of plan’s crops.

Parameters:
  • plan – from plan_crops().

  • start – first row of the page.

  • stop – one past its last row.

  • shape – (height, width, channels) every crop must come back as. None takes the first crop of the page as the template, which is how load_crops() fixes the stack’s shape from page one.

Returns:

(page, conformed) – the (n, height, width, channels) array, and how many of its crops had to be padded or trimmed to fit.

Raises:

CropLoadError – a crop could not be read, or came back with a different number of channels than the page’s template – which is a real difference between two crops and is not silently padded away.

spacr.crop_loader.object_classes(db_path: str, table: str = CROP_TABLE) → Tuple[str, ...][source]

Which object classes this database actually has crops for.

A class is offered only when its id column is present AND holds at least one value, because a png_list written for cells carries an empty nucleus_id column on some schemas and offering it would put a choice in the panel that can only ever return nothing.

Parameters:
  • db_path – the measurements.db.

  • table – the crop table; 'png_list'.

Returns:

class names in spacr.png_list.PNG_LIST_ID_COLUMNS order. Empty when the database has no crop table.

spacr.crop_loader.plan_crops(query: CropQuery) → CropPlan[source]

Answer query without reading a pixel.

Parameters:

query – what to select.

Returns:

a CropPlan. Check CropPlan.is_empty before loading – an empty plan is an answer, not a failure, and it carries the sentence explaining itself.

Raises:

CropLoadError – an unknown source, or a path that is not there.

spacr.crop_loader.plan_from_database(query: CropQuery) → CropPlan[source]

Choose crops out of a measurements.db, reading no pixels.

The rows come back carrying png_path – what spacr.crops.PngCropSource reads – and, through spacr.png_list.crop_rows_from_png_list(), path_name and an integer object_label, which is what spacr.crops.MergedCropSource needs. So the plan is good for whichever route spacr.crops.resolve_crop_source() picks, and the route is picked here, once, rather than per crop.

Parameters:

query – what to select. query.source is not re-checked.

Returns:

a CropPlan, possibly an empty one carrying CropPlan.empty_reason.

Raises:

CropLoadError – the database is missing, has no crop table, or refused the selection.

spacr.crop_loader.plan_from_folder(query: CropQuery) → CropPlan[source]

Choose crops out of a folder of *.png, reading no pixels.

Parameters:

query – what to select. Only path and limit apply: a folder records no class and no plate, and saying so in CropPlan.source_label is better than filtering on a filename convention that no run guarantees.

Returns:

a CropPlan, possibly an empty one carrying CropPlan.empty_reason.

Raises:

CropLoadError – no folder was chosen, or it is not a directory.

spacr.crop_loader.plates(db_path: str, table: str = CROP_TABLE) → Tuple[str, ...][source]

Every plate the crop table names, sorted.

Parameters:
  • db_path – the measurements.db.

  • table – the crop table; 'png_list'.

Returns:

plate identifiers as strings. Empty when the table has no plateID column, which is how a database written before plate keys existed reads rather than an error.