spacr.plate_measurements

One merged measurements frame from the plate rows of the input table.

A row of the regression input table is one plate: its score CSV, its count CSV and now its measurements database. This module is the headless half of that feature – given {plate: database path}, the tables the user ticked and the anchor they chose, it returns the merged frame.

NOTHING HERE IS NEW MERGE LOGIC, AND THAT IS THE POINT. spaCR already has four places that join measurement tables and one of them (io._read_and_join_tables) aggregates every numeric column with mean, which turns four pathogens’ total area into an average area and a MINIMUM into a mean of minima. A fifth would be a fifth answer. So:

There is no sum() or mean() of a measurement column written in this file, and there must never be one.

WHY THE TWO EXISTING FUNCTIONS DO NOT COMPOSE BY THEMSELVES

read_merged is many databases, one table. merge_tables is one database, many tables – and it takes a PATH, so it cannot be handed a frame that has already been stacked across databases. This feature needs both, so the composition is: read each chosen table across every database, then roll the children onto the anchor.

The roll-up keys are the load-bearing detail. They are the identity columns PLUS screenID PLUS source_database PLUS the child’s anchor column. Leave screenID out and two screens that legitimately share plate1 (which describe_merge() deliberately permits) collapse into one parent – reintroducing, one layer up, exactly the pooling multi_database exists to prevent.

WHAT A CALLER MUST SHOW THE USER

A merge that silently changed how a measurement was combined produces a number that is wrong and looks fine, so PlateMerge carries what the panel has to say: the anchor and the row count, the rows each source contributed, the measurements columns='common' dropped, the plate ids shared across screens – and every column that fell through to spacr.merge_tables.DEFAULT_AGGREGATION because no rule matched it. That last set is recomputed by re-walking AGGREGATION_RULES (default_aggregated_columns()) rather than listed anywhere, so it cannot drift out of step with the rules.

Classes

PlateDatabase

One plate row of the input table and the database attached to it.

PlateMerge

The merged frame and the Measurements tab's audit information.

TableMerge

What one chosen table contributed to the merge, and how.

Functions

ambiguous_identifiers(→ Dict[str, Dict[str, Any]])

Text identifiers that are NOT constant within their roll-up group.

available_tables(→ Tuple[str, ...])

The object tables present in EVERY attached database.

classify_default_columns(→ Dict[str, Tuple[str, ...]])

Split the columns no rule names into what will ACTUALLY happen to them.

default_aggregated_columns(→ Tuple[str, ...])

The columns in plan that got the default because NO rule matched.

describe_identifier_refusal(→ str)

One line saying why column was left out, with an example.

merge_plate_databases(, *, anchor, policy, screens, ...)

Merge every attached plate database into one frame, one row per anchor.

missing_databases(→ Tuple[PlateDatabase, ...])

Attached databases that are not on disk, named before the run starts.

plate_databases(→ Tuple[PlateDatabase, ...])

The plate rows that HAVE a database, in the order the table lists them.

unattached_plates(→ Tuple[str, ...])

The plates with no database, which is legal and must be said out loud.

Module Contents

class spacr.plate_measurements.PlateDatabase[source]

One plate row of the input table and the database attached to it.

Parameters:
  • plate – input-table plate label used to map screen metadata and identify missing or duplicate attachments in diagnostics. A blank label normalised from an input row becomes "row N".

  • path – filesystem path to the plate’s measurements database. An empty value means no database is attached; a nonempty path is checked for existence before table discovery and merging.

property exists: bool[source]

Whether the file is still where the input table says it is.

Checked before the run rather than during it: the design asks that a database that has been moved is named up front, not four minutes into a regression.

class spacr.plate_measurements.PlateMerge[source]

The merged frame and the Measurements tab’s audit information.

Parameters:
  • frame – final measurements frame, with one row per surviving anchor object.

  • anchor – normalized object-table name whose rows define the output cardinality.

  • attachments – validated plate/database attachments included in the merge, in input order.

  • tables – per-table audit records in merge order, beginning with the anchor and including any child table skipped because it could not be linked.

describe() → str[source]

The disclosure, as lines – what this merge did and what it cost.

property default_aggregation_columns: Tuple[str, ...][source]

Merged columns whose aggregation is the default, no rule matching.

spacr.merge_tables.DEFAULT_AGGREGATION is MEAN and the comment beside it says why – but a measurement nobody thought about is exactly the one worth naming, because it is the one where the default is most likely to be answering a different question.

property dropped_columns: Tuple[str, ...][source]

Measurements columns='common' left out, as merged names.

A dropped measurement is a measurement the user came to compare, so it is reported rather than merely defaulted.

property rows: int[source]

How many anchor objects survived the merge.

property rows_per_source: Dict[str, int][source]

Anchor rows each database contributed to the FINAL frame.

Less than rows_read_per_source wherever an inner join removed objects – a nucleus-less cell, or a pathogen-less one when keep_uninfected=False. The pair is what lets a panel say how many rows were dropped instead of only how many are left.

property rows_read_per_source: Dict[str, int][source]

Anchor rows each database HELD, before any join dropped any.

property shared_plates_across_screens: Dict[str, Tuple[str, ...]][source]

Plate ids that appear in more than one screen. NOT a collision.

Two screens sharing a guide library both have plate1, and once screenID is in the frame those are two identities. Reported anyway, because a user who did not mean to run two screens still needs to see that they did.

property sources: Tuple[str, ...][source]

The database labels, as spacr.multi_database named them.

class spacr.plate_measurements.TableMerge[source]

What one chosen table contributed to the merge, and how.

One of these per table the user ticked, including the anchor. It is the record a panel reads to answer “what happened to my numbers”: which aggregation each column got, whether the table was rolled up or joined directly, and which join mode the merge policy selected.

Parameters:
  • table – Chosen source object-table name, including the anchor table.

  • plan – Multi-database read plan used to read this table.

  • rows – Number of source rows read across the attached databases before joining or child roll-up.

  • keys – Source identity, join, or grouping columns retained without a table prefix.

  • how – Child-table join mode returned by spacr.merge_tables.MergePolicy.how_for(); empty for the anchor or a skipped table.

  • rolled_up – Whether this child table’s rows were aggregated onto the anchor; false for the anchor, a directly joined one-row-per-cell table, or a skipped table.

  • aggregations – Source-column-to-aggregation mapping actually applied during child roll-up; empty for the anchor, a direct join, or a skipped table.

  • default_columns – Source columns in aggregations that used the default because no named aggregation rule matched and no override applied.

  • dropped – Source columns absent from at least one attached database and therefore omitted by the common-column read.

  • note – Explanation when the table was skipped and contributed no merged columns; empty otherwise.

merged_column(column: str) → str[source]

The name column carries in the merged frame.

Parameters:

column – source-table column name to express in merged-frame form.

The same rule spacr.merge_tables.roll_up() applies: a join key keeps its name, a column that already starts with the table’s name is left alone – nucleus_area must not become nucleus_nucleus_area – and everything else is prefixed with the table it measures.

spacr.plate_measurements.ambiguous_identifiers(child: pandas.DataFrame, keys: Sequence[str], *, plan: Mapping[str, str] | None = None, overrides: Mapping[str, str] | None = None, examples: int = 1) → Dict[str, Dict[str, Any]][source]

Text identifiers that are NOT constant within their roll-up group.

arrived at from the other side: when two things that must agree do not, that is a real inconsistency and it is named rather than silently resolved. A cell with three pathogens has ONE path_name if all three came off the same image; if it has three, first names one of them and the merged row then claims a provenance the data does not support.

Only TEXT columns are examined, and this is deliberate. object_label also takes first, by the identity rule, and it is SUPPOSED to differ across the children – spacr.merge_tables.AGGREGATION_RULES says so in as many words: the label is carried verbatim only so a row can be traced back. A numeric label that varies is the normal case; a text identifier that varies is the ambiguity.

Parameters:
  • child – the many-per-parent table, before the roll-up.

  • keys – the group keys the roll-up will use.

  • plan – the aggregation plan, from spacr.merge_tables.aggregation_plan(). Recomputed when omitted.

  • overrides – the caller’s explicit choices. A column they set is theirs, and is not second-guessed here.

  • examples – how many offending group keys to record per column.

Returns:

{column: {'groups': n, 'examples': [(key, [values])]}} – empty when every identifier is constant, which is the normal case.

spacr.plate_measurements.available_tables(attachments: Any) → Tuple[str, ...][source]

The object tables present in EVERY attached database.

Parameters:

attachments – input-table attachment rows whose existing databases define the table intersection.

The intersection, not the union, and this is not a nicety: spacr.multi_database.describe_merge() raises a bare sqlite3.OperationalError: no such table when one database lacks the chosen table, which reaches a user as a crash rather than as a choice they were never offered.

Returns:

the tables in spacr.merge_tables.OBJECT_TABLES order, so the picker’s order is the registry’s order rather than sqlite’s, and so a sixth object kind reaches this list by being declared once. png_list is not offered: it holds one row per CROP rather than per object, and a crop is not a measurement to aggregate.

spacr.plate_measurements.classify_default_columns(columns: Sequence[str], kinds: Mapping[str, str] | None = None, *, overrides: Mapping[str, str] | None = None) → Dict[str, Tuple[str, ...]][source]

Split the columns no rule names into what will ACTUALLY happen to them.

Parameters:
  • columns – the column names about to be aggregated.

  • kinds – {column: 'numeric' | 'text' | 'unknown'}, as spacr.multi_database.column_kinds() returns it. A column absent from the mapping is 'unknown'.

  • overrides – the caller’s explicit choices. A column they set is not a fall-through and appears in neither bucket – they chose it.

Returns:

{'mean': (...), 'identifier': (...), 'unknown': (...)}:

mean

numeric, no rule matched, so spacr.merge_tables.DEFAULT_AGGREGATION is what it takes.

identifier

text. It takes spacr.merge_tables.TEXT_AGGREGATION, not a mean, and whether it is CARRIED or REFUSED depends on the data – see ambiguous_identifiers(), which can only be answered by reading the rows.

unknown

the database declared no type, so neither can be promised. Named separately rather than folded into either: an absent answer that reads as a definite one is the failure this module exists to avoid.

spacr.plate_measurements.default_aggregated_columns(plan: Mapping[str, str], *, overrides: Mapping[str, str] | None = None) → Tuple[str, ...][source]

The columns in plan that got the default because NO rule matched.

Recomputed by re-walking spacr.merge_tables.AGGREGATION_RULES, so this cannot drift out of step with them: a rule added for a measurement tomorrow removes that measurement from this list the same day, and a rule deleted puts it back.

A column the caller overrode is not reported however it was set – they chose it. Nor is a text column, which takes spacr.merge_tables.TEXT_AGGREGATION because text does not add up, not because nobody thought about it.

Parameters:
spacr.plate_measurements.describe_identifier_refusal(table: str, column: str, detail: Mapping[str, Any]) → str[source]

One line saying why column was left out, with an example.

Parameters:
  • table – child table whose identifier could not be carried safely.

  • column – varying text-identifier column that was omitted.

  • detail – ambiguity record from ambiguous_identifiers(), including the affected group count and optional examples.

Written here rather than in the panel so the headless merge and the Qt one say the same sentence.

spacr.plate_measurements.merge_plate_databases(attachments: Any, tables: Sequence[str] = (), *, anchor: str | None = None, policy: spacr.merge_tables.MergePolicy | None = None, screens: Any = None, columns: str = 'common', report: Callable[[str], None] | None = None, on_ambiguous_identifier: str = 'refuse') → PlateMerge[source]

Merge every attached plate database into one frame, one row per anchor.

Calls spacr.multi_database.read_merged() once per chosen table and spacr.merge_tables.roll_up() once per child table. It performs no aggregation of its own, and the join type comes from spacr.merge_tables.MergePolicy.how_for() per table rather than from a blanket how.

Parameters:
  • attachments – any shape plate_databases() accepts. Plates with no database are skipped, not fatal.

  • tables – the object tables to join. The anchor is added if absent.

  • anchor – the object a row of the result MEANS. cell by default (spacr.merge_tables.DEFAULT_PRIMARY), overriding policy.primary when both are given.

  • policy – the aggregation and join policy – na, overrides, consolidate_on_cell, keep_uninfected.

  • screens – optional screen label per plate: a mapping keyed by plate, or a sequence parallel to the attached databases.

  • columns – 'common' or 'union', passed to read_merged.

  • report – called with one line per thing the merge cost, the table named.

  • on_ambiguous_identifier – 'refuse' (default) leaves out a text identifier that differs within a roll-up group and names it on report; 'first' restores the old silent pick. See ambiguous_identifiers() – and note that this is the same default the Measurements tab applies, because two answers to one question is how two merges of one screen come to disagree.

Returns:

the PlateMerge.

Raises:

MergeRefused – nothing attached, a database that has moved, one database attached to two plates, an anchor that is not one row per cell, a table missing from some database, or – from read_merged itself – a plate id repeated inside one screen.

spacr.plate_measurements.missing_databases(attachments: Any) → Tuple[PlateDatabase, ...][source]

Attached databases that are not on disk, named before the run starts.

Parameters:

attachments – input-table attachment rows in any shape accepted by plate_databases().

spacr.plate_measurements.plate_databases(attachments: Any) → Tuple[PlateDatabase, ...][source]

The plate rows that HAVE a database, in the order the table lists them.

Parameters:

attachments – {plate: path}, a sequence of PlateDatabase, a sequence of (plate, path) pairs, or the input table’s own rows (mappings carrying plate and database).

Returns:

one PlateDatabase per attached row. Rows with no database are left out rather than carried as blanks – see unattached_plates(), which is where they are named.

spacr.plate_measurements.unattached_plates(attachments: Any) → Tuple[str, ...][source]

The plates with no database, which is legal and must be said out loud.

Parameters:

attachments – input-table attachment rows in any shape accepted by plate_databases().

The regression runs on scores and counts; the database is what makes the Measurements tab possible for that plate. Its absence disables that plate there rather than failing the run – so the plate is listed, and the listing is what stops it looking like an omission.