spacr.plate_measurements¶
One merged measurements frame from the plate rows of the input table.
A row of the regression input table is one plate: its score
CSV, its count CSV and now its measurements database. This module is the
headless half of that feature – given {plate: database path}, the tables
the user ticked and the anchor they chose, it returns the merged frame.
NOTHING HERE IS NEW MERGE LOGIC, AND THAT IS THE POINT. spaCR already has
four places that join measurement tables and one of them
(io._read_and_join_tables) aggregates every numeric column with mean,
which turns four pathogens’ total area into an average area and a MINIMUM
into a mean of minima. A fifth would be a fifth answer. So:
spacr.multi_database.describe_merge()/read_merged()stack the databases – one table at a time, carryingsource_databaseandscreenID, refusing a plate id repeated inside one screen;spacr.merge_tables.roll_up()aggregates each child table onto the anchor, one rule per column out ofspacr.merge_tables.AGGREGATION_RULES;spacr.merge_tables.MergePolicy.how_for()decides the join PER TABLE, because the cardinality differs – a cell has one cytoplasm and many pathogens, and an uninfected cell is still a cell.
There is no sum() or mean() of a measurement column written in this
file, and there must never be one.
WHY THE TWO EXISTING FUNCTIONS DO NOT COMPOSE BY THEMSELVES¶
read_merged is many databases, one table. merge_tables is one
database, many tables – and it takes a PATH, so it cannot be handed a frame
that has already been stacked across databases. This feature needs both, so
the composition is: read each chosen table across every database, then roll
the children onto the anchor.
The roll-up keys are the load-bearing detail. They are the identity columns
PLUS screenID PLUS source_database PLUS the child’s anchor column.
Leave screenID out and two screens that legitimately share plate1
(which describe_merge() deliberately permits)
collapse into one parent – reintroducing, one layer up,
exactly the pooling multi_database exists to prevent.
WHAT A CALLER MUST SHOW THE USER¶
A merge that silently changed how a measurement was combined produces a number
that is wrong and looks fine, so PlateMerge carries what the panel has
to say: the anchor and the row count, the rows each source contributed, the
measurements columns='common' dropped, the plate ids shared across screens
– and every column that fell through to
spacr.merge_tables.DEFAULT_AGGREGATION because no rule matched it. That
last set is recomputed by re-walking AGGREGATION_RULES
(default_aggregated_columns()) rather than listed anywhere, so it cannot
drift out of step with the rules.
Classes¶
One plate row of the input table and the database attached to it. |
|
The merged frame and the Measurements tab's audit information. |
|
What one chosen table contributed to the merge, and how. |
Functions¶
|
Text identifiers that are NOT constant within their roll-up group. |
|
The object tables present in EVERY attached database. |
|
Split the columns no rule names into what will ACTUALLY happen to them. |
|
The columns in |
|
One line saying why |
|
Merge every attached plate database into one frame, one row per anchor. |
|
Attached databases that are not on disk, named before the run starts. |
|
The plate rows that HAVE a database, in the order the table lists them. |
|
The plates with no database, which is legal and must be said out loud. |
Module Contents¶
- class spacr.plate_measurements.PlateDatabase[source]¶
One plate row of the input table and the database attached to it.
- Parameters:
plate – input-table plate label used to map screen metadata and identify missing or duplicate attachments in diagnostics. A blank label normalised from an input row becomes
"row N".path – filesystem path to the plate’s measurements database. An empty value means no database is attached; a nonempty path is checked for existence before table discovery and merging.
- class spacr.plate_measurements.PlateMerge[source]¶
The merged frame and the Measurements tab’s audit information.
- Parameters:
frame – final measurements frame, with one row per surviving anchor object.
anchor – normalized object-table name whose rows define the output cardinality.
attachments – validated plate/database attachments included in the merge, in input order.
tables – per-table audit records in merge order, beginning with the anchor and including any child table skipped because it could not be linked.
- property default_aggregation_columns: Tuple[str, ...][source]¶
Merged columns whose aggregation is the default, no rule matching.
spacr.merge_tables.DEFAULT_AGGREGATIONis MEAN and the comment beside it says why – but a measurement nobody thought about is exactly the one worth naming, because it is the one where the default is most likely to be answering a different question.
- property dropped_columns: Tuple[str, ...][source]¶
Measurements
columns='common'left out, as merged names.A dropped measurement is a measurement the user came to compare, so it is reported rather than merely defaulted.
- property rows_per_source: Dict[str, int][source]¶
Anchor rows each database contributed to the FINAL frame.
Less than
rows_read_per_sourcewherever an inner join removed objects – a nucleus-less cell, or a pathogen-less one whenkeep_uninfected=False. The pair is what lets a panel say how many rows were dropped instead of only how many are left.
- property rows_read_per_source: Dict[str, int][source]¶
Anchor rows each database HELD, before any join dropped any.
Plate ids that appear in more than one screen. NOT a collision.
Two screens sharing a guide library both have
plate1, and oncescreenIDis in the frame those are two identities. Reported anyway, because a user who did not mean to run two screens still needs to see that they did.
- property sources: Tuple[str, ...][source]¶
The database labels, as
spacr.multi_databasenamed them.
- class spacr.plate_measurements.TableMerge[source]¶
What one chosen table contributed to the merge, and how.
One of these per table the user ticked, including the anchor. It is the record a panel reads to answer “what happened to my numbers”: which aggregation each column got, whether the table was rolled up or joined directly, and which join mode the merge policy selected.
- Parameters:
table – Chosen source object-table name, including the anchor table.
plan – Multi-database read plan used to read this table.
rows – Number of source rows read across the attached databases before joining or child roll-up.
keys – Source identity, join, or grouping columns retained without a table prefix.
how – Child-table join mode returned by
spacr.merge_tables.MergePolicy.how_for(); empty for the anchor or a skipped table.rolled_up – Whether this child table’s rows were aggregated onto the anchor; false for the anchor, a directly joined one-row-per-cell table, or a skipped table.
aggregations – Source-column-to-aggregation mapping actually applied during child roll-up; empty for the anchor, a direct join, or a skipped table.
default_columns – Source columns in
aggregationsthat used the default because no named aggregation rule matched and no override applied.dropped – Source columns absent from at least one attached database and therefore omitted by the common-column read.
note – Explanation when the table was skipped and contributed no merged columns; empty otherwise.
- merged_column(column: str) str[source]¶
The name
columncarries in the merged frame.- Parameters:
column – source-table column name to express in merged-frame form.
The same rule
spacr.merge_tables.roll_up()applies: a join key keeps its name, a column that already starts with the table’s name is left alone –nucleus_areamust not becomenucleus_nucleus_area– and everything else is prefixed with the table it measures.
- spacr.plate_measurements.ambiguous_identifiers(child: pandas.DataFrame, keys: Sequence[str], *, plan: Mapping[str, str] | None = None, overrides: Mapping[str, str] | None = None, examples: int = 1) Dict[str, Dict[str, Any]][source]¶
Text identifiers that are NOT constant within their roll-up group.
arrived at from the other side: when two things that must agree do not, that is a real inconsistency and it is named rather than silently resolved. A cell with three pathogens has ONE
path_nameif all three came off the same image; if it has three,firstnames one of them and the merged row then claims a provenance the data does not support.Only TEXT columns are examined, and this is deliberate.
object_labelalso takesfirst, by the identity rule, and it is SUPPOSED to differ across the children –spacr.merge_tables.AGGREGATION_RULESsays so in as many words: the label is carried verbatim only so a row can be traced back. A numeric label that varies is the normal case; a text identifier that varies is the ambiguity.- Parameters:
child – the many-per-parent table, before the roll-up.
keys – the group keys the roll-up will use.
plan – the aggregation plan, from
spacr.merge_tables.aggregation_plan(). Recomputed when omitted.overrides – the caller’s explicit choices. A column they set is theirs, and is not second-guessed here.
examples – how many offending group keys to record per column.
- Returns:
{column: {'groups': n, 'examples': [(key, [values])]}}– empty when every identifier is constant, which is the normal case.
- spacr.plate_measurements.available_tables(attachments: Any) Tuple[str, ...][source]¶
The object tables present in EVERY attached database.
- Parameters:
attachments – input-table attachment rows whose existing databases define the table intersection.
The intersection, not the union, and this is not a nicety:
spacr.multi_database.describe_merge()raises a baresqlite3.OperationalError: no such tablewhen one database lacks the chosen table, which reaches a user as a crash rather than as a choice they were never offered.- Returns:
the tables in
spacr.merge_tables.OBJECT_TABLESorder, so the picker’s order is the registry’s order rather than sqlite’s, and so a sixth object kind reaches this list by being declared once.png_listis not offered: it holds one row per CROP rather than per object, and a crop is not a measurement to aggregate.
- spacr.plate_measurements.classify_default_columns(columns: Sequence[str], kinds: Mapping[str, str] | None = None, *, overrides: Mapping[str, str] | None = None) Dict[str, Tuple[str, ...]][source]¶
Split the columns no rule names into what will ACTUALLY happen to them.
- Parameters:
columns – the column names about to be aggregated.
kinds –
{column: 'numeric' | 'text' | 'unknown'}, asspacr.multi_database.column_kinds()returns it. A column absent from the mapping is'unknown'.overrides – the caller’s explicit choices. A column they set is not a fall-through and appears in neither bucket – they chose it.
- Returns:
{'mean': (...), 'identifier': (...), 'unknown': (...)}:meannumeric, no rule matched, so
spacr.merge_tables.DEFAULT_AGGREGATIONis what it takes.identifiertext. It takes
spacr.merge_tables.TEXT_AGGREGATION, not a mean, and whether it is CARRIED or REFUSED depends on the data – seeambiguous_identifiers(), which can only be answered by reading the rows.unknownthe database declared no type, so neither can be promised. Named separately rather than folded into either: an absent answer that reads as a definite one is the failure this module exists to avoid.
- spacr.plate_measurements.default_aggregated_columns(plan: Mapping[str, str], *, overrides: Mapping[str, str] | None = None) Tuple[str, ...][source]¶
The columns in
planthat got the default because NO rule matched.Recomputed by re-walking
spacr.merge_tables.AGGREGATION_RULES, so this cannot drift out of step with them: a rule added for a measurement tomorrow removes that measurement from this list the same day, and a rule deleted puts it back.A column the caller overrode is not reported however it was set – they chose it. Nor is a text column, which takes
spacr.merge_tables.TEXT_AGGREGATIONbecause text does not add up, not because nobody thought about it.- Parameters:
plan –
{column: aggregation}fromspacr.merge_tables.aggregation_plan().overrides – the caller’s explicit choices, which are excluded.
- spacr.plate_measurements.describe_identifier_refusal(table: str, column: str, detail: Mapping[str, Any]) str[source]¶
One line saying why
columnwas left out, with an example.- Parameters:
table – child table whose identifier could not be carried safely.
column – varying text-identifier column that was omitted.
detail – ambiguity record from
ambiguous_identifiers(), including the affected group count and optional examples.
Written here rather than in the panel so the headless merge and the Qt one say the same sentence.
- spacr.plate_measurements.merge_plate_databases(attachments: Any, tables: Sequence[str] = (), *, anchor: str | None = None, policy: spacr.merge_tables.MergePolicy | None = None, screens: Any = None, columns: str = 'common', report: Callable[[str], None] | None = None, on_ambiguous_identifier: str = 'refuse') PlateMerge[source]¶
Merge every attached plate database into one frame, one row per anchor.
Calls
spacr.multi_database.read_merged()once per chosen table andspacr.merge_tables.roll_up()once per child table. It performs no aggregation of its own, and the join type comes fromspacr.merge_tables.MergePolicy.how_for()per table rather than from a blankethow.- Parameters:
attachments – any shape
plate_databases()accepts. Plates with no database are skipped, not fatal.tables – the object tables to join. The anchor is added if absent.
anchor – the object a row of the result MEANS.
cellby default (spacr.merge_tables.DEFAULT_PRIMARY), overridingpolicy.primarywhen both are given.policy – the aggregation and join policy –
na,overrides,consolidate_on_cell,keep_uninfected.screens – optional screen label per plate: a mapping keyed by plate, or a sequence parallel to the attached databases.
columns –
'common'or'union', passed toread_merged.report – called with one line per thing the merge cost, the table named.
on_ambiguous_identifier –
'refuse'(default) leaves out a text identifier that differs within a roll-up group and names it onreport;'first'restores the old silent pick. Seeambiguous_identifiers()– and note that this is the same default the Measurements tab applies, because two answers to one question is how two merges of one screen come to disagree.
- Returns:
the
PlateMerge.- Raises:
MergeRefused – nothing attached, a database that has moved, one database attached to two plates, an anchor that is not one row per cell, a table missing from some database, or – from
read_mergeditself – a plate id repeated inside one screen.
- spacr.plate_measurements.missing_databases(attachments: Any) Tuple[PlateDatabase, ...][source]¶
Attached databases that are not on disk, named before the run starts.
- Parameters:
attachments – input-table attachment rows in any shape accepted by
plate_databases().
- spacr.plate_measurements.plate_databases(attachments: Any) Tuple[PlateDatabase, ...][source]¶
The plate rows that HAVE a database, in the order the table lists them.
- Parameters:
attachments –
{plate: path}, a sequence ofPlateDatabase, a sequence of(plate, path)pairs, or the input table’s own rows (mappings carryingplateanddatabase).- Returns:
one
PlateDatabaseper attached row. Rows with no database are left out rather than carried as blanks – seeunattached_plates(), which is where they are named.
- spacr.plate_measurements.unattached_plates(attachments: Any) Tuple[str, ...][source]¶
The plates with no database, which is legal and must be said out loud.
- Parameters:
attachments – input-table attachment rows in any shape accepted by
plate_databases().
The regression runs on scores and counts; the database is what makes the Measurements tab possible for that plate. Its absence disables that plate there rather than failing the run – so the plate is listed, and the listing is what stops it looking like an omission.