spacr.multi_database

Load several measurement databases as one frame, without pooling them.

A screen acquired as three plates used to require three sessions in the Gate Editor and three separate UMAPs, and the comparison a user actually wants – all of it in one embedding, one gate, one figure – could not be made inside spaCR at all.

WHY THIS IS ONE MODULE AND NOT TWO SCREENS’ PRIVATE CODE. The hard part of merging databases is not concatenation, it is the three ways it goes wrong quietly, and each of them has exactly one right answer that both screens must give:

  • two plates named plate1 in two files are two plates, and plate1_r1_c1_f1_o1 is one key. Pool them and every per-well number afterwards is computed over two experiments at once, with nothing on screen to say so.

  • databases written by different spaCR versions have different columns, and “just intersect” silently drops measurements the user came to compare.

  • a row that has forgotten which file it came from cannot answer the single most valuable question a merged view can ask, which is whether the clusters are biology or batch.

Two private implementations would disagree about all three, and the third module to need this would make a third set of answers.

THE DESIGN, in one line: describe the merge before performing it, refuse ambiguity rather than resolving it silently, and never lose provenance.

TWO SCREENS, AND A COLLISION THAT IS NOT ONE

There is also a case where the first bullet above is backwards. Two screens that share a guide library both have plate1..plate4, and there they are not two plates claiming one identity – they are two different plates whose identity was only ever partly written down. The missing part is the screen.

So the check is no longer “does this plate id appear twice?” but “does this plate id appear twice INSIDE ONE SCREEN?”. The first is normal and must be silent about being fine while still saying it happened (MergePlan.shared_plates_across_screens); the second is the original error, unchanged and just as fatal.

on_collision='qualify' still exists for callers who want it, but it is no longer the recommended answer for two screens: rewriting plate1 to kd-plate1 makes the keys unique, which is all it was built to do, and leaves the screen un-analysable – you cannot block on it, test for a screen effect, or colour by it without parsing a string back apart.

Exceptions

MergeCancelled

The user stopped the merge before it finished.

MergeRefused

The merge would have silently changed what the numbers mean.

Classes

MergeDecision

One merge, and what was decided about it.

MergePlan

What a merge WOULD do, computed before anything is concatenated.

SourceSummary

What one database contributes to a merge.

Functions

column_kinds(→ Dict[str, str])

{column: 'numeric' | 'text' | 'unknown'} for one table, WITHOUT

decision_for(→ MergeDecision)

Build the record for what plan was asked to do.

decision_log_path(→ str)

Where merge decisions are appended.

describe_merge(→ MergePlan)

Work out what merging paths would produce, WITHOUT reading rows.

normalise_plate_ids(→ pandas.DataFrame)

Collapse a doubled p prefix in every column that carries a plate.

read_merged(→ pandas.DataFrame)

Read table from every path and return one frame.

record_decision(→ str)

Append decision to the merge log and return the file it went to.

source_labels(→ Tuple[str, ...])

One short, unique, human name per database, decided for the whole set.

Module Contents

exception spacr.multi_database.MergeCancelled[source]

Bases: RuntimeError

The user stopped the merge before it finished.

NOT a MergeRefused, and the distinction is the whole reason this is its own class. A refusal is an ANSWER about the data and the caller shows it; a cancellation is the user changing their mind and there is nothing to report about their databases. A caller that caught both as one would put “the merge was refused” in front of somebody who pressed Stop.

Nothing is half-written when this is raised: every merge in this module builds a frame in memory and returns it at the end, so abandoning one leaves the previous result exactly where it was.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.multi_database.MergeRefused[source]

Bases: RuntimeError

The merge would have silently changed what the numbers mean.

Initialize self. See help(type(self)) for accurate signature.

class spacr.multi_database.MergeDecision[source]

One merge, and what was decided about it.

Deliberately flat and JSON-safe: this is written to a log that outlives the session, so it holds strings and numbers rather than objects whose class may not exist by the time somebody reads it back.

Parameters:
  • table – database table targeted by the recorded merge.

  • sources – source database paths in requested merge order.

  • labels – provenance labels corresponding positionally to sources.

  • rows – pre-merge row count keyed by source label.

  • columns – column-selection rule supplied for the merge, normally "common" or "union".

  • dropped_columns – columns omitted by the selected rule; empty for a union merge.

  • colliding_plates – within-screen duplicate plate identifiers mapped to their contributing source labels.

  • outcome – caller-supplied outcome label, such as "merged", "refused", or "resolved".

  • resolution – operator explanation of how a collision or other decision was resolved, or "" when no explanation was recorded.

  • when – ISO-formatted timestamp attached to the decision; decision_for() generates local time to second precision when none is supplied.

as_dict() → Dict[str, Any][source]

The record as plain JSON-safe data.

class spacr.multi_database.MergePlan[source]

What a merge WOULD do, computed before anything is concatenated.

This exists so a user is told what they are about to lose. The column set a merge produces IS the analysis they are about to run, and finding out afterwards that half the measurements were dropped is finding out too late.

Parameters:
  • sources – per-database summaries in requested merge order.

  • common_columns – canonical columns present in every source.

  • partial_columns – non-common columns mapped to the source labels that contain them.

  • colliding_plates – plate identifiers duplicated within one screen, mapped to contributing source labels.

  • colliding_identities – duplicate (screen, plate) identities mapped to contributing source labels.

  • shared_plates_across_screens – plate identifiers reused by distinct screens; reported for review but deliberately not treated as collisions.

describe() → str[source]

A human-readable summary, for a dialog or a log line.

property dropped_columns: Tuple[str, ...][source]

Return measurements discarded by a columns='common' merge.

Identity columns in _ALWAYS_KEPT are excluded because the merge retains them even when only some sources contain them.

property has_collisions: bool[source]

Return whether a plate identity collides within one screen.

property screens: Tuple[str, ...][source]

Every screen this merge would produce, in sorted order.

property screens_were_named: bool[source]

Whether the caller is working in screens at all.

True as soon as ONE database was given a screen label. It is not len(screens) > 1: labelling two databases as the same screen is a deliberate statement (they are two halves of one experiment), and a refusal about them still has to say which screen, or the user cannot tell it apart from the two-screen case that is allowed.

property total_rows: int[source]

Return the sum of pre-merge row counts across all sources.

class spacr.multi_database.SourceSummary[source]

What one database contributes to a merge.

Parameters:
  • path – filesystem path of the source database, preserved in the form supplied to describe_merge().

  • label – short provenance label assigned uniquely across the requested source set.

  • table – database table inspected for the merge.

  • rows – number of rows in the inspected source table before any per-source preview limit.

  • columns – stored source-table column names in database order.

  • plates – sorted canonical plate identifiers contributed by the source.

  • screen – caller-supplied screen override, or None when a stored screenID or spacr.schema.DEFAULT_SCREEN determines it.

  • screen_plates – sorted distinct canonical (screenID, plateID) identities contributed after applying the screen override or stored screen values.

  • stored_plates – sorted distinct plate identifiers exactly as stored, retained so non-canonical spellings can be reported.

property name: str[source]

Short name for a legend or a chip.

property odd_plates: Tuple[str, ...][source]

Stored plate ids whose spelling is not the canonical one.

Empty in the normal case, which is why a caller can print a line per entry and stay silent when there is nothing to say.

property screens: Tuple[str, ...][source]

Distinct screens this source contributes, in sorted order.

spacr.multi_database.column_kinds(path: str, table: str) → Dict[str, str][source]

{column: 'numeric' | 'text' | 'unknown'} for one table, WITHOUT reading a row.

WHY THIS EXISTS, and it is a correctness answer rather than a convenience. A pre-merge plan has to say how each column will be combined, and the merge decides that from the column’s pandas dtype (spacr.merge_tables.aggregation_plan() asks is_numeric_dtype). A plan that matched only on the column NAME told users that file_name and path_name “would take the default (mean)”, which is not what happens and is not a thing that can happen to a string. The declared affinity is what predicts the dtype, and reading it costs one PRAGMA.

'unknown' is returned rather than guessed for a column declared with no type or as a BLOB – an absent answer that reads as a definite one is the failure this module exists to avoid.

Parameters:
  • path – the database.

  • table – the table.

Returns:

one entry per column, in the table’s own column order.

spacr.multi_database.decision_for(plan: MergePlan, *, outcome: str, columns: str = 'common', resolution: str = '', when: str | None = None) → MergeDecision[source]

Build the record for what plan was asked to do.

Parameters:
  • plan – the plan the decision is about.

  • outcome – caller-supplied outcome label, such as "merged", "refused", or "resolved".

  • columns – the column rule that was used.

  • resolution – what the user chose, in words.

  • when – ISO timestamp; None takes the current local time.

spacr.multi_database.decision_log_path() → str[source]

Where merge decisions are appended.

Beside ~/.spacr/runs, which is where this application already keeps the record of what it was asked to do.

spacr.multi_database.describe_merge(paths: Sequence[str], table: str, *, screens: Any = None) → MergePlan[source]

Work out what merging paths would produce, WITHOUT reading rows.

Reads only sqlite metadata and the distinct plate ids, so this is cheap enough to run while the user is still choosing files – which is the point, because the answer has to arrive before they commit.

Parameters:
  • paths – measurement databases, in the order the user added them.

  • table – the table to merge, e.g. 'cell'.

  • screens – optional screen label per database – a sequence parallel to paths, or a mapping from path. Two databases in different screens may share a plate id; two in the same screen may not.

Returns:

the MergePlan.

spacr.multi_database.normalise_plate_ids(frame: pandas.DataFrame) → pandas.DataFrame[source]

Collapse a doubled p prefix in every column that carries a plate.

WHY THIS IS NOT COSMETIC. A measurements database stamped pplate1 produces merged rows stamped pplate1, while the score and count CSVs have already been normalised to plate1. The two then do not meet. Every join INSIDE the merge is unaffected, because both sides read the same stored value – which is exactly what makes it hard to see: the merge succeeds, the row counts are right, and the failure appears later and somewhere else, as a gene half that is missing for no visible reason.

THE COMPOSED KEYS MATTER AS MUCH AS THE PLATE COLUMN. prc is <plate>_<row>_<column>, so a doubled prefix rides in its first component; rewriting plateID alone would leave prc unjoinable and the two columns disagreeing about the same plate.

Applied on READ, so nothing on disk is rewritten and an old database keeps working – the standing rule is to correct the format going forward and migrate the content, and a measurements database is the user’s data rather than ours to edit.

A thin name over spacr.schema.normalise_plate_columns(), which is the one implementation.

Parameters:

frame – any frame read from a measurements database.

Returns:

the same frame, with the plate-bearing columns normalised in place. Columns it does not have are skipped.

spacr.multi_database.read_merged(paths: Sequence[str], table: str, *, plan: MergePlan | None = None, columns: str = 'common', on_collision: str = 'refuse', screens: Any = None, report: Callable[[str], None] | None = None, limit_per_source: int | None = None, progress: Callable[[str, int, int], None] | None = None, cancelled: Callable[[], bool] | None = None, rows_done: int = 0, rows_total: int | None = None) → pandas.DataFrame[source]

Read table from every path and return one frame.

Parameters:
  • paths – measurement databases.

  • table – the table to read from each.

  • plan – a plan from describe_merge(); recomputed if omitted. Pass screens= to describe_merge() and here alike, or the plan and the read disagree about which screen a database is.

  • columns – 'common' keeps only columns present in every source (safe, and drops); 'union' keeps everything and leaves nulls where a source did not have it (keeps, and changes what “missing” means). 'common' is the default because measurement tables are wide and differ between spaCR versions – but a dropped measurement is a measurement the user came to compare, so the set is reported rather than merely defaulted (see report and .attrs).

  • on_collision – what to do when a plate id is duplicated within one screen. 'refuse' raises MergeRefused; 'qualify' prefixes each colliding plate with its source label. There is deliberately no option that pools them. Two different screens sharing a plate id is not a collision and reaches neither branch.

  • screens – optional screen label per database – a sequence parallel to paths, or a mapping from path. This is the recommended answer for two screens: the label is written into SCREEN_COLUMN, so it stays a dimension you can block on, rather than into the plate id, where it becomes a string to be parsed back apart.

  • report – called with one human-readable line per thing the merge cost – currently the dropped measurements.

  • limit_per_source – row cap per database, for previews.

  • progress – called progress(stage, done, total) before and after each database is read. stage is a sentence naming the table and the database it is on; done and total are ROWS, counted against rows_total so a caller reading several tables can show one bar across all of them. Called from whatever thread this runs on, so a GUI caller must relay it rather than touch a widget in it.

  • cancelled – called before each database; a true answer raises MergeCancelled and nothing is returned. Checked between sources rather than inside one, because a half-read source is not a thing this function can hand back.

  • rows_done – rows already counted by an earlier call, for the multi-table case.

  • rows_total – the denominator for progress. Defaults to this plan’s own total, which is right for a single table and too small for a caller merging several – so a caller merging several passes the grand total it computed from all their plans.

Returns:

one frame carrying SOURCE_COLUMN and SCREEN_COLUMN, with frame.attrs['dropped_columns'] naming the measurements that did not survive.

Raises:
  • MergeRefused – on a within-screen plate collision under on_collision='refuse', or an unknown option.

  • MergeCancelled – when cancelled() answered true.

spacr.multi_database.record_decision(decision: MergeDecision, path: str | None = None) → str[source]

Append decision to the merge log and return the file it went to.

JSON lines, appended: a merge decision is an event, and rewriting a whole document to add one would lose the others if two screens decided at once.

Never raises for an unwritable log – a read-only home directory must not take a screen down for the sake of an audit line – but returns "" so a caller that wants to say the record was not kept can.

Parameters:

decision – JSON-safe merge decision to append.

spacr.multi_database.source_labels(paths: Sequence[str]) → Tuple[str, ...][source]

One short, unique, human name per database, decided for the whole set.

Public because a chip, a legend and the SOURCE_COLUMN value have to be the SAME string: a screen that labels a chip plate1 while the provenance column says measurements (2) has provenance the user cannot follow.

Decided across all the paths at once rather than one at a time, because “is this name ambiguous?” is a question about the set. Three rules, in order:

  1. the file stems, when they differ – that is what the user called them;

  2. the nearest MEANINGFUL parent folder (_meaningful_parent()), which is the plate folder in spaCR’s own <plate>/measurements/ measurements.db layout;

  3. the historical parent/stem rule with a numeric tail, for the case where even the folders repeat.

Parameters:

paths – database paths to label together in their given order.