spacr.multi_database¶
Load several measurement databases as one frame, without pooling them.
A screen acquired as three plates used to require three sessions in the Gate Editor and three separate UMAPs, and the comparison a user actually wants – all of it in one embedding, one gate, one figure – could not be made inside spaCR at all.
WHY THIS IS ONE MODULE AND NOT TWO SCREENS’ PRIVATE CODE. The hard part of merging databases is not concatenation, it is the three ways it goes wrong quietly, and each of them has exactly one right answer that both screens must give:
two plates named
plate1in two files are two plates, andplate1_r1_c1_f1_o1is one key. Pool them and every per-well number afterwards is computed over two experiments at once, with nothing on screen to say so.databases written by different spaCR versions have different columns, and “just intersect” silently drops measurements the user came to compare.
a row that has forgotten which file it came from cannot answer the single most valuable question a merged view can ask, which is whether the clusters are biology or batch.
Two private implementations would disagree about all three, and the third module to need this would make a third set of answers.
THE DESIGN, in one line: describe the merge before performing it, refuse ambiguity rather than resolving it silently, and never lose provenance.
TWO SCREENS, AND A COLLISION THAT IS NOT ONE¶
There is also a case where the first bullet above is backwards. Two
screens that share a guide library both have plate1..plate4, and there
they are not two plates claiming one identity – they are two different
plates whose identity was only ever partly written down. The missing part is
the screen.
So the check is no longer “does this plate id appear twice?” but “does this
plate id appear twice INSIDE ONE SCREEN?”. The first is normal and must be
silent about being fine while still saying it happened
(MergePlan.shared_plates_across_screens); the second is the original
error, unchanged and just as fatal.
on_collision='qualify' still exists for callers who want it, but it is no
longer the recommended answer for two screens: rewriting plate1 to
kd-plate1 makes the keys unique, which is all it was built to do, and
leaves the screen un-analysable – you cannot block on it, test for a screen
effect, or colour by it without parsing a string back apart.
Exceptions¶
The user stopped the merge before it finished. |
|
The merge would have silently changed what the numbers mean. |
Classes¶
One merge, and what was decided about it. |
|
What a merge WOULD do, computed before anything is concatenated. |
|
What one database contributes to a merge. |
Functions¶
|
|
|
Build the record for what |
|
Where merge decisions are appended. |
|
Work out what merging |
|
Collapse a doubled |
|
Read |
|
Append |
|
One short, unique, human name per database, decided for the whole set. |
Module Contents¶
- exception spacr.multi_database.MergeCancelled[source]¶
Bases:
RuntimeErrorThe user stopped the merge before it finished.
NOT a
MergeRefused, and the distinction is the whole reason this is its own class. A refusal is an ANSWER about the data and the caller shows it; a cancellation is the user changing their mind and there is nothing to report about their databases. A caller that caught both as one would put “the merge was refused” in front of somebody who pressed Stop.Nothing is half-written when this is raised: every merge in this module builds a frame in memory and returns it at the end, so abandoning one leaves the previous result exactly where it was.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.multi_database.MergeRefused[source]¶
Bases:
RuntimeErrorThe merge would have silently changed what the numbers mean.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.multi_database.MergeDecision[source]¶
One merge, and what was decided about it.
Deliberately flat and JSON-safe: this is written to a log that outlives the session, so it holds strings and numbers rather than objects whose class may not exist by the time somebody reads it back.
- Parameters:
table – database table targeted by the recorded merge.
sources – source database paths in requested merge order.
labels – provenance labels corresponding positionally to
sources.rows – pre-merge row count keyed by source label.
columns – column-selection rule supplied for the merge, normally
"common"or"union".dropped_columns – columns omitted by the selected rule; empty for a union merge.
colliding_plates – within-screen duplicate plate identifiers mapped to their contributing source labels.
outcome – caller-supplied outcome label, such as
"merged","refused", or"resolved".resolution – operator explanation of how a collision or other decision was resolved, or
""when no explanation was recorded.when – ISO-formatted timestamp attached to the decision;
decision_for()generates local time to second precision when none is supplied.
- class spacr.multi_database.MergePlan[source]¶
What a merge WOULD do, computed before anything is concatenated.
This exists so a user is told what they are about to lose. The column set a merge produces IS the analysis they are about to run, and finding out afterwards that half the measurements were dropped is finding out too late.
- Parameters:
sources – per-database summaries in requested merge order.
common_columns – canonical columns present in every source.
partial_columns – non-common columns mapped to the source labels that contain them.
colliding_plates – plate identifiers duplicated within one screen, mapped to contributing source labels.
colliding_identities – duplicate
(screen, plate)identities mapped to contributing source labels.shared_plates_across_screens – plate identifiers reused by distinct screens; reported for review but deliberately not treated as collisions.
- property dropped_columns: Tuple[str, ...][source]¶
Return measurements discarded by a
columns='common'merge.Identity columns in
_ALWAYS_KEPTare excluded because the merge retains them even when only some sources contain them.
- property screens_were_named: bool[source]¶
Whether the caller is working in screens at all.
True as soon as ONE database was given a screen label. It is not
len(screens) > 1: labelling two databases as the same screen is a deliberate statement (they are two halves of one experiment), and a refusal about them still has to say which screen, or the user cannot tell it apart from the two-screen case that is allowed.
- class spacr.multi_database.SourceSummary[source]¶
What one database contributes to a merge.
- Parameters:
path – filesystem path of the source database, preserved in the form supplied to
describe_merge().label – short provenance label assigned uniquely across the requested source set.
table – database table inspected for the merge.
rows – number of rows in the inspected source table before any per-source preview limit.
columns – stored source-table column names in database order.
plates – sorted canonical plate identifiers contributed by the source.
screen – caller-supplied screen override, or
Nonewhen a storedscreenIDorspacr.schema.DEFAULT_SCREENdetermines it.screen_plates – sorted distinct canonical
(screenID, plateID)identities contributed after applying the screen override or stored screen values.stored_plates – sorted distinct plate identifiers exactly as stored, retained so non-canonical spellings can be reported.
- spacr.multi_database.column_kinds(path: str, table: str) Dict[str, str][source]¶
{column: 'numeric' | 'text' | 'unknown'}for one table, WITHOUT reading a row.WHY THIS EXISTS, and it is a correctness answer rather than a convenience. A pre-merge plan has to say how each column will be combined, and the merge decides that from the column’s pandas dtype (
spacr.merge_tables.aggregation_plan()asksis_numeric_dtype). A plan that matched only on the column NAME told users thatfile_nameandpath_name“would take the default (mean)”, which is not what happens and is not a thing that can happen to a string. The declared affinity is what predicts the dtype, and reading it costs onePRAGMA.'unknown'is returned rather than guessed for a column declared with no type or as a BLOB – an absent answer that reads as a definite one is the failure this module exists to avoid.- Parameters:
path – the database.
table – the table.
- Returns:
one entry per column, in the table’s own column order.
- spacr.multi_database.decision_for(plan: MergePlan, *, outcome: str, columns: str = 'common', resolution: str = '', when: str | None = None) MergeDecision[source]¶
Build the record for what
planwas asked to do.- Parameters:
plan – the plan the decision is about.
outcome – caller-supplied outcome label, such as
"merged","refused", or"resolved".columns – the column rule that was used.
resolution – what the user chose, in words.
when – ISO timestamp;
Nonetakes the current local time.
- spacr.multi_database.decision_log_path() str[source]¶
Where merge decisions are appended.
Beside
~/.spacr/runs, which is where this application already keeps the record of what it was asked to do.
- spacr.multi_database.describe_merge(paths: Sequence[str], table: str, *, screens: Any = None) MergePlan[source]¶
Work out what merging
pathswould produce, WITHOUT reading rows.Reads only sqlite metadata and the distinct plate ids, so this is cheap enough to run while the user is still choosing files – which is the point, because the answer has to arrive before they commit.
- Parameters:
paths – measurement databases, in the order the user added them.
table – the table to merge, e.g.
'cell'.screens – optional screen label per database – a sequence parallel to
paths, or a mapping from path. Two databases in different screens may share a plate id; two in the same screen may not.
- Returns:
the
MergePlan.
- spacr.multi_database.normalise_plate_ids(frame: pandas.DataFrame) pandas.DataFrame[source]¶
Collapse a doubled
pprefix in every column that carries a plate.WHY THIS IS NOT COSMETIC. A measurements database stamped
pplate1produces merged rows stampedpplate1, while the score and count CSVs have already been normalised toplate1. The two then do not meet. Every join INSIDE the merge is unaffected, because both sides read the same stored value – which is exactly what makes it hard to see: the merge succeeds, the row counts are right, and the failure appears later and somewhere else, as a gene half that is missing for no visible reason.THE COMPOSED KEYS MATTER AS MUCH AS THE PLATE COLUMN.
prcis<plate>_<row>_<column>, so a doubled prefix rides in its first component; rewritingplateIDalone would leaveprcunjoinable and the two columns disagreeing about the same plate.Applied on READ, so nothing on disk is rewritten and an old database keeps working – the standing rule is to correct the format going forward and migrate the content, and a measurements database is the user’s data rather than ours to edit.
A thin name over
spacr.schema.normalise_plate_columns(), which is the one implementation.- Parameters:
frame – any frame read from a measurements database.
- Returns:
the same frame, with the plate-bearing columns normalised in place. Columns it does not have are skipped.
- spacr.multi_database.read_merged(paths: Sequence[str], table: str, *, plan: MergePlan | None = None, columns: str = 'common', on_collision: str = 'refuse', screens: Any = None, report: Callable[[str], None] | None = None, limit_per_source: int | None = None, progress: Callable[[str, int, int], None] | None = None, cancelled: Callable[[], bool] | None = None, rows_done: int = 0, rows_total: int | None = None) pandas.DataFrame[source]¶
Read
tablefrom every path and return one frame.- Parameters:
paths – measurement databases.
table – the table to read from each.
plan – a plan from
describe_merge(); recomputed if omitted. Passscreens=todescribe_merge()and here alike, or the plan and the read disagree about which screen a database is.columns –
'common'keeps only columns present in every source (safe, and drops);'union'keeps everything and leaves nulls where a source did not have it (keeps, and changes what “missing” means).'common'is the default because measurement tables are wide and differ between spaCR versions – but a dropped measurement is a measurement the user came to compare, so the set is reported rather than merely defaulted (seereportand.attrs).on_collision – what to do when a plate id is duplicated within one screen.
'refuse'raisesMergeRefused;'qualify'prefixes each colliding plate with its source label. There is deliberately no option that pools them. Two different screens sharing a plate id is not a collision and reaches neither branch.screens – optional screen label per database – a sequence parallel to
paths, or a mapping from path. This is the recommended answer for two screens: the label is written intoSCREEN_COLUMN, so it stays a dimension you can block on, rather than into the plate id, where it becomes a string to be parsed back apart.report – called with one human-readable line per thing the merge cost – currently the dropped measurements.
limit_per_source – row cap per database, for previews.
progress – called
progress(stage, done, total)before and after each database is read.stageis a sentence naming the table and the database it is on;doneandtotalare ROWS, counted againstrows_totalso a caller reading several tables can show one bar across all of them. Called from whatever thread this runs on, so a GUI caller must relay it rather than touch a widget in it.cancelled – called before each database; a true answer raises
MergeCancelledand nothing is returned. Checked between sources rather than inside one, because a half-read source is not a thing this function can hand back.rows_done – rows already counted by an earlier call, for the multi-table case.
rows_total – the denominator for
progress. Defaults to this plan’s own total, which is right for a single table and too small for a caller merging several – so a caller merging several passes the grand total it computed from all their plans.
- Returns:
one frame carrying
SOURCE_COLUMNandSCREEN_COLUMN, withframe.attrs['dropped_columns']naming the measurements that did not survive.- Raises:
MergeRefused – on a within-screen plate collision under
on_collision='refuse', or an unknown option.MergeCancelled – when
cancelled()answered true.
- spacr.multi_database.record_decision(decision: MergeDecision, path: str | None = None) str[source]¶
Append
decisionto the merge log and return the file it went to.JSON lines, appended: a merge decision is an event, and rewriting a whole document to add one would lose the others if two screens decided at once.
Never raises for an unwritable log – a read-only home directory must not take a screen down for the sake of an audit line – but returns
""so a caller that wants to say the record was not kept can.- Parameters:
decision – JSON-safe merge decision to append.
- spacr.multi_database.source_labels(paths: Sequence[str]) Tuple[str, ...][source]¶
One short, unique, human name per database, decided for the whole set.
Public because a chip, a legend and the
SOURCE_COLUMNvalue have to be the SAME string: a screen that labels a chipplate1while the provenance column saysmeasurements (2)has provenance the user cannot follow.Decided across all the paths at once rather than one at a time, because “is this name ambiguous?” is a question about the set. Three rules, in order:
the file stems, when they differ – that is what the user called them;
the nearest MEANINGFUL parent folder (
_meaningful_parent()), which is the plate folder in spaCR’s own<plate>/measurements/ measurements.dblayout;the historical parent/stem rule with a numeric tail, for the case where even the folders repeat.
- Parameters:
paths – database paths to label together in their given order.