spacr.merge_tables¶
Merging object tables, aggregating each measurement by what it MEASURES.
A cell with four pathogens in it has four rows in pathogen and one in
cell. Putting a pathogen measurement on the same axis as a cell
measurement means rolling those four up into one number – and which number
depends entirely on what is being measured:
area, perimeter, integrated intensity, counts -> SUM minimum intensity -> MIN of the four maximum intensity -> MAX of the four mean, median -> the object mean/median shape descriptors, positions -> mean text -> the first
spacr.io._read_and_join_tables already does this join, and aggregates
every numeric column with mean. That silently answers a different question
per column: four pathogens’ total area becomes an average area, a count
becomes an average count, and a MINIMUM becomes a mean of minima, which is not
a minimum of anything. The join is otherwise sound – it is where the parent
link, the timelapse key and the cardinality checks live – so this module
changes what the aggregation is, not how the tables find each other.
Naming. A merged column carries its table: area from nucleus
becomes nucleus_area. Prefix rather than suffix so every column of one
object sorts together in the axis picker, which is how anyone looks for them.
The primary object. Everything is rolled up onto ONE table – the cell by default. That choice decides what a row means, so it is a setting rather than an assumption: rolling cells onto pathogens is a legitimate thing to want and gives a different table.
Exceptions¶
One column, two tables, two different values for the same object. |
|
A merge that cannot be done, and why. |
|
A reduction that cannot be computed, and why. |
Classes¶
How a merge is performed. Every field is a user-facing setting. |
Functions¶
|
How |
|
Resolve global rules and table-qualified per-column overrides. |
|
The aggregation chosen for every column -- what the user gets shown. |
|
Calculate each column group's share of variance in the input matrix. |
|
One table with every chosen object's measurements on it. |
|
The object tables in this database, in preference order. |
|
Does "was this object measured" predict where it landed? |
|
Object identifiers as integers, whatever spelling they arrived in. |
|
Collapse |
|
Reduce many measurements to a few, for gating in xD. |
|
Aggregate |
|
Every table in the database, excluding SQLite's own internals. |
Module Contents¶
- exception spacr.merge_tables.ColumnConflict[source]¶
Bases:
MergeErrorOne column, two tables, two different values for the same object.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.merge_tables.MergeError[source]¶
Bases:
ValueErrorA merge that cannot be done, and why.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.merge_tables.ReductionError[source]¶
Bases:
ValueErrorA reduction that cannot be computed, and why.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.merge_tables.MergePolicy[source]¶
How a merge is performed. Every field is a user-facing setting.
- Parameters:
primary – the object everything is rolled up onto. Decides what a row of the merged table MEANS.
na – what happens to a primary object with no children –
keepleaves NaN,zerofills with 0,dropremoves the row. Not interchangeable: a cell with no pathogens genuinely has a pathogen COUNT of zero, and genuinely has no pathogen mean intensity at all.overrides – column -> aggregation, beating the rules. The rules are right most of the time, and a default that is right most of the time is a wrong answer nobody can find the rest of it.
consolidate_on_cell – whether many-per-cell child tables restrict output to cells that contributed a child.
keep_uninfected – preserve cells without pathogens or organelles as the uninfected control population even while consolidating.
- how_for(table: str) str[source]¶
Whether
tablekeeps cells it contributed no rows for.- Parameters:
table – child object table whose join mode is requested.
A cell has exactly one cytoplasm and may have multiple nuclei, pathogens, or organelles. The many-per-cell tables are rolled up to one row per cell first (see
roll_up()), andconsolidate_on_celldecides what happens to a cell the roll-up found nothing for:- consolidate_on_cell=True the analysis is about cells that HAVE
the child, so the join is inner
- consolidate_on_cell=False keep the cell and leave the child’s
columns NA
Pathogen and organelle tables follow
keep_uninfectedso the default retains uninfected control cells. Set it toFalseto restrict the merged table to infected cells.A one-row-per-cell table keeps whatever
object_roles.JOIN_HOWdeclares, because there is no consolidation to decide about.
- spacr.merge_tables.aggregation_for(column: str, *, numeric: bool = True, overrides: Mapping[str, str] | None = None) str[source]¶
How
columncombines when several children roll up into one parent.- Parameters:
column – measurement column name matched against the ordered
AGGREGATION_RULES.numeric – text columns take the first value whatever their name.
overrides – explicit choices, which always win.
- Returns:
one of
AGGREGATIONS.
- spacr.merge_tables.aggregation_overrides(policy: MergePolicy, table: str) Dict[str, str][source]¶
Resolve global rules and table-qualified per-column overrides.
- Parameters:
policy – Shared merge policy.
table – Source table whose aggregation rules are requested.
- Returns:
Unqualified column-to-method mapping for this table.
- spacr.merge_tables.aggregation_plan(frame: pandas.DataFrame, *, overrides: Mapping[str, str] | None = None, skip: Sequence[str] = ()) Dict[str, str][source]¶
The aggregation chosen for every column – what the user gets shown.
- Parameters:
frame – child-object table whose columns will be rolled up.
Returned rather than applied silently so the settings panel can display it and the user can override any of it.
Calculate each column group’s share of variance in the input matrix.
The matrix is prepared with the same missing-value and scaling procedure used by
reduce_dimensions(). The result therefore characterizes the inputs supplied to PCA, UMAP, t-SNE, and other reducers without requiring method-specific loadings.- Parameters:
frame – Measurement frame containing the candidate feature columns.
groups –
{group_name: columns}. A column named by two groups is counted in both; shares may therefore sum to more than one, and this condition is recorded inresult.attrs['overlapping'].scale – Whether to standardize features before calculating variance, matching the corresponding reducer option.
min_coverage – Minimum non-missing fraction required for a feature to enter the prepared matrix.
- Returns:
DataFrame indexed by group with
shareandcolumnsfields, sorted by decreasing share.
- spacr.merge_tables.merge_tables(db_path: str, tables: Sequence[str], *, policy: MergePolicy | None = None) pandas.DataFrame[source]¶
One table with every chosen object’s measurements on it.
This is what makes “a cell measurement on one axis, nuclear on another and pathogen on a third” possible: each table’s columns arrive prefixed with the object they measure, so they can be told apart and picked separately.
- Parameters:
db_path – path to the SQLite measurements database.
tables – which object tables to include. The primary must be one of them, and is added if it is not.
policy – how to aggregate and what to do with childless parents.
- Returns:
one row per primary object.
- Raises:
MergeError – the primary table is not in the database, or a child cannot be linked to it.
- spacr.merge_tables.mergeable_tables(db_path: str) Tuple[str, ...][source]¶
The object tables in this database, in preference order.
- Parameters:
db_path – path to the SQLite database to inspect.
- spacr.merge_tables.missingness_leak(components: pandas.DataFrame, frame: pandas.DataFrame, columns: Sequence[str], *, min_objects: int = 30) pandas.DataFrame[source]¶
Does “was this object measured” predict where it landed?
For each column, the objects that HAVE a value and the objects that do not are compared by the distance between their centroids in the projection, expressed in map radii so it is comparable across runs and across methods.
A gap near 1 means the projection has separated the two groups about as far as the map is wide – on the fact of measurement, not on a measurement. In spaCR that is usually infected against uninfected, and it is exactly the kind of split a user would otherwise write up.
- Parameters:
components – the reducer’s output, indexed like
frame.frame – original measurement table used to determine which component rows had or lacked each input measurement.
columns – the columns that went into the projection.
min_objects – skip a column unless both sides have at least this many objects. A gap computed from four objects is noise, and reporting it would bury the real ones.
- Returns:
one row per checked column, worst
severityfirst. Empty – WITH ITS COLUMNS – when nothing was checkable, so a caller can sort it without a KeyError.
TWO ARTEFACTS, NOT ONE, and this is where spaCR differs from the tool the idea came from. Which one appears depends on how the gap was filled:
centroid_gapthe missing objects sit SOMEWHERE ELSE. Near 1 means the projection has moved them about as far as the map is wide.
dispersion_ratiothe missing objects COLLAPSE.
reduce_dimensionsfills with the column median, so every uninfected cell gets the SAME value on every pathogen column and they land on one point. Near 0 means they have no spread of their own.
Measured on a synthetic infected/uninfected table with twelve pathogen columns, the median fill produced a centroid gap of 0.06 – almost nothing – and a dispersion ratio of 0.11. THE CENTROID STATISTIC ALONE WOULD HAVE MISSED IT, because a median fill puts the missing objects in the middle of the present ones rather than away from them. Both are reported, and
severityis whichever is worse.
- spacr.merge_tables.object_keys(values: pandas.Series) pandas.Series[source]¶
Object identifiers as integers, whatever spelling they arrived in.
- Parameters:
values – object-label series to coerce to nullable integers.
The object key is an integer in every object table and TEXT in
png_list–'o5'– so merging the two raisedyou are trying to merge on int64 and object columns for key object_label
which names the dtypes and not the tables, and stopped the whole merge. The
'o5'form is translated by the one function that already knows every way it goes wrong ('omulti','onone','error', NULL); plain numeric text is converted directly. Anything left becomes NA, so those rows do not match rather than causing the merge to fail.
- spacr.merge_tables.reconcile_duplicates(frame: pandas.DataFrame, suffix: str, *, key: str = 'prcfo', left_name: str = 'the primary table', right_name: str = 'the joined table', on_conflict: str = 'warn') pandas.DataFrame[source]¶
Collapse
col/col+suffixpairs that agree; report those that do not.Joining two measurement tables gives every shared column twice –
plateIDandplateID_cytoplasmhold the same plate written by two stages of the same run. Carrying both doubles the width of the frame and invites a downstream reader to pick the wrong one.So the pair is COMPARED, object by object, rather than assumed: identical columns collapse to one, and a column that disagrees means the two tables describe different objects under the same identity, which is a defect in the data no analysis should quietly average over.
- Parameters:
frame – the merged frame, modified only by dropping columns.
suffix – what the merge appended to the right-hand duplicates.
key – the identity the comparison is reported against.
on_conflict –
warnkeeps the left-hand column and prints;raisestops withColumnConflict.
- Returns:
the frame with agreeing duplicates dropped.
- Raises:
ColumnConflict – a pair disagrees and
on_conflict='raise'.
- spacr.merge_tables.reduce_dimensions(frame: pandas.DataFrame, columns: Sequence[str], *, method: str = 'pca', components: int = 2, scale: bool = True, min_coverage: float = 0.5, seed: int = 0, n_neighbors: int = 15, min_dist: float = 0.1, perplexity: float = 30.0) pandas.DataFrame[source]¶
Reduce many measurements to a few, for gating in xD.
Gating in more dimensions than can be drawn means drawing something else: a projection. The components come back as ORDINARY COLUMNS (
PC1,PC2, …), so every existing gate tool works on them unchanged – a gate on PC1 vs PC2 is the same kind of object as a gate on area vs intensity, and saves, re-applies and exports identically.- Parameters:
frame – object-by-measurement table to project. The returned frame is reindexed to this table’s complete index.
columns – the measurements to reduce. At least two.
method –
pcaalways available; umap and t-SNE if installed.scale – standardise first. Without it a measurement whose numbers are larger dominates every component regardless of what it means.
min_coverage – a column with fewer than this fraction of real values is left out. What remains is median-filled rather than row-dropped – see the comment in the body, which is the difference between xD working on a real table and returning nothing at all.
n_neighbors – UMAP only. How much of the data each point is placed against: small values keep local structure and fragment the map, large ones preserve the global shape and merge populations. Ignored by PCA and t-SNE, which have no such parameter – hence the greying in the xD tab rather than a control that silently does nothing.
min_dist – UMAP only. How tightly points may pack. Ignored elsewhere.
perplexity – t-SNE only, and CLAMPED to
(n - 1) / 3: sklearn raises outright when it exceeds the sample size, which would turn a legitimate setting into a failed projection on a small selection.
- Returns:
a frame of components, indexed like
frame.- Raises:
ReductionError – too few columns, too few rows, nothing numeric, or a method whose package is not installed.
- spacr.merge_tables.roll_up(child: pandas.DataFrame, keys: Sequence[str], *, name: str, policy: MergePolicy) pandas.DataFrame[source]¶
Aggregate
childonto its parent, one rule per column.- Parameters:
child – child-object rows to group and aggregate.
keys – the parent’s identity in the child – the identity columns plus the parent link.
name – the child table’s name, used to prefix its columns.
policy – merge policy supplying per-column aggregation overrides.
- Returns:
one row per parent, columns prefixed with
name.- Raises:
MergeError – the child has none of the keys.
An unmeasured group is missing, never a measured zero. Regression and both interactive merge consumers use this same rule.