spacr.lineage

V9 B20 — cell → nucleus → pathogen, as the tree it already is.

Every object table in measurements.db is flat, and the relationships between them are already stored: spacr.schema gives nucleus and pathogen a cell_id pointing at the cell they sit inside (spacr.schema.CHILD_OBJECT_TABLES). Nothing has ever shown that. “Cell 41 in field 3 has one nucleus and four pathogens, and one of those pathogens has an area that would be impossible inside that nucleus” is a question the database can answer and the GUI could not ask.

This module is the answer, in plain pandas with no Qt, so the tree can be built in a notebook and tested without a display.

Why a tree and not a join

A join answers “give me every pathogen with its cell’s area”. A tree answers “what is inside this cell”, which is a different question and the one a person looking at a crop actually has. The difference shows up in the failures: a join silently drops a cell with no pathogens and silently drops a pathogen whose cell_id names a cell that is not there. Both of those are findings — the first is the negative control working, the second is a segmentation bug — so build_forest() keeps childless parents and orphans() returns the unattached children rather than discarding them.

Identity is the shared one

Nodes are keyed by spacr.selection.object_keys(), the same string the UMAP, the plate view and the crop grid use, so selecting a node in the tree publishes something every other view already understands. A child’s parent is resolved within its own field: cell_id is a label, not a key, and label 7 exists in every field on the plate.

The shared key carries the object type — this module is why

spacr.selection.OBJECT_KEY_COLUMNS used to be the field plus the object label, with no table in it, so a nucleus labelled 1 and a pathogen labelled 1 in the same field had the same key. A lineage tree is where that became visible, because a cell’s own children are exactly the objects most likely to collide: four objects opened as three crops and which one you got depended on the row order of png_list.

spacr.selection.object_keys() now writes the object type into the key, so a node’s shared key already says which table it came from and LineageNode.node_id is the same identity in a different spelling. Both are kept. node_id is what the tree addresses its rows by and cannot collide by construction; LineageNode.key_collisions() compares the two and is now a regression test rather than a warning — it finds nothing on a correctly keyed forest, and build_forest() still takes typed=False so the collapse can be reproduced on demand rather than only remembered.

Exceptions

LineageError

A set of tables that cannot be assembled into a lineage.

Classes

LineageNode

One object and everything inside it.

Functions

build_forest(→ Tuple[LineageNode, ...])

Assemble object tables into one tree per root object.

child_tables(→ Tuple[str, ...])

The tables whose rows carry a parent link, from the schema itself.

describe_forest(→ str)

The shape of a whole forest in words.

field_key(→ str)

The prcf of one row — the field a label is unique within.

forest_key_collisions(→ Dict[str, Tuple[str, ...]])

Every shared key in forest that names more than one object.

lineage_frame(→ pandas.DataFrame)

The forest flattened: one row per node, with its parent and depth.

node_key(→ str)

The shared object key of one row: field, then the typed object id.

orphans(→ pandas.DataFrame)

Child rows whose parent link names no row in the root table.

read_object_tables(→ Dict[str, pandas.DataFrame])

Read the object tables a lineage needs. Safe on a worker thread.

tree_for(→ Optional[LineageNode])

The tree containing key, or None.

Module Contents

exception spacr.lineage.LineageError[source]

Bases: ValueError

A set of tables that cannot be assembled into a lineage.

Raised rather than returning an empty forest: “this cell has no children” and “the cell table has no object_label column so nothing could be matched” render identically as a leaf node, and only one of them is a result.

Initialize self. See help(type(self)) for accurate signature.

class spacr.lineage.LineageNode[source]

One object and everything inside it.

Parameters:
  • key – shared typed object key used by spaCR’s other views and selections.

  • table – object-table name from which the source row came.

  • label – integer object label within its field.

  • field – prcf field identifier containing the object.

  • children – contained objects, grouped in LINEAGE_TABLES order and then by label for deterministic trees.

  • row – copied source row retained so views can display measurements without reopening the frame.

counts() → Dict[str, int][source]

How many of each table are inside this node (itself included).

descendants() → Tuple[LineageNode, ...][source]

Everything below this node, depth-first, excluding itself.

describe() → str[source]

One line: what this is, and what is inside it.

find(key: str) → LineageNode | None[source]

The node with this key, anywhere below (or at) this one.

Parameters:

key – shared object key to look for, compared as a string with each node’s key in depth-first order; the first match wins.

key_collisions() → Dict[str, Tuple[str, ...]][source]

{shared key: the tables that share it}, for the ones that do.

Empty on a correctly keyed forest — that is now the assertion this method exists to make, not a hope. Non-empty means this subtree holds objects that every other view will treat as one: the crop grid shows one of them, and which one depends on the order png_list happens to be in. That was the ordinary case before the object type went into the key; it is a regression now, and it is still detectable because node_id cannot collide even when key can.

keys() → Tuple[str, ...][source]

This node’s key and every descendant’s, depth-first, de-duplicated.

The order a selection made on this node publishes in, so the parent comes first and the crops open with the cell at the front of the grid.

As long as the subtree, now that the shared key carries the object type. It used not to be: a nucleus 1 and a pathogen 1 inside the same cell were one key, de-duplicating was the only honest thing to do (sending the same key twice would draw the same crop twice), and four objects opened as three. key_collisions() is the check that this no longer happens — it returns nothing on a typed forest.

node_ids() → Tuple[str, ...][source]

Every node_id below (and at) this node, depth-first.

Always distinct — unlike keys(), which cannot be.

walk(depth: int = 0) → Iterable[Tuple[int, LineageNode]][source]

This node then its descendants, depth-first, with their depth.

property node_id: str[source]

'pathogen:plate1_r1_c1_f1_pathogen1' — the table, then the key.

Distinct by construction, which key was not until the object type went into it. Now that it has, this is the same identity said twice — and it stays, because it is what proves the other one: key_collisions() is exactly the comparison between them, and a second identity that cannot collide is what makes the first one’s collisions detectable rather than invisible.

spacr.lineage.build_forest(frames: Mapping[str, pandas.DataFrame], *, root: str = ROOT_TABLE, typed: bool = True) → Tuple[LineageNode, ...][source]

Assemble object tables into one tree per root object.

Parameters:
  • frames – {table name: rows}. Only the tables named in LINEAGE_TABLES are read; anything else is ignored, so a caller can hand over everything it loaded.

  • root – the table whose rows become the roots.

  • typed – put each node’s object table into its shared key, so a cell’s nucleus 1 and its pathogen 1 are two keys. False rebuilds the untyped keys spaCR wrote before object types existed — kept so the collapse can be reproduced rather than only remembered, which is what makes LineageNode.key_collisions() a test with two sides to it.

Returns:

one LineageNode per root row, in field order then label order.

Raises:

LineageError – when the root table is absent or unusable. A missing child table is not an error — a run that measured cells and not pathogens produces a forest of childless cells, which is the truth about that run.

Childless roots are kept. A cell with no pathogens in an infection assay is the negative control working, and a tree that dropped it would show the infected population as if it were the whole plate.

spacr.lineage.child_tables() → Tuple[str, ...][source]

The tables whose rows carry a parent link, from the schema itself.

Read off spacr.schema.OBJECT_TABLE_SCHEMAS rather than written out here, so a table that gains a parent_column joins the tree without an edit in this file. organelle is in spacr.schema.CHILD_OBJECT_TABLES but has no declared schema entry, so it is included by name — its per-object table is optional and its rows carry cell_id when it exists.

spacr.lineage.describe_forest(forest: Sequence[LineageNode]) → str[source]

The shape of a whole forest in words.

Says the thing a tree of ten thousand rows cannot: how many parents have nothing inside them. In an infection assay that number is the readout, and having to count it by scrolling is how it gets estimated instead.

Parameters:

forest – root nodes as returned by build_forest(). The first root’s table is taken as the root table; an empty forest gets a sentence saying there is nothing to show.

spacr.lineage.field_key(row: Mapping[str, Any]) → str[source]

The prcf of one row — the field a label is unique within.

cell_id is an object label, and label 7 exists in every field of every plate. Matching children to parents on the label alone attaches every field’s nuclei to every field’s cell 7, which produces a tree that looks plausible and is wrong everywhere.

Parameters:

row – an object-table row carrying every column of spacr.schema.FIELD_KEY_COLUMNS; their values are joined with the schema key separator.

spacr.lineage.forest_key_collisions(forest: Sequence[LineageNode]) → Dict[str, Tuple[str, ...]][source]

Every shared key in forest that names more than one object.

Collisions are counted within a family, not across the whole forest: two different cells’ pathogens both labelled 1 in the same field cannot happen (the label is unique per mask), while a nucleus 1 and a pathogen 1 inside one cell is the ordinary case. Merging the two would report a collision on every plate.

Parameters:

forest – root nodes as returned by build_forest(); each root’s LineageNode.key_collisions() are merged into one mapping.

spacr.lineage.lineage_frame(forest: Sequence[LineageNode]) → pandas.DataFrame[source]

The forest flattened: one row per node, with its parent and depth.

For export, for a table view, and for the tests — a tree is awkward to assert on and this is the same information in a shape pandas can compare.

Parameters:

forest – root nodes as returned by build_forest(), flattened depth-first in order; roots get an empty parent_key and depth 0.

spacr.lineage.node_key(row: Mapping[str, Any], object_type: str | None = None) → str[source]

The shared object key of one row: field, then the typed object id.

Parameters:
  • row – an object-table row.

  • object_type – which table it came from. None builds the untyped key spaCR wrote before object types existed — the one that gives a cell’s nucleus 1 and its pathogen 1 the same name.

spacr.lineage.orphans(frames: Mapping[str, pandas.DataFrame], *, root: str = ROOT_TABLE) → pandas.DataFrame[source]

Child rows whose parent link names no row in the root table.

A finding, not an error. A nucleus whose cell_id is 12 in a field whose cell table has no object 12 means the two masks disagree — the nucleus segmentation found something the cell segmentation did not — and that is worth showing rather than dropping on the way into a tree.

Parameters:

frames – {table name: rows}. The root table is required; each other table in LINEAGE_TABLES that is present and has a parent-link column is checked, and anything else is ignored.

Returns:

the offending rows with table and parent_id columns added, in table then field then label order. Empty when everything attaches, which is the healthy case.

spacr.lineage.read_object_tables(db_path: str, tables: Sequence[str] | None = None, *, limit: int = 200000) → Dict[str, pandas.DataFrame][source]

Read the object tables a lineage needs. Safe on a worker thread.

Missing tables are absent from the result — a run that measured cells and nuclei but no pathogens is a legitimate experiment, not a broken database.

Parameters:
  • db_path – path to the measurements SQLite database, opened read-only so the loader is safe to run on a worker thread.

  • limit – per-table row cap, so a mis-aimed path cannot turn into a two-minute read behind a spinner.

Raises:

LineageError – when there is no database at db_path.

spacr.lineage.tree_for(frames: Mapping[str, pandas.DataFrame], key: str, *, root: str = ROOT_TABLE, typed: bool = True) → LineageNode | None[source]

The tree containing key, or None.

key may name the root or anything inside it: asked about a pathogen, this returns the cell it lives in, because the useful view of a pathogen is the cell around it. A caller that wants only the subtree calls LineageNode.find() on the result.

The whole forest is rebuilt and rescanned on every call, and the cost is the same whether the key is the first root or absent entirely, so a run of lookups should call build_forest() once and use LineageNode.find() on what it gets back.

Parameters:
  • frames – {table name: rows}, forwarded to build_forest(). Only the tables in LINEAGE_TABLES are read, so anything else handed over is ignored, and a run that measured no pathogens searches childless cells.

  • key – a shared object key, compared against LineageNode.key after str(). It is not compared against LineageNode.node_id, so the table-prefixed spelling of the same identity finds nothing and returns None.

  • root – the table whose rows are the roots being searched. It is not checked against what the children’s parent column actually points at: with root='nucleus', a pathogen whose cell_id equals a nucleus label is attached to that nucleus, and a cell key then finds nothing because no cell is in the forest at all.

  • typed – must agree with how key is spelled, and nothing checks that it does. A typed key searched with typed=False — or a legacy untyped one searched with the default — returns None, the same answer as “no such object”. With typed=False, collisions inside one family still identify that family, but a key found in different families raises rather than choosing whichever root sorts first.

Raises:

LineageError – when root is missing from frames, or when a table that is present cannot be named. The forest is built before the search, so a nucleus table without its field columns raises even when the wanted key belongs to a pathogen. A key that is merely not there is None, not an exception. Also raised when key names objects in more than one family; use typed keys to disambiguate them.

Nested helpers

lineage_frame.visit(node: LineageNode, parent: str, depth: int) → None

Append one node and recursively flatten its descendants.

Parameters:
  • node – lineage node to record before visiting its children.

  • parent – key of the caller-supplied parent, or an empty string for a root.

  • depth – zero-based depth to store for this node.

Returns:

None. A row containing node metadata and its current child count is appended to the captured list, then children are visited in their existing order with this node as parent.

spacr/lineage.py:504