spacr.selection¶
A shared selection and filter model, so the views can talk to each other.
spaCR ships four views over the same measurement table — the UMAP, the plate heatmap, the database browser and the annotation grid — and until now none of them knew the others existed. Lassoing a cluster in the UMAP told you nothing about where those cells sat on the plate, and narrowing to high-count wells in one view narrowed nothing anywhere else.
This module is the piece underneath that. It is deliberately pure pandas and numpy, with no Qt anywhere:
it can be tested without a display, which the Qt half cannot;
the same filter can be applied headless in
spacr-runor a notebook, so a selection made in the GUI is expressible as something reproducible;and the expensive part — evaluating a filter over a million-row table — is kept away from the event loop by construction.
spacr.qt.linked_selection wraps it in a QObject with signals.
Three ideas, kept separate on purpose¶
A filter narrows the population everyone is looking at: “wells with at least 200 cells, plate 3 only”. It is declarative, cheap to describe, and survives being written to a settings file.
A selection is the subset the user has pointed at inside that population: the cluster they lassoed. It is transient and arbitrary — there is no predicate that describes it — so it is carried as explicit keys.
Conflating the two is the usual mistake. A filter can be re-applied to a different table (another plate, a re-run) and still mean something; a selection cannot, because it names individual objects.
A request (ObjectRequest) is neither: it is one act of routing.
“Open exactly these twelve objects, because they are the cells this model
called infected and the annotator called uninfected.” A filter and a selection
are state that every view reads; a request is an event that travels once,
from the view that made it to the view that can show it. It carries the reason
with it, because a view showing twelve crops out of ninety thousand has to be
able to say why those twelve — otherwise they read as the whole dataset.
Identity¶
Objects are identified by object_keys() — the schema’s own row key
(plateID, rowID, columnID, fieldID, object_label), joined
with the schema separator. That is the one identity every table in
measurements.db already agrees on, so a key from the UMAP means the same
row in the plate view without a lookup table in between.
The object type is part of that identity. It did not used to be, and the
collapse was real: object tables are one type per table, so a nucleus
labelled 1 and a pathogen labelled 1 in the same field composed to the
identical key. A cell’s own children are exactly the objects most likely to
collide, which is where object linking is most useful — four objects opened
as three crops, and which one you got depended on the row order of
png_list (spacr.active_learning.crops_for_object_keys() keeps the
first). The type now goes into the object component of the key, exactly as
spacr.schema.object_id() writes it:
plate1_r1_c1_f1_7 # type not stated (what spaCR always wrote)
plate1_r1_c1_f1_nucleus7 # type stated
plate1_r1_c1_f1_pathogen7 # a different object, and now a different key
A frame states its type either by carrying OBJECT_TYPE_COLUMN or by
the reader passing object_type= — readers know which table they read, and
with_object_type() is the one line that puts it on the frame.
Two rules keep the change from breaking anything already written:
Untyped keys do not move. The untyped form is byte for byte what it was, so
every key in every stored selection, exported .h5ad and prediction bundle
still composes and parses as it did. There is nothing on disk to migrate.
An untyped key is LESS SPECIFIC, not wrong. plate1_r1_c1_f1_7 means
“the object labelled 7 in that field”, which is exactly what it always meant
— so it matches that object whatever its type, rather than silently matching
one of them. Symmetrically a typed key still matches a row that has not said
what it is. Selection.mask_for() and ObjectRequest.select_from()
both apply that rule, and it is the whole migration: an old key is read as
the thing it always said, never re-read as something narrower.
Exceptions¶
A filter that cannot be applied to the frame it was handed. |
Classes¶
Keep rows whose |
|
An AND of |
|
One "open exactly these objects" act, on its way to whatever shows them. |
|
Keep rows whose |
|
The keys the user has pointed at, plus the filter they sit inside. |
Functions¶
|
Coerce whatever names a set of objects into object keys. |
|
Percent-escape one key component the way |
|
The object type |
|
Boolean mask over |
|
Return one stable key per row of |
|
|
|
The keys |
|
Return a copy of |
Module Contents¶
- exception spacr.selection.FilterError[source]¶
Bases:
ValueErrorA filter that cannot be applied to the frame it was handed.
Raised rather than silently dropped. A filter naming a column that is not there has almost always been carried over from a different table, and quietly ignoring it would narrow the population by less than the user asked for while the UI still showed the filter as active — which reads as “these are all the cells that match” when it is not.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.selection.CategoryFilter[source]¶
Keep rows whose
columnis one ofvalues.- Parameters:
column – categorical column evaluated by the filter.
values – accepted values; an empty tuple accepts no rows.
An EMPTY
valueskeeps nothing, and that is deliberate: unticking every box in a category list means “show me none of these”, and quietly reinterpreting it as “show me all of them” would silently widen the population the user is looking at. The UI is expected to make an empty selection visible; this class will not paper over it.- mask(df: pd) numpy.ndarray[source]¶
Return which rows of
dfhave an accepted category.- Parameters:
df – data frame whose configured column is evaluated.
- class spacr.selection.DataFilter[source]¶
An AND of
RangeFilterandCategoryFilterclauses.Declarative and re-appliable: the same filter means something on a re-run’s table, which is what separates it from a selection.
- Parameters:
clauses – range and category predicates combined with logical AND.
- add(clause) DataFilter[source]¶
Add a clause, replacing any existing one on the same column.
- Parameters:
clause – range or category clause to add.
Replacing rather than appending is what makes a slider a slider: a widget that emits on every drag would otherwise stack a hundred near-identical range clauses and turn an O(1) filter into an O(n) one.
- apply(df: pd) pd[source]¶
dfnarrowed to the rows this filter keeps.- Parameters:
df – data frame to filter.
- clear() DataFilter[source]¶
Remove every clause and return this filter for fluent chaining.
- describe() str[source]¶
One human line, for the header of whatever view is filtered.
A filtered view that does not say it is filtered is how someone reports a result computed on a fifth of their data.
- mask(df: pd) numpy.ndarray[source]¶
Return a boolean mask over
df’s rows.- Parameters:
df – data frame to test against every clause.
An empty filter keeps everything, which is the identity a “no filter” state should have.
- remove(column: str) DataFilter[source]¶
Drop the clause on
column, if any. Unknown columns are fine.- Parameters:
column – configured column whose clause should be removed.
- class spacr.selection.ObjectRequest[source]¶
One “open exactly these objects” act, on its way to whatever shows them.
Built by the view that asked and routed to the opener registered for
kind— seespacr.qt.linked_selection.open_objects(). Openers take this one object rather than a handful of arguments so the request can grow a field without breaking every registered opener.- Parameters:
keys – anything
as_key_index()accepts. Normalised on construction to an Index of strings in the caller’s order, with duplicates removed. Existing strings are neither validated nor recomposed, so they need not useOBJECT_KEY_COLUMNS.reason – why these objects, in the words the receiving view will put on screen (“predicted infected, annotated uninfected”). Required and non-blank: a grid showing twelve crops out of ninety thousand and not saying why is read as the whole dataset.
source – the view that asked (“umap”, “classifier_evaluation”).
kind – the destination. Left empty by callers who take the default; the router stamps the kind it actually dispatched to, so an opener registered for two kinds can tell which one it was reached through.
timelapse – whether
keyscarry a timepoint, so the receiver resolves them against its own table the same way they were built.context – free-form extras for the destination — per-key scores to sort by, a column to annotate into. Shallow-copied into a read-only outer mapping; nested mutable values remain shared with the caller.
- Raises:
ValueError – on a blank
reason.
An EMPTY request is legal. A confusion-matrix cell holding no errors is a real answer, and the destination saying “0 objects · no errors in this cell” is more use than an exception the caller has to catch.
- as_selection() Selection[source]¶
The same objects as a
Selection, to publish as a highlight.Opening a subset and highlighting it everywhere else are two acts, and this is the seam between them: a receiver that wants the plate view to light up the crops it just opened publishes this, rather than the router doing it behind the user’s back and wiping the lasso they made it with.
- select_from(df: pd, *, object_type: Any = None) pd[source]¶
The rows of
dfthis request names, in the request’s order.Not a mask, and not
df’s order: the caller’s order is the answer for a request built worst-first, and a boolean mask would silently re-sort it back into table order. Keys with no row indfare dropped — a request can name objects a narrower table does not carry, and that is a smaller result, not an error.Types are matched by specificity, exactly as
Selection.mask_for()describes: an exact match first, then a key that states no type, then a typed key against a row that states none. Trying them in that order is what keeps a request naming both a nucleus 1 and a pathogen 1 opening as two rows rather than one.- Parameters:
df – object table to filter and reorder according to the request’s key order.
object_type – the table
dfcame from, when the frame does not carryOBJECT_TYPE_COLUMNitself.
- class spacr.selection.RangeFilter[source]¶
Keep rows whose
columnlies within[low, high].- Parameters:
column – numeric column evaluated by the range filter.
low – inclusive lower bound, or
Nonefor no lower bound.high – inclusive upper bound, or
Nonefor no upper bound.
Noneon either bound means unbounded on that side, which is what a slider dragged to its end should mean — not “exclude everything”.NaN never passes. A measurement that could not be computed is not a measurement inside the range, and letting it through would put objects with no value into a population the user defined by value.
- mask(df: pd) numpy.ndarray[source]¶
Return which rows of
dfpass this range.- Parameters:
df – data frame whose configured column is evaluated.
- class spacr.selection.Selection[source]¶
The keys the user has pointed at, plus the filter they sit inside.
keysofNonemeans “nothing selected”, which is different from an empty index meaning “an explicit selection that happens to be empty” — a lasso around blank space. Views draw those two differently: the first is the resting state, the second is a result.- Parameters:
keys – selected object keys, or
Nonefor the resting state.source – view name that published the selection, used to avoid echo.
- classmethod from_frame(df: pd, source: str = '', *, timelapse: bool = False, object_type: Any = None) Selection[source]¶
Select exactly the rows of
df.- Parameters:
df – the rows to select, keyed by
object_keys(). One key per row and no de-duplication — unlikefrom_keys()— so two rows naming the same object give that key twice. An empty frame gives an active selection of nothing, which is the lasso-caught-nothing state and notnone().source – the name of the view that made this. Stored verbatim, never stripped or validated.
spacr.qt.linked_selection.LinkedViewdrops a selection whosesourceequals its own name, so the default""is delivered back to the publishing view along with everyone else.timelapse – key each timepoint of an object separately. Left False on a timelapse frame, every frame of an object composes to the same key.
object_type – the object table
dfcame from. OverridesOBJECT_TYPE_COLUMNwhen the frame carries one as well, and is case-folded.Noneleaves the keys untyped rather than assumingcell.
- Raises:
FilterError – if a key column is missing — including
timeIDwhentimelapseis set.spacr.schema.KeyParseError – if
object_typeis not one spaCR keys objects by.with_object_type()treats such a table as a no-op; this does not. An emptydfreturns before the check, so the same argument raises or not depending on the row count.
- classmethod from_keys(keys: Any, source: str = '', *, timelapse: bool = False, object_type: Any = None) Selection[source]¶
Select exactly
keys— anythingas_key_index()accepts.The counterpart to
from_frame()for a view that never had the frame: a scatter plot holding an array of keys, or a screen restoring a selection from a settings file.- Parameters:
keys – a frame, another
Selection, ONE keystr, or an iterable of keys coerced withstr. Order is kept and duplicates dropped, whichfrom_frame()does not do.source – as
from_frame(). NOT inherited from aSelectionpassed askeys— the re-publisher is the source now, and carrying the original name over would suppress the echo in the wrong view.timelapse – forwarded to
object_keys(), so it is read only whenkeysis a frame. Key strings already carry whatever timepoint they were composed with, so passing this with a list or an Index is silently a no-op.object_type – likewise frame-only: a list of untyped keys stays untyped however this is set, and a type spaCR cannot key objects by raises for a frame while being ignored for everything else.
- Raises:
TypeError – if
keysis not something that names objects —Noneincluded.ValueError – for a resting
Selection, which names nothing to select.
- mask_for(df: pd, *, timelapse: bool = False, object_type: Any = None) numpy.ndarray[source]¶
Boolean mask of the rows of
dfthat are in this selection.With no selection every row is in it — the resting state highlights nothing rather than everything, but a caller asking “which rows are selected” when nothing is gets the whole frame rather than an empty one, which is what keeps
df[sel.mask_for(df)]meaning “the data the user is looking at”.Types are matched by specificity, not by equality, and that is the whole of the object-type migration:
two typed keys match when they agree — a nucleus 1 is not a pathogen 1, which is the collapse this exists to end;
a key stating no type matches a row of any type. It says “the object labelled 7 in that field”, which is exactly what it said before types existed, so every selection ever saved goes on naming what it named. It is deliberately not narrowed to one type: an old key that quietly resolved to one of four objects is the bug, and replacing it with a different silent choice would not be a fix.
a typed key matches a row stating no type. The row has not contradicted it; it has said nothing.
- Parameters:
df – object table to match; the returned NumPy mask has one boolean value per row in this frame.
object_type – the table
dfcame from, when the frame does not carryOBJECT_TYPE_COLUMNitself.
- spacr.selection.as_key_index(keys: Any, *, timelapse: bool = False, object_type: Any = None) pd[source]¶
Coerce whatever names a set of objects into object keys.
- Parameters:
keys –
one of
a
Selection— its keys. A resting selection (keys is None) raises: “open nothing” is not a request, and turning it into an empty one would silently open an empty view.a
pandas.DataFramecarryingOBJECT_KEY_COLUMNS—object_keys()of it.a single
str— ONE key. Special-cased on purpose: a string is iterable, so falling through to the iterable branch would open one object per character, which is the sort of thing that produces a grid of 23 empty tiles rather than an error.any other iterable of keys — an Index, a Series, a list — coerced with
str.
timelapse – passed to
object_keys()for the frame case, so a timelapse table keys each frame of an object separately.object_type – likewise — the object table a frame came from.
- Returns:
a
pandas.Indexofstr.- Raises:
TypeError – if
keysis not something that can name objects.ValueError – for a resting
Selection.
Order is preserved and duplicates are dropped. Order is load-bearing: it is what carries “worst errors first” from a confusion-matrix cell through to whatever opens them, and a duplicated key would draw the same crop twice in the grid.
Key strings are passed through untouched, including untyped ones. It is tempting to normalise them here, and it would be wrong: a caller may legitimately hand over a
prcfo, a crop path or a file name (seespacr.active_learning.crops_for_object_keys()), and rewriting those into something that looks like an object key would break the resolution they were relying on. Untyped and typed keys are reconciled where they are compared —match_keys()— not where they are collected.
- spacr.selection.escape_key_component(value: Any) str[source]¶
Percent-escape one key component the way
object_keys()does.Public because a resolver has to compose the same spelling the producer did.
spacr.active_learning.crops_for_object_keys()builds its lookup keys out ofpng_list’s own metadata columns, and afieldIDof'f_1'composes to'f%5F1'here while a bare'_'.joincomposes it to'f_1'— so a routed selection’s key matched nothing and that crop was silently dropped from the crops it opened, while its neighbours in the same selection opened normally.- Parameters:
value – one component of an object key.
- Returns:
the component as it appears inside a composed key. Unchanged when it contains neither
%nor the separator, which is the overwhelmingly common case.
- spacr.selection.key_object_type(key: Any) str | None[source]¶
The object type
keystates, orNonewhen it states none.- Parameters:
key – object key to inspect.
Nonemeans not stated. It does not mean “cell”, and nothing here will ever guess: every key spaCR wrote before today is untyped, so a default would put a type on the whole world’s existing data.
- spacr.selection.match_keys(keys: Any, wanted: Iterable[Any]) numpy.ndarray[source]¶
Boolean mask over
keysof the oneswantednames.The key-to-key form of
Selection.mask_for(), for a view that holds an array of keys rather than the frame they came from — a scatter plot, a UMAP that derived its point identity once at load, a tree. Those views reached for a bareIndex.isin, which asks for exact equality, and exact equality is the wrong question the moment one side of a link states an object type and the other does not: a table publishing…_f1_cell1highlighted nothing at all in a UMAP whose points are keyed…_f1_1. Silence, in the one place a linked view has to be loud.- Parameters:
keys – the view’s own keys, in its own order.
wanted – the keys to match against — a selection’s, a request’s.
- Returns:
a boolean
numpy.ndarrayaligned tokeys.
- spacr.selection.object_keys(df: pd, *, timelapse: bool = False, object_type: Any = None) pd[source]¶
Return one stable key per row of
df.- Parameters:
df – any frame carrying the object key columns.
timelapse – include the timepoint, so the same object at two frames is two keys rather than one. A timelapse table that leaves this False collapses every frame of an object onto a single key, which is the bug that has bitten this codebase repeatedly in the other direction (see
schema.parse_prcfo).object_type – the object table these rows came from. Overrides
OBJECT_TYPE_COLUMNwhen the frame also carries one.Nonemeans “read it off the frame, and leave the keys untyped if it does not say” — not “these are cells”.
- Returns:
a
pandas.Indexofstr, aligned todf.index.- Raises:
FilterError – if any key column is missing.
spacr.schema.KeyParseError – if a stated type is not one spaCR keys objects by.
A typed key is exactly the
prcfospacr.schema.compose_prcfo()writes for the same object, so the two identities converge as soon as the type is known. They differ only in the untyped case, where this joins the label bare andprcfowrites'o7'— kept that way on purpose, since changing it would move every key that already exists.
- spacr.selection.untyped_object_key(key: Any) str[source]¶
keywith the object type taken back off, for a looser comparison.- Parameters:
key – object key to reduce to its untyped form.
'p_r1_c1_f1_nucleus7'→'p_r1_c1_f1_7', and a key that already states no type is returned unchanged. Aprcforeduces the same way ('p_r1_c1_f1_o7'→'p_r1_c1_f1_7'), which is what lets a key copied out of a crop table match a key built from a measurement table.Anything this cannot read as an object key — a crop path, a file name — is returned untouched rather than mangled: those travel through the same routing contract and must survive it.
- spacr.selection.untyped_object_keys(df: pd, *, timelapse: bool = False) pd[source]¶
The keys
dfwould have had before object types existed.- Parameters:
df – object rows whose legacy untyped keys are requested.
Not a legacy shim: it is the less specific name for the same rows, and it is what makes an old key go on meaning what it always meant. A key naming no type says “the object labelled 7 in that field” and has to match that object whatever its type — see
Selection.mask_for().
- spacr.selection.with_object_type(df: pd, object_type: Any) pd[source]¶
Return a copy of
dfstamped with the object table it came from.The one line a reader adds after loading a table, so every key built from the frame afterwards says which of a cell’s children it names. Readers are where this belongs: a frame does not know what it is, but whatever ran
SELECT * FROM nucleusdoes.An
object_typespaCR does not key objects by —png_list, a summary, a user’s own table — is a no-op, not an error. Such a table’s rows are keyed the way they always were, which is the correct answer for something that is not one of the four analysis compartments.- Parameters:
df – any frame.
object_type – an object table name, or
None.
- Returns:
dfunchanged, or a copy carryingOBJECT_TYPE_COLUMN.
Nested helpers¶
- ObjectRequest.select_from.rank(typed: str, plain: str) int¶
Find a row key’s position in the captured request.
- Parameters:
typed – row key including its object type when available.
plain – the same row key without an object type.
- Returns:
request position using exact typed, loose untyped, then narrowed typed-request precedence. Narrowing is allowed only for an untyped row (
typed == plain); absent keys return -1.
spacr/selection.py:991