spacr.selection

A shared selection and filter model, so the views can talk to each other.

spaCR ships four views over the same measurement table — the UMAP, the plate heatmap, the database browser and the annotation grid — and until now none of them knew the others existed. Lassoing a cluster in the UMAP told you nothing about where those cells sat on the plate, and narrowing to high-count wells in one view narrowed nothing anywhere else.

This module is the piece underneath that. It is deliberately pure pandas and numpy, with no Qt anywhere:

  • it can be tested without a display, which the Qt half cannot;

  • the same filter can be applied headless in spacr-run or a notebook, so a selection made in the GUI is expressible as something reproducible;

  • and the expensive part — evaluating a filter over a million-row table — is kept away from the event loop by construction.

spacr.qt.linked_selection wraps it in a QObject with signals.

Three ideas, kept separate on purpose

A filter narrows the population everyone is looking at: “wells with at least 200 cells, plate 3 only”. It is declarative, cheap to describe, and survives being written to a settings file.

A selection is the subset the user has pointed at inside that population: the cluster they lassoed. It is transient and arbitrary — there is no predicate that describes it — so it is carried as explicit keys.

Conflating the two is the usual mistake. A filter can be re-applied to a different table (another plate, a re-run) and still mean something; a selection cannot, because it names individual objects.

A request (ObjectRequest) is neither: it is one act of routing. “Open exactly these twelve objects, because they are the cells this model called infected and the annotator called uninfected.” A filter and a selection are state that every view reads; a request is an event that travels once, from the view that made it to the view that can show it. It carries the reason with it, because a view showing twelve crops out of ninety thousand has to be able to say why those twelve — otherwise they read as the whole dataset.

Identity

Objects are identified by object_keys() — the schema’s own row key (plateID, rowID, columnID, fieldID, object_label), joined with the schema separator. That is the one identity every table in measurements.db already agrees on, so a key from the UMAP means the same row in the plate view without a lookup table in between.

The object type is part of that identity. It did not used to be, and the collapse was real: object tables are one type per table, so a nucleus labelled 1 and a pathogen labelled 1 in the same field composed to the identical key. A cell’s own children are exactly the objects most likely to collide, which is where object linking is most useful — four objects opened as three crops, and which one you got depended on the row order of png_list (spacr.active_learning.crops_for_object_keys() keeps the first). The type now goes into the object component of the key, exactly as spacr.schema.object_id() writes it:

plate1_r1_c1_f1_7          # type not stated  (what spaCR always wrote)
plate1_r1_c1_f1_nucleus7   # type stated
plate1_r1_c1_f1_pathogen7  # a different object, and now a different key

A frame states its type either by carrying OBJECT_TYPE_COLUMN or by the reader passing object_type= — readers know which table they read, and with_object_type() is the one line that puts it on the frame.

Two rules keep the change from breaking anything already written:

Untyped keys do not move. The untyped form is byte for byte what it was, so every key in every stored selection, exported .h5ad and prediction bundle still composes and parses as it did. There is nothing on disk to migrate.

An untyped key is LESS SPECIFIC, not wrong. plate1_r1_c1_f1_7 means “the object labelled 7 in that field”, which is exactly what it always meant — so it matches that object whatever its type, rather than silently matching one of them. Symmetrically a typed key still matches a row that has not said what it is. Selection.mask_for() and ObjectRequest.select_from() both apply that rule, and it is the whole migration: an old key is read as the thing it always said, never re-read as something narrower.

Exceptions

FilterError

A filter that cannot be applied to the frame it was handed.

Classes

CategoryFilter

Keep rows whose column is one of values.

DataFilter

An AND of RangeFilter and CategoryFilter clauses.

ObjectRequest

One "open exactly these objects" act, on its way to whatever shows them.

RangeFilter

Keep rows whose column lies within [low, high].

Selection

The keys the user has pointed at, plus the filter they sit inside.

Functions

as_key_index(→ pd)

Coerce whatever names a set of objects into object keys.

escape_key_component(→ str)

Percent-escape one key component the way object_keys() does.

key_object_type(→ Optional[str])

The object type key states, or None when it states none.

match_keys(→ numpy.ndarray)

Boolean mask over keys of the ones wanted names.

object_keys(→ pd)

Return one stable key per row of df.

untyped_object_key(→ str)

key with the object type taken back off, for a looser comparison.

untyped_object_keys(→ pd)

The keys df would have had before object types existed.

with_object_type(→ pd)

Return a copy of df stamped with the object table it came from.

Module Contents

exception spacr.selection.FilterError[source]

Bases: ValueError

A filter that cannot be applied to the frame it was handed.

Raised rather than silently dropped. A filter naming a column that is not there has almost always been carried over from a different table, and quietly ignoring it would narrow the population by less than the user asked for while the UI still showed the filter as active — which reads as “these are all the cells that match” when it is not.

Initialize self. See help(type(self)) for accurate signature.

class spacr.selection.CategoryFilter[source]

Keep rows whose column is one of values.

Parameters:
  • column – categorical column evaluated by the filter.

  • values – accepted values; an empty tuple accepts no rows.

An EMPTY values keeps nothing, and that is deliberate: unticking every box in a category list means “show me none of these”, and quietly reinterpreting it as “show me all of them” would silently widen the population the user is looking at. The UI is expected to make an empty selection visible; this class will not paper over it.

describe() → str[source]

Return a compact description showing at most three accepted values.

mask(df: pd) → numpy.ndarray[source]

Return which rows of df have an accepted category.

Parameters:

df – data frame whose configured column is evaluated.

class spacr.selection.DataFilter[source]

An AND of RangeFilter and CategoryFilter clauses.

Declarative and re-appliable: the same filter means something on a re-run’s table, which is what separates it from a selection.

Parameters:

clauses – range and category predicates combined with logical AND.

add(clause) → DataFilter[source]

Add a clause, replacing any existing one on the same column.

Parameters:

clause – range or category clause to add.

Replacing rather than appending is what makes a slider a slider: a widget that emits on every drag would otherwise stack a hundred near-identical range clauses and turn an O(1) filter into an O(n) one.

apply(df: pd) → pd[source]

df narrowed to the rows this filter keeps.

Parameters:

df – data frame to filter.

clear() → DataFilter[source]

Remove every clause and return this filter for fluent chaining.

describe() → str[source]

One human line, for the header of whatever view is filtered.

A filtered view that does not say it is filtered is how someone reports a result computed on a fifth of their data.

mask(df: pd) → numpy.ndarray[source]

Return a boolean mask over df’s rows.

Parameters:

df – data frame to test against every clause.

An empty filter keeps everything, which is the identity a “no filter” state should have.

remove(column: str) → DataFilter[source]

Drop the clause on column, if any. Unknown columns are fine.

Parameters:

column – configured column whose clause should be removed.

property is_empty: bool[source]

Return whether this filter contains no clauses.

class spacr.selection.ObjectRequest[source]

One “open exactly these objects” act, on its way to whatever shows them.

Built by the view that asked and routed to the opener registered for kind — see spacr.qt.linked_selection.open_objects(). Openers take this one object rather than a handful of arguments so the request can grow a field without breaking every registered opener.

Parameters:
  • keys – anything as_key_index() accepts. Normalised on construction to an Index of strings in the caller’s order, with duplicates removed. Existing strings are neither validated nor recomposed, so they need not use OBJECT_KEY_COLUMNS.

  • reason – why these objects, in the words the receiving view will put on screen (“predicted infected, annotated uninfected”). Required and non-blank: a grid showing twelve crops out of ninety thousand and not saying why is read as the whole dataset.

  • source – the view that asked (“umap”, “classifier_evaluation”).

  • kind – the destination. Left empty by callers who take the default; the router stamps the kind it actually dispatched to, so an opener registered for two kinds can tell which one it was reached through.

  • timelapse – whether keys carry a timepoint, so the receiver resolves them against its own table the same way they were built.

  • context – free-form extras for the destination — per-key scores to sort by, a column to annotate into. Shallow-copied into a read-only outer mapping; nested mutable values remain shared with the caller.

Raises:

ValueError – on a blank reason.

An EMPTY request is legal. A confusion-matrix cell holding no errors is a real answer, and the destination saying “0 objects · no errors in this cell” is more use than an exception the caller has to catch.

__len__() → int[source]

Return the number of normalised keys in this request.

__post_init__() → None[source]

Normalise keys and reason, then freeze a shallow context copy.

as_selection() → Selection[source]

The same objects as a Selection, to publish as a highlight.

Opening a subset and highlighting it everywhere else are two acts, and this is the seam between them: a receiver that wants the plate view to light up the crops it just opened publishes this, rather than the router doing it behind the user’s back and wiping the lasso they made it with.

describe() → str[source]

One line for the receiving view’s header.

select_from(df: pd, *, object_type: Any = None) → pd[source]

The rows of df this request names, in the request’s order.

Not a mask, and not df’s order: the caller’s order is the answer for a request built worst-first, and a boolean mask would silently re-sort it back into table order. Keys with no row in df are dropped — a request can name objects a narrower table does not carry, and that is a smaller result, not an error.

Types are matched by specificity, exactly as Selection.mask_for() describes: an exact match first, then a key that states no type, then a typed key against a row that states none. Trying them in that order is what keeps a request naming both a nucleus 1 and a pathogen 1 opening as two rows rather than one.

Parameters:
  • df – object table to filter and reorder according to the request’s key order.

  • object_type – the table df came from, when the frame does not carry OBJECT_TYPE_COLUMN itself.

class spacr.selection.RangeFilter[source]

Keep rows whose column lies within [low, high].

Parameters:
  • column – numeric column evaluated by the range filter.

  • low – inclusive lower bound, or None for no lower bound.

  • high – inclusive upper bound, or None for no upper bound.

None on either bound means unbounded on that side, which is what a slider dragged to its end should mean — not “exclude everything”.

NaN never passes. A measurement that could not be computed is not a measurement inside the range, and letting it through would put objects with no value into a population the user defined by value.

describe() → str[source]

Return a compact description of this filter’s inclusive bounds.

mask(df: pd) → numpy.ndarray[source]

Return which rows of df pass this range.

Parameters:

df – data frame whose configured column is evaluated.

class spacr.selection.Selection[source]

The keys the user has pointed at, plus the filter they sit inside.

keys of None means “nothing selected”, which is different from an empty index meaning “an explicit selection that happens to be empty” — a lasso around blank space. Views draw those two differently: the first is the resting state, the second is a result.

Parameters:
  • keys – selected object keys, or None for the resting state.

  • source – view name that published the selection, used to avoid echo.

__len__() → int[source]

Return the selected-key count, or zero in the resting state.

classmethod from_frame(df: pd, source: str = '', *, timelapse: bool = False, object_type: Any = None) → Selection[source]

Select exactly the rows of df.

Parameters:
  • df – the rows to select, keyed by object_keys(). One key per row and no de-duplication — unlike from_keys() — so two rows naming the same object give that key twice. An empty frame gives an active selection of nothing, which is the lasso-caught-nothing state and not none().

  • source – the name of the view that made this. Stored verbatim, never stripped or validated. spacr.qt.linked_selection.LinkedView drops a selection whose source equals its own name, so the default "" is delivered back to the publishing view along with everyone else.

  • timelapse – key each timepoint of an object separately. Left False on a timelapse frame, every frame of an object composes to the same key.

  • object_type – the object table df came from. Overrides OBJECT_TYPE_COLUMN when the frame carries one as well, and is case-folded. None leaves the keys untyped rather than assuming cell.

Raises:
  • FilterError – if a key column is missing — including timeID when timelapse is set.

  • spacr.schema.KeyParseError – if object_type is not one spaCR keys objects by. with_object_type() treats such a table as a no-op; this does not. An empty df returns before the check, so the same argument raises or not depending on the row count.

classmethod from_keys(keys: Any, source: str = '', *, timelapse: bool = False, object_type: Any = None) → Selection[source]

Select exactly keys — anything as_key_index() accepts.

The counterpart to from_frame() for a view that never had the frame: a scatter plot holding an array of keys, or a screen restoring a selection from a settings file.

Parameters:
  • keys – a frame, another Selection, ONE key str, or an iterable of keys coerced with str. Order is kept and duplicates dropped, which from_frame() does not do.

  • source – as from_frame(). NOT inherited from a Selection passed as keys — the re-publisher is the source now, and carrying the original name over would suppress the echo in the wrong view.

  • timelapse – forwarded to object_keys(), so it is read only when keys is a frame. Key strings already carry whatever timepoint they were composed with, so passing this with a list or an Index is silently a no-op.

  • object_type – likewise frame-only: a list of untyped keys stays untyped however this is set, and a type spaCR cannot key objects by raises for a frame while being ignored for everything else.

Raises:
  • TypeError – if keys is not something that names objects — None included.

  • ValueError – for a resting Selection, which names nothing to select.

mask_for(df: pd, *, timelapse: bool = False, object_type: Any = None) → numpy.ndarray[source]

Boolean mask of the rows of df that are in this selection.

With no selection every row is in it — the resting state highlights nothing rather than everything, but a caller asking “which rows are selected” when nothing is gets the whole frame rather than an empty one, which is what keeps df[sel.mask_for(df)] meaning “the data the user is looking at”.

Types are matched by specificity, not by equality, and that is the whole of the object-type migration:

  • two typed keys match when they agree — a nucleus 1 is not a pathogen 1, which is the collapse this exists to end;

  • a key stating no type matches a row of any type. It says “the object labelled 7 in that field”, which is exactly what it said before types existed, so every selection ever saved goes on naming what it named. It is deliberately not narrowed to one type: an old key that quietly resolved to one of four objects is the bug, and replacing it with a different silent choice would not be a fix.

  • a typed key matches a row stating no type. The row has not contradicted it; it has said nothing.

Parameters:
  • df – object table to match; the returned NumPy mask has one boolean value per row in this frame.

  • object_type – the table df came from, when the frame does not carry OBJECT_TYPE_COLUMN itself.

classmethod none() → Selection[source]

Return the resting selection with no keys and no source.

property is_active: bool[source]

Return whether this is an explicit, possibly empty, selection.

spacr.selection.as_key_index(keys: Any, *, timelapse: bool = False, object_type: Any = None) → pd[source]

Coerce whatever names a set of objects into object keys.

Parameters:
  • keys –

    one of

    • a Selection — its keys. A resting selection (keys is None) raises: “open nothing” is not a request, and turning it into an empty one would silently open an empty view.

    • a pandas.DataFrame carrying OBJECT_KEY_COLUMNS — object_keys() of it.

    • a single str — ONE key. Special-cased on purpose: a string is iterable, so falling through to the iterable branch would open one object per character, which is the sort of thing that produces a grid of 23 empty tiles rather than an error.

    • any other iterable of keys — an Index, a Series, a list — coerced with str.

  • timelapse – passed to object_keys() for the frame case, so a timelapse table keys each frame of an object separately.

  • object_type – likewise — the object table a frame came from.

Returns:

a pandas.Index of str.

Raises:

Order is preserved and duplicates are dropped. Order is load-bearing: it is what carries “worst errors first” from a confusion-matrix cell through to whatever opens them, and a duplicated key would draw the same crop twice in the grid.

Key strings are passed through untouched, including untyped ones. It is tempting to normalise them here, and it would be wrong: a caller may legitimately hand over a prcfo, a crop path or a file name (see spacr.active_learning.crops_for_object_keys()), and rewriting those into something that looks like an object key would break the resolution they were relying on. Untyped and typed keys are reconciled where they are compared — match_keys() — not where they are collected.

spacr.selection.escape_key_component(value: Any) → str[source]

Percent-escape one key component the way object_keys() does.

Public because a resolver has to compose the same spelling the producer did. spacr.active_learning.crops_for_object_keys() builds its lookup keys out of png_list’s own metadata columns, and a fieldID of 'f_1' composes to 'f%5F1' here while a bare '_'.join composes it to 'f_1' — so a routed selection’s key matched nothing and that crop was silently dropped from the crops it opened, while its neighbours in the same selection opened normally.

Parameters:

value – one component of an object key.

Returns:

the component as it appears inside a composed key. Unchanged when it contains neither % nor the separator, which is the overwhelmingly common case.

spacr.selection.key_object_type(key: Any) → str | None[source]

The object type key states, or None when it states none.

Parameters:

key – object key to inspect.

None means not stated. It does not mean “cell”, and nothing here will ever guess: every key spaCR wrote before today is untyped, so a default would put a type on the whole world’s existing data.

spacr.selection.match_keys(keys: Any, wanted: Iterable[Any]) → numpy.ndarray[source]

Boolean mask over keys of the ones wanted names.

The key-to-key form of Selection.mask_for(), for a view that holds an array of keys rather than the frame they came from — a scatter plot, a UMAP that derived its point identity once at load, a tree. Those views reached for a bare Index.isin, which asks for exact equality, and exact equality is the wrong question the moment one side of a link states an object type and the other does not: a table publishing …_f1_cell1 highlighted nothing at all in a UMAP whose points are keyed …_f1_1. Silence, in the one place a linked view has to be loud.

Parameters:
  • keys – the view’s own keys, in its own order.

  • wanted – the keys to match against — a selection’s, a request’s.

Returns:

a boolean numpy.ndarray aligned to keys.

spacr.selection.object_keys(df: pd, *, timelapse: bool = False, object_type: Any = None) → pd[source]

Return one stable key per row of df.

Parameters:
  • df – any frame carrying the object key columns.

  • timelapse – include the timepoint, so the same object at two frames is two keys rather than one. A timelapse table that leaves this False collapses every frame of an object onto a single key, which is the bug that has bitten this codebase repeatedly in the other direction (see schema.parse_prcfo).

  • object_type – the object table these rows came from. Overrides OBJECT_TYPE_COLUMN when the frame also carries one. None means “read it off the frame, and leave the keys untyped if it does not say” — not “these are cells”.

Returns:

a pandas.Index of str, aligned to df.index.

Raises:

A typed key is exactly the prcfo spacr.schema.compose_prcfo() writes for the same object, so the two identities converge as soon as the type is known. They differ only in the untyped case, where this joins the label bare and prcfo writes 'o7' — kept that way on purpose, since changing it would move every key that already exists.

spacr.selection.untyped_object_key(key: Any) → str[source]

key with the object type taken back off, for a looser comparison.

Parameters:

key – object key to reduce to its untyped form.

'p_r1_c1_f1_nucleus7' → 'p_r1_c1_f1_7', and a key that already states no type is returned unchanged. A prcfo reduces the same way ('p_r1_c1_f1_o7' → 'p_r1_c1_f1_7'), which is what lets a key copied out of a crop table match a key built from a measurement table.

Anything this cannot read as an object key — a crop path, a file name — is returned untouched rather than mangled: those travel through the same routing contract and must survive it.

spacr.selection.untyped_object_keys(df: pd, *, timelapse: bool = False) → pd[source]

The keys df would have had before object types existed.

Parameters:

df – object rows whose legacy untyped keys are requested.

Not a legacy shim: it is the less specific name for the same rows, and it is what makes an old key go on meaning what it always meant. A key naming no type says “the object labelled 7 in that field” and has to match that object whatever its type — see Selection.mask_for().

spacr.selection.with_object_type(df: pd, object_type: Any) → pd[source]

Return a copy of df stamped with the object table it came from.

The one line a reader adds after loading a table, so every key built from the frame afterwards says which of a cell’s children it names. Readers are where this belongs: a frame does not know what it is, but whatever ran SELECT * FROM nucleus does.

An object_type spaCR does not key objects by — png_list, a summary, a user’s own table — is a no-op, not an error. Such a table’s rows are keyed the way they always were, which is the correct answer for something that is not one of the four analysis compartments.

Parameters:
  • df – any frame.

  • object_type – an object table name, or None.

Returns:

df unchanged, or a copy carrying OBJECT_TYPE_COLUMN.

Nested helpers

ObjectRequest.select_from.rank(typed: str, plain: str) → int

Find a row key’s position in the captured request.

Parameters:
  • typed – row key including its object type when available.

  • plain – the same row key without an object type.

Returns:

request position using exact typed, loose untyped, then narrowed typed-request precedence. Narrowing is allowed only for an untyped row (typed == plain); absent keys return -1.

spacr/selection.py:991