spacr.qt.widgets.data_filter_panel

Local data filter — narrow every open view at once.

JMP’s Local Data Filter, which is the feature its users reach for most: a panel of live controls that subsets every plot, table and image grid simultaneously, so “does this hit survive if I drop the low-count wells?” is a second’s work rather than a re-run.

It writes into spacr.qt.linked_selection.linked_selection(), so the panel knows nothing about the views and the views know nothing about the panel.

Choosing what to offer

A spaCR measurement table has hundreds of columns, so offering all of them in one list is the same as offering none. The panel classifies them instead:

  • categorical — few enough distinct values to tick (MAX_CATEGORY_VALUES), which is what plate, row, column, gene and class look like;

  • numeric — anything pandas reads as a number, offered as a range;

  • skipped — high-cardinality text, which is neither tickable nor rangeable, and object keys, which identify rows rather than describe them.

The classification is a suggestion: the picker lists everything it can filter, and the user chooses. Nothing is filtered until they do.

Cost

Every clause change re-evaluates the filter over the whole frame, so the panel debounces. A dragged spinbox emits per keystroke, and re-filtering a million rows per keystroke would make the control unusable — which is why spacr.selection.DataFilter also replaces rather than appends a clause on the same column.

Classes

DataFilterPanel

The panel. Call set_frame(), then let the user drive.

Functions

classify_columns(→ Dict[str, str])

Sort frame's columns into 'category', 'range' or 'skip'.

Module Contents

class spacr.qt.widgets.data_filter_panel.DataFilterPanel(parent=None, *, link=None)[source]

Bases: PySide6.QtWidgets.QWidget

The panel. Call set_frame(), then let the user drive.

Emits filter_changed after publishing, for a host that wants to update a count label without subscribing to the shared model itself.

Parameters:
  • parent – parent widget.

  • link – the LinkedSelection the filter publishes into, so filtering here narrows every view on it. None joins the shared one; pass a private one in a test.

Build the shared-filter panel.

Parameters:
  • parent – parent widget, or None.

  • link – the selection link to publish into. Injectable so a test drives a private one rather than the process-wide link every other open view is listening to.

add_column(column: str) → None[source]

Add a clause row for column. Adding twice is a no-op.

Parameters:

column – column of the current frame; a range row is added for a "range" column and a category row for a "category" one. Anything else, or no frame yet, adds nothing.

available_columns() → List[str][source]

The columns the picker is offering, in order.

clear() → None[source]

Drop every clause and publish an empty filter.

current_filter() → spacr.selection.DataFilter[source]

The filter the controls currently describe.

flush() → None[source]

Publish immediately instead of waiting out the debounce.

For tests, and for a host that needs the filter applied before it does something else — closing the panel, or running an export.

load(path: str) → List[str][source]

Read a filter set from path. Returns the missing columns.

Parameters:

path – JSON file written by save(); its contents are applied with restore().

remove_column(column: str) → None[source]

Drop one column’s filter.

Parameters:

column – the column’s name.

restore(state: dict) → List[str][source]

Apply a saved set to the CURRENT frame.

Parameters:

state – dict as state() returns it, {"version": 1, "filters": [...]}; any other version raises ValueError. Existing clauses are cleared first.

Returns:

the columns that could not be restored, so the caller can say so. A filter set saved against one table and loaded against another is a normal thing to do – what must not happen is it appearing to work while quietly filtering on nothing.

save(path: str) → str[source]

Write the filter set to path as JSON.

Parameters:

path – destination file, overwritten with state(); it is also the return value.

set_frame(frame: pandas.DataFrame) → None[source]

Point the panel at a table and offer its filterable columns.

Existing clauses are dropped rather than carried over: a clause naming a column the new frame does not have would raise on the next apply, and silently keeping only the ones that still resolve would narrow by less than the panel claims to.

Parameters:

frame – the table to filter; every column that classify_columns() does not mark "skip" is offered.

state() → dict[source]

The whole panel, as plain data.

Versioned because the row kinds will grow. A reader that meets a version it does not know refuses rather than guessing, since a half-applied filter set silently selects the wrong rows.

spacr.qt.widgets.data_filter_panel.classify_columns(frame: pandas.DataFrame) → Dict[str, str][source]

Sort frame’s columns into 'category', 'range' or 'skip'.

Pure and Qt-free so the rule can be tested directly — the panel is only a rendering of it.

Memoised per frame object. The classification runs Series.nunique() over every column, which on a 200 000-row x 48-column measurement table costs 230 ms. Loading one table used to pay that four times inside a single set_frame — this function, plus spacr.qt.widgets.graph_spec.column_kinds(), plus GraphSpec.kinds_for, plus plottable_columns, each re-deriving the same answer from the same object — 0.9 s of the ~1.9 s the GUI thread spent delivering a freshly loaded frame.

The key is the frame’s identity, not its contents: a subset has fewer rows and can genuinely classify differently, so re-deriving for a filtered frame is correct rather than wasteful. id() alone would be unsound because CPython reuses addresses, so the entry also holds a weakref and the hit is confirmed with is — a collected frame cannot produce a false hit, because its weakref resolves to None. The shape is checked too, since a frame can be mutated in place.

Parameters:

frame – any measurement-shaped frame.

Returns:

column name → kind. A fresh dict on every call, because callers (GraphSpec.kinds_for) update it in place.