spacr.column_groups

Choose measurement columns individually or by named groups.

Three ways of naming the same set, because a measurement table names a column three ways at once – cell_channel_1_mean_intensity is a CELL measurement, a CHANNEL 1 measurement and an INTENSITY measurement, and which of those a user means depends on the question:

object cell, nucleus, pathogen, cytoplasm, organelle channel channel_0, channel_1, … family morphology, intensity, texture, correlation, moment

THE FAMILIES ARE NOT INVENTED HERE. They come from spacr.feature_dict.FEATURE_FAMILIES, which is the dictionary that already documents every column, and the classification is spacr.feature_dict.parse_column(). A second taxonomy would be a second thing to keep in step with the measurement code, and it would disagree first in exactly the corners nobody checks.

Qt-free, so the picker’s logic is testable without a display and the same grouping can serve the CLI, a notebook, or a future screen.

Functions

classify(→ Dict[str, Dict[str, List[str]]])

Group columns by object, by channel and by family.

columns_in(→ List[str])

Every column in one named group.

group_names(→ Dict[str, List[str]])

The group names a picker should offer, per kind, sorted.

resolve() → List[str])

The columns a selection actually means, de-duplicated and in order.

summarise() → str)

One line saying what is selected, for the picker to show.

Module Contents

spacr.column_groups.classify(columns: Iterable[str]) → Dict[str, Dict[str, List[str]]][source]

Group columns by object, by channel and by family.

Parameters:

columns – the table’s column names.

Returns:

{kind: {group name: [column, ...]}} for the kinds in GROUP_KINDS. A column can appear in one group of each kind, which is the point – cell_channel_1_mean_intensity is in object/cell, channel/channel_1 and family/intensity.

spacr.column_groups.columns_in(columns: Iterable[str], kind: str, name: str) → List[str][source]

Every column in one named group.

Raises:

KeyError – an unknown kind, naming the ones there are. A typo here would otherwise select nothing and read as “this table has no intensity measurements”.

Parameters:
  • columns – available measurement column names to classify.

  • kind – grouping dimension, one of GROUP_KINDS.

  • name – group name within the selected grouping dimension.

spacr.column_groups.group_names(columns: Iterable[str]) → Dict[str, List[str]][source]

The group names a picker should offer, per kind, sorted.

Channels sort NUMERICALLY – channel_2 before channel_10 – which a plain string sort gets wrong the moment a run has more than ten.

Parameters:

columns – available measurement column names to classify.

spacr.column_groups.resolve(columns: Iterable[str], selection: Mapping[str, Sequence[str]] | None = None, *, explicit: Sequence[str] = ()) → List[str][source]

The columns a selection actually means, de-duplicated and in order.

Parameters:
  • columns – available column names. Their input order is preserved in the resolved result, regardless of group or checkbox order.

  • selection – {kind: [group name, ...]} – the groups ticked.

  • explicit – individual columns ticked, which are added to whatever the groups select. Both halves of the request are the same list in the end, so a user can tick “intensity” and then add one morphology column without the two mechanisms fighting.

Returns:

the columns, in the order they appear in columns, so a reduction’s input order does not depend on which checkbox was clicked first – a UMAP whose axes depend on click order is not reproducible.

spacr.column_groups.summarise(columns: Iterable[str], selection: Mapping[str, Sequence[str]] | None = None, *, explicit: Sequence[str] = ()) → str[source]

One line saying what is selected, for the picker to show.

A reduction over 400 columns and one over 4 look identical in a dialog until something says which it is.

Parameters:

columns – available column names whose selected fraction is reported.