spacr.qt.screens.pca

Workflow inputs and outputs

PCA

Inspect principal components and feature loadings for the selected measured objects.

Open: Image UMAP → PCA.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route: cell, nucleus, pathogen, cytoplasm. Relevant columns, depending on the route: plateID, rowID, columnID, fieldID.

Outputs

  • Projection and clusters — Image UMAP/PCA coordinate tables, selected clusters and figures for the loaded measurement data.

  • Figures and table exports — The output location chosen by the tool; exports describe the selected data and filters.

API reference.

Module tutorial.

The PCA screen — hundreds of features, two axes, and which ones did it.

Assembles four things that already exist:

The one thing worth knowing before reading a chart from this screen is in spacr.qt.widgets.pca_model: features are standardised by default, because cell_area in px² would otherwise be PC1 of every table in the project; NaN is never quietly imputed, because a pathogen_* NaN means “no pathogen” and not “value unknown”; and the report under the plot always says how much of the table the picture is actually about.

Filter, then decompose

A filter change recomputes the decomposition rather than just re-drawing it, and that is the interesting decision on this screen. PCA is a property of a population: the centre, the scale and the component directions are all computed from the rows in it. Keeping the old components and dropping the filtered points would draw a plot whose axes belong to a population the user is no longer looking at — the scores would still be in the old basis, and a cluster that separates in the filtered subset would not appear. So the filter is upstream of the maths, the recomputation is debounced, and the report under the plot re-states the row count every time.

register() is not called at import; see its docstring, and the same note on spacr.qt.screens.graph_builder.register(), for the registration collateral that is still owned by app.py.

Classes

PCAScreen

Load a measurement table, decompose it, and brush the result.

Functions

make_pca_screen(→ PySide6.QtWidgets.QWidget)

Build the screen. The one constructor every caller goes through.

Module Contents

class spacr.qt.screens.pca.PCAScreen(parent=None, *, link=None, threaded: bool = True)[source]

Bases: PySide6.QtWidgets.QWidget

Load a measurement table, decompose it, and brush the result.

Parameters:
  • link – a private LinkedSelection for tests. None joins the process-wide one, which is the point of the screen in normal use.

  • parent – parent widget; ownership only.

  • threaded – False runs every table read inline instead of on the job runner’s thread. A TEST NEEDS THE RESULT ON THE LINE AFTER THE CALL; a user needs the window to keep painting while a large table loads. The jobs are the same either way – they still register, still report failure through job_failed – so only the waiting differs.

Build the screen: the PCA panel beside the filter, with a re-filter timer.

The filter sits upstream of the maths, so the screen listens for it itself rather than leaving the canvas to redraw components computed on rows the filter has since removed.

Parameters:
  • parent – parent widget, or None.

  • link – shared selection link, passed to the panel and the filter so both answer to the same selection.

  • threaded – run reads and the fit on a worker thread. Set False in tests so a load finishes before it returns.

active_jobs() → int[source]

How many worker threads are still winding down.

The panel’s own count is included: the read and the decomposition are two jobs a caller cannot tell apart, and a waitUntil(active_jobs() == 0) that stopped at the read would return before the plot exists.

choose_table() → None[source]

Ask which table in the project to use.

closeEvent(event)[source]

Stop background work and unlink before going away.

Parameters:

event – the Qt close event.

export_csv() → None[source]

Write scores, loadings and explained variance beside each other.

Three files rather than one, because they have three different row meanings — an object, a feature and a component — and a single sheet that mixed them would have to be unpicked before anyone could use it.

is_busy() → bool[source]

True while a table read or a decomposition is in flight.

load_path(path: str, table: str | None = None) → None[source]

Load a CSV or one table of a SQLite measurement database.

The read runs on a worker thread. SELECT * FROM cell into pandas measures 1.5 s for a 200 000-row measurement table on a warm local SSD, and this method used to run it inline: the whole window stopped redrawing for the read. Listing the table names stays inline – it is one sqlite_master query, measured at 0.4 ms – because the picker has to be populated before the read is dispatched, to know which table to read.

Returns as soon as the read is dispatched; _on_frame_loaded() finishes on the GUI thread.

Parameters:

path – CSV, TSV or TXT file, or a SQLite measurement database whose table names fill the table picker.

set_frame(frame: pandas.DataFrame, *, label: str = '') → None[source]

Decompose frame. The one call a host needs.

Parameters:

frame – measurement table, one row per object; it is passed through the screen’s filters before the decomposition.

spacr.qt.screens.pca.make_pca_screen(app_key: str | None = None) → PySide6.QtWidgets.QWidget[source]

Build the screen. The one constructor every caller goes through.

app_key is accepted and ignored: it is the shape spacr.qt.app.register_app() called a factory with, and callers written against that shape still work.