"""Regression, and the three modules that are the rest of the visit.
A regression run is not finished when the coefficients appear. Three
screens carry on from there, and none of them duplicates anything the
results panel does:
* **Volcano Explorer** is the only publication-figure path in spaCR. Its
volcano is a matplotlib render behind a 56-field style that can be
saved and reloaded as JSON, every called point labelled, colour or
shape driven by an annotation file merged on an inferred key, arbitrary
x/y columns, and a vector re-render at journal size. The panel's own
volcano is pyqtgraph and can honour none of it, so the explorer is
offered from that plot's own menu as "Publication figure…", seeded with
the frame already on screen.
* **Hit List** displays one ranked row per gene for the loaded regression
run. Backends that report p-values receive Benjamini–Hochberg q-values
computed across the genes tested; penalised backends instead rank by
bootstrap selection frequency and do not report q-values. Each row includes
the effect estimate, a 95% interval when a standard error is available, and
gRNA sign agreement when guide-level coefficients are available. Annotation
CSVs are collapsed to one row per gene before validated many-to-one joins.
The complete :class:`~spacr.qt.screens.hit_list.HitListScreen` is installed
as the **Hits** tab after Guide support, follows the run loaded by the
results panel, and exports the displayed filtered list as CSV, Markdown, or
self-contained HTML.
* **Methods & Results** builds the run digest -- package versions,
timings, seed and error policy, per-module parameters parsed out of the
emitted macro, the segmentation verdict, artifact counts, held-out
metrics -- drafts the two sections from it and then mechanically checks
every number in the draft back against the digest. It opens seeded with
the project the regression screen is already pointed at, so the path it
otherwise asks the user to type is filled in.
Each of the three is also a button on the Regression masthead: the
module's own icon with no text, its one-line description as the tooltip,
lit on hover in the maturity colour its tile used -- see
:class:`spacr.qt.widgets.fold_strip.FoldStrip`. The button and the place
the capability lives are the same door: the Hit List button raises the
Hits tab rather than opening a second hit list, and the Volcano Explorer
button opens exactly what "Publication figure…" opens.
ONE CORRECTION FAMILY ON ONE VOLCANO. A guide permutation writes every
minimum-support family stacked into one long table, and fits one family
per response when several outcomes were fitted. Each of those is its own
Benjamini-Hochberg family, so drawing the table unfiltered puts a guide on
the plot two to four times at different heights and pools two corrections
into one picture. :func:`single_correction_family` is the cut, and
:func:`install_correction_families` puts it in front of the panel's own
``set_frame`` so every route into the panel -- a finished run, a folder
opened by hand, a dropped bundle -- draws one family.
The shared half of a fold -- opening a module in a window, wiring the host
signals a sidebar row used to wire, and hanging the strip off the
masthead -- lives in :mod:`spacr.qt.screens.map_barcodes` and is imported
rather than repeated.
"""
from __future__ import annotations
import logging
import os
from functools import partial
from typing import Callable, Dict, Optional, Tuple
from PySide6.QtCore import QUrl
from PySide6.QtGui import QDesktopServices
from PySide6.QtWidgets import QWidget
from ..i18n import tr
from ..widgets.fold_strip import FoldStrip
from .map_barcodes import (FoldOpener, build_registered_screen,
restate_fold_button)
LOG = logging.getLogger(__name__)
#: Registry key of the screen this module hangs its strip on.
HOST_KEY = "regression"
#: The fold key for the diagnostics page. Not a registered app -- there is
#: no module behind it, only the panels a finished run wrote -- so it takes
#: a key of its own and `fold_description` answers for it from
#: `FOLD_FALLBACK` below.
DIAGNOSTICS_KEY = "regression_diagnostics"
#: Where `spacr.ml._write_regression_diagnostics` puts its panels, relative
#: to the run's results folder. One spelling, shared with the writer, so the
#: button cannot look somewhere the writer does not use.
DIAGNOSTICS_DIRNAME = "diagnostics"
#: Registry keys of the modules folded into it, in the order the strip
#: draws them: the figure, then the list, then the write-up -- which is
#: the order the three are wanted in after a run finishes.
FOLDED_APPS: Tuple[str, ...] = ("volcano_explorer", "hit_list",
"methods_export", "investigate_hit",
"profiler", DIAGNOSTICS_KEY)
#: What the "Hits" tab is called, and the tab it is inserted after.
HITS_TAB_TITLE = "Hits"
HITS_TAB_AFTER = "Guide support"
#: The directory a regression writes its runs into, inside the project
#: -- ``<project>/results/<score>/<kind>``. Named so
#: :func:`project_path` can walk back up to the project the digest
#: wants without a second spelling of the layout.
RESULTS_DIRNAME = "results"
#: The entry the regression volcano's own right-click menu grows.
PUBLICATION_FIGURE_LABEL = "Publication figure…"
#: The heading that entry sits under. Its own section, like the re-fit
#: entry: everything above restyles the plot on screen, and this hands the
#: same rows to a different renderer.
PUBLICATION_FIGURE_SECTION = "Publication figure"
[docs]
def single_correction_family(frame):
"""``frame`` reduced to ONE multiple-testing family, ready to plot.
A guide permutation fits the same guides once per minimum-support
threshold and once per response, stacks the lot into one long table,
and corrects each stack separately. Two things follow, and both are
wrong on a plot: the same guide is drawn two to four times at
different heights, and two Benjamini-Hochberg corrections share one
axis -- so a q-value read off the picture belongs to whichever family
the point came from, which nothing on the picture says.
The primary family is the smallest ``minimum_wells_threshold``, which
is what ``perform_regression`` writes ``results.csv`` from when
``guide_primary_min_wells`` is left blank; where several responses
were fitted the first is kept, because each response is its own
correction family and the explorer's column controls can switch.
A frame with neither column is returned unchanged, which is every
parametric run -- so this costs an ordinary table nothing.
:param frame: a coefficient table, or None.
:returns: the same object when there was nothing to cut, otherwise a
new frame holding one family.
"""
if frame is None:
return frame
columns = list(getattr(frame, "columns", ()))
if not columns or not len(frame):
return frame
cut = frame
if "minimum_wells_threshold" in columns:
thresholds = cut["minimum_wells_threshold"].dropna().unique()
if len(thresholds) > 1:
cut = cut.loc[cut["minimum_wells_threshold"] == min(thresholds)]
if "outcome" in columns and cut["outcome"].nunique() > 1:
cut = cut.loc[cut["outcome"] == cut["outcome"].iloc[0]]
if cut is frame:
return frame
return cut.reset_index(drop=True)
[docs]
def install_correction_families(panel) -> bool:
"""Put :func:`single_correction_family` in front of ``panel.set_frame``.
Every route into the results panel ends in ``set_frame`` -- a finished
run, a folder opened through "Load results…", a dropped bundle, a
frame handed straight in -- so one wrapper there is what makes the
volcano show one correction family whichever way the table arrived.
Idempotent, so installing twice does not stack two cuts.
:param panel: the :class:`RegressionResultsPanel`, or None.
:returns: True when this call installed the cut.
"""
if panel is None or getattr(panel, "_one_correction_family", False):
return False
original = getattr(panel, "set_frame", None)
if not callable(original):
return False
def set_frame(frame, source: str = "") -> bool:
"""Narrow the frame to one correction family before setting it."""
return original(single_correction_family(frame), source=source)
panel.set_frame = set_frame
panel._one_correction_family = True
return True
def _run_folder(panel) -> str:
"""The run folder behind ``panel``, or "" when it came from nowhere."""
reader = getattr(panel, "run_folder", None)
if not callable(reader):
return ""
try:
return str(reader() or "")
except Exception:
LOG.debug("Could not read the panel's run folder", exc_info=True)
return ""
[docs]
def install_hits_tab(panel):
"""Add the Hit List to ``panel`` as a tab, beside Guide support.
The whole screen goes in, not a copy of its table: the filter bar, the
metadata picker and the three export buttons ARE the capability the
panel has none of, and a tab that reimplemented the list would keep
whichever parts the person doing the folding thought of.
The tab follows the panel: whenever a run is loaded the hit list is
pointed at the same folder, so it is never showing one run's hits
beside another run's coefficients.
Idempotent.
:param panel: the results panel.
:returns: the :class:`~spacr.qt.screens.hit_list.HitListScreen`, or
None when the tab could not be built.
"""
if panel is None:
return None
existing = getattr(panel, "hits", None)
if existing is not None:
return existing
tabs = getattr(panel, "tabs", None)
if tabs is None:
return None
from .hit_list import HitListScreen
try:
hits = HitListScreen(parent=panel, folder=_run_folder(panel))
except Exception:
LOG.exception("Could not build the Hits tab")
return None
index = tabs.count()
for position in range(tabs.count()):
if tabs.tabText(position) == HITS_TAB_AFTER:
index = position + 1
break
tabs.insertTab(index, hits, HITS_TAB_TITLE)
tabs.setTabToolTip(
tabs.indexOf(hits),
"One row per gene rather than per coefficient: the effect with its "
"95% interval, a q-value recomputed over the genes actually tested, "
"how many of the gene's own guides agree in sign, and any annotation "
"CSV joined on. Export the exact list on screen as CSV, Markdown or "
"a self-contained HTML page.")
panel.hits = hits
loaded = getattr(panel, "loaded", None)
if loaded is not None:
try:
loaded.connect(lambda _path, p=panel: _follow_the_run(p))
except Exception:
LOG.debug("The Hits tab could not follow the panel", exc_info=True)
return hits
def _follow_the_run(panel) -> None:
"""Point the Hits tab at whatever run the panel just loaded."""
hits = getattr(panel, "hits", None)
folder = _run_folder(panel)
if hits is None or not folder:
return
try:
hits.load_folder(folder)
except Exception:
LOG.debug("The Hits tab could not follow the run", exc_info=True)
[docs]
def raise_hits_tab(panel) -> bool:
"""Bring the Hits tab to the front. False when there is none.
:param panel: regression results panel carrying ``hits`` and ``tabs``
attributes; None, or a panel missing either, returns False.
"""
hits = getattr(panel, "hits", None) if panel is not None else None
tabs = getattr(panel, "tabs", None) if panel is not None else None
if hits is None or tabs is None:
return False
tabs.setCurrentWidget(hits)
return True
[docs]
def project_path(screen) -> str:
"""Resolve the project directory associated with a regression screen.
The function first examines the run loaded in the results panel and
returns the parent of its nearest ``results`` directory. If no project can
be resolved from that run, it reads ``src`` from the screen's settings
model. A sequence-valued ``src`` contributes its first entry.
:param screen: Regression screen or compatible host.
:returns: Project path, or ``""`` when it cannot be determined.
"""
if screen is None:
return ""
folder = _run_folder(_results_panel_if_built(screen))
while folder and folder != os.path.dirname(folder):
parent = os.path.dirname(folder)
if os.path.basename(folder) == RESULTS_DIRNAME:
return parent
folder = parent
model = getattr(screen, "_settings_model", None)
collect = getattr(model, "collect", None) if model is not None else None
if not callable(collect):
return ""
try:
source = (collect() or {}).get("src")
except Exception:
LOG.debug("Could not read the regression source", exc_info=True)
return ""
if isinstance(source, (list, tuple)):
source = source[0] if source else ""
source = str(source or "").strip()
if not source:
return ""
return os.path.abspath(os.path.expanduser(source))
[docs]
def build_methods_export(host_window: Optional[QWidget] = None,
screen: Optional[QWidget] = None) -> QWidget:
"""Methods & Results' own screen, seeded with the run it will describe.
Both sources the regression screen knows are filled in: the project,
which supplies the provenance summary and the segmentation verdict,
and the results folder, which supplies the hit statistics. The other
two -- the run journal and the classifier checkpoint -- are the user's
to name, and every one of them is optional on that screen.
"""
from .methods_export import MethodsExportScreen
panel = results_panel(screen)
return MethodsExportScreen(project=project_path(screen),
results_folder=_run_folder(panel))
[docs]
def results_panel(screen):
"""``screen``'s results panel, or None on a screen that has none.
:param screen: app screen whose ``_results_panel`` attribute is returned;
None is accepted.
"""
return getattr(screen, "_results_panel", None) if screen is not None \
else None
def _results_panel_if_built(screen):
"""``screen``'s results panel if it exists yet; never builds one.
A Regression screen builds its results panel on first use. A panel
that is not built has no run loaded, so a question about the loaded
run gets the same answer from its absence.
"""
probe = getattr(screen, "_results_panel_if_built", None)
if callable(probe):
return probe()
return results_panel(screen)
#: One builder per folded module that opens a WINDOW. Each takes the main
#: window and the host screen; :func:`install_folds` binds the screen, so a
#: builder still has the one-argument shape
#: :class:`spacr.qt.screens.map_barcodes.FoldOpener` calls it with.
#:
#: ``hit_list`` is not here: its home is a tab on the results panel, so its
#: button raises that tab rather than opening anything -- see
#: :class:`HitsOpener`, which falls back to a window on a screen that has no
#: panel to hold a tab.
def _build_investigate_hit(host_window: Optional[QWidget] = None,
screen: Optional[QWidget] = None
) -> Optional[QWidget]:
"""Investigate Hit, as the window builds it.
Folded here because a hit is the thing a regression produces: the
module answers "which cells are behind this row", which is a question
you can only have once a regression has run. It had a Home tile,
where it read as a place to start.
"""
return build_registered_screen("investigate_hit", host_window)
def _build_profiler(host_window: Optional[QWidget] = None,
screen: Optional[QWidget] = None) -> Optional[QWidget]:
"""Prediction Profiler, as the window builds it.
Same reasoning: it profiles a FITTED model's response, so it belongs
behind the module that fits one rather than beside it on Home.
"""
return build_registered_screen("profiler", host_window)
BUILDERS: Dict[str, Callable[..., Optional[QWidget]]] = {
"volcano_explorer": open_publication_figure,
"methods_export": build_methods_export,
"investigate_hit": _build_investigate_hit,
"profiler": _build_profiler,
}
#: What the diagnostics button says, since no registry row answers for it.
#: `fold_description` reads the registry, then the declared catalogue, then
#: the shared fold records -- and this is in none of them, because it is a
#: view of files a run wrote rather than a module.
FOLD_FALLBACK: Dict[str, Tuple[str, str, str]] = {
DIAGNOSTICS_KEY: (
"Diagnostics",
"Show the diagnostic panels the last regression run wrote beside "
"its results",
"alpha"),
}
[docs]
class DiagnosticsOpener:
"""The Diagnostics button: show the panels the last run wrote.
NOT A MODULE. `spacr.ml` writes these panels into the run's results
folder as it finishes -- see `_write_regression_diagnostics` -- so
there is nothing to compute here and nothing to configure. The button
opens what is already on disk.
Which is also why it can be pressed when there is nothing to show. A
run that has not happened, or one whose backend has no residuals, has
no folder or a partial one, and saying so is more useful than a
button that appears to do nothing.
:param screen: the host screen the button sits on.
"""
key = DIAGNOSTICS_KEY
def __init__(self, screen: QWidget) -> None:
"""Record the screen the diagnostics page is opened on.
:param screen: the host screen.
"""
self.screen = screen
def _folder(self) -> Optional[str]:
"""The diagnostics folder of the run on screen, if it is there."""
root = project_path(self.screen)
if not root:
return None
candidate = os.path.join(str(root), RESULTS_DIRNAME)
if not os.path.isdir(candidate):
return None
newest, newest_at = None, -1.0
for base, dirs, _files in os.walk(candidate):
if DIAGNOSTICS_DIRNAME in dirs:
found = os.path.join(base, DIAGNOSTICS_DIRNAME)
stamped = os.path.getmtime(found)
if stamped > newest_at:
newest, newest_at = found, stamped
return newest
[docs]
def verdict(self) -> tuple:
"""``(level, detail)`` for the badge, from what is on disk.
The button is badged with the worst of score_design /
score_residuals / score_inference. That worst is
already computed and written -- ``diagnostic_summary.csv`` carries
a ``suite``/``verdict_level`` row -- so this is a READ, and it has
to stay one. Recomputing it here would let the button and the
panels disagree about the same run.
``("", "")`` when there is nothing to read, which is not the same
as a pass: a run that has not happened has no verdict, and a green
dot would claim it did and was fine.
"""
folder = self._folder()
if folder is None:
return "", ""
level = detail = ""
try:
import csv
for name in sorted(os.listdir(folder)):
if not name.startswith("diagnostic_summary"):
continue
with open(os.path.join(folder, name), newline="",
encoding="utf-8") as handle:
for row in csv.DictReader(handle):
if row.get("section") != "suite":
continue
if row.get("metric") == "verdict_level":
level = str(row.get("value") or "")
elif row.get("metric") == "verdict":
detail = str(row.get("value") or "")
except Exception: # noqa: BLE001
LOG.debug("could not read the diagnostics summary", exc_info=True)
return "", ""
return level, detail
[docs]
def open(self) -> None:
"""Open the folder, or say why there is nothing to open."""
from PySide6.QtWidgets import QMessageBox
folder = self._folder()
if folder is None:
QMessageBox.information(
self.screen, tr("Diagnostics"),
tr("This project has no regression diagnostics yet. They "
"are written when a regression finishes."))
return
note = os.path.join(folder, "residual_panels_not_available.txt")
if os.path.isfile(note):
try:
with open(note, encoding="utf-8") as handle:
QMessageBox.information(
self.screen, tr("Diagnostics"), handle.read().strip())
except Exception: # noqa: BLE001
LOG.debug("could not read the diagnostics note",
exc_info=True)
QDesktopServices.openUrl(QUrl.fromLocalFile(folder))
[docs]
class HitsOpener:
"""The Hits button: raise the tab, or open the module when there is none.
ONE HIT LIST, not two. Where the results panel exists the list is
already a tab on it, loaded with the run on screen, so the button goes
there -- and it brings the Results page forward first, because a button
that silently changed a tab behind the settings form would read as a
button that does nothing.
A bare settings screen has no panel and so no tab; then the module
opens in a window of its own, like every other fold, with the run
folder seeded.
:param screen: the host screen the button sits on.
"""
key = "hit_list"
def __init__(self, screen: QWidget) -> None:
"""Record the screen the hit list is opened on, and how to build it.
:param screen: the host screen; the page itself is built lazily, on
first open.
"""
self.screen = screen
self._window = FoldOpener(screen, self.key, self._build_window)
def _build_window(self, host_window) -> QWidget:
"""The Hit List module itself, seeded with the run on screen."""
from .hit_list import HitListScreen, connect_investigation
hits = HitListScreen(folder=_run_folder(results_panel(self.screen)))
connect_investigation(hits, host_window)
return hits
[docs]
def open(self, _checked: bool = False) -> Optional[QWidget]:
"""Show the hit list, wherever this screen keeps it."""
panel = results_panel(self.screen)
raise_results = getattr(self.screen, "_raise_the_results_tab", None)
if panel is not None and callable(raise_results):
try:
raise_results()
except Exception:
LOG.debug("Could not raise the results page", exc_info=True)
if raise_hits_tab(panel):
return panel.hits
return self._window.open()
[docs]
def publication_opener(screen) -> FoldOpener:
"""Return the shared Volcano Explorer opener for ``screen``.
The opener is created on first use and retained on the screen. Both the
**Publication figure…** command and the masthead action use this instance,
so reopening the explorer raises the existing window rather than creating
a duplicate with independent state.
:param screen: Regression screen that owns the publication workflow.
:returns: Persistent :class:`~spacr.qt.screens.map_barcodes.FoldOpener`.
"""
opener = getattr(screen, "_publication_opener", None)
if not isinstance(opener, FoldOpener):
opener = FoldOpener(screen, "volcano_explorer",
partial(open_publication_figure, screen=screen))
screen._publication_opener = opener
return opener
[docs]
def install_folds(screen: QWidget) -> Optional[FoldStrip]:
"""Put Regression's fold strip on ``screen``'s masthead.
Built here rather than through
:func:`spacr.qt.screens.map_barcodes.install_fold_strip` because one of
the three buttons does not open a window: the Hits button raises a tab
on the screen the user is already looking at.
Idempotent, and defensive by design: a screen that opens without its
fold buttons is a smaller screen, while an exception raised here would
be no regression screen at all.
:param screen: app screen whose masthead receives the strip; it must have
``app_key`` ``"regression"`` and a ``_header`` with ``add_trailing``.
:returns: the strip, or None when this screen cannot carry one -- it is
not the host, it has no masthead, or one is already installed.
"""
if getattr(screen, "app_key", None) != HOST_KEY:
return None
existing = getattr(screen, "_fold_strip", None)
if isinstance(existing, FoldStrip):
return existing
header = getattr(screen, "_header", None)
if header is None or not hasattr(header, "add_trailing"):
return None
openers = []
for key in FOLDED_APPS:
if key == DiagnosticsOpener.key:
openers.append(DiagnosticsOpener(screen))
elif key == HitsOpener.key:
openers.append(HitsOpener(screen))
elif key == "volcano_explorer":
openers.append(publication_opener(screen))
else:
openers.append(FoldOpener(screen, key,
partial(BUILDERS[key], screen=screen)))
install_extras(screen)
try:
strip = FoldStrip([(o.key, o.open) for o in openers], header)
for opener in openers:
button = strip.button_for(opener.key)
restate_fold_button(button, opener.key)
if hasattr(opener, "verdict") and hasattr(button, "set_verdict"):
try:
level, detail = opener.verdict()
button.set_verdict(level, detail)
except Exception: # noqa: BLE001
LOG.debug("could not badge the diagnostics button",
exc_info=True)
header.add_trailing(strip)
except Exception:
LOG.debug("Could not build the regression fold strip", exc_info=True)
return None
screen._fold_openers = openers
screen._fold_strip = strip
return strip