Source code for spacr.anndata_export

"""Export a spaCR ``measurements.db`` as AnnData (``.h5ad``).

spaCR already produces the exact shape AnnData was designed for: N objects
x M features, plus per-object metadata and an embedding. Writing that out as
``.h5ad`` puts scanpy, scvi-tools and squidpy within reach of a spaCR user
for the cost of one function call, instead of a bespoke join script per lab.

Nothing here invents an identity, a filter or a feature/metadata boundary.
All four already exist in spaCR and are reused verbatim:

``obs_names``
    :func:`spacr.selection.object_keys` over
    :data:`spacr.selection.OBJECT_KEY_COLUMNS` -- the schema's own row key
    (``plateID``, ``rowID``, ``columnID``, ``fieldID``, ``object_label``),
    joined on :data:`spacr.schema.KEY_SEPARATOR`. A key from a UMAP lasso
    names the same object in an exported ``.h5ad``.
``X`` / ``obs`` split
    :func:`spacr.schema.is_provenance_column`, which is the boundary
    :mod:`spacr.feature_dict` and every model path already use, plus
    :mod:`spacr.agreement`'s knowledge of which columns a *model* wrote and
    which a *human* did.
``var``
    :func:`spacr.feature_dict.parse_column`, resolved against the row's own
    ``measurement_units`` so a 3-D run's ``cell_area`` is documented as the
    volume it is rather than the area it is named after.
Filtering
    :class:`spacr.selection.DataFilter` and :class:`spacr.selection.Selection`
    -- the same objects the GUI's linked views publish, so "export what I am
    looking at" is one call and not a second filter language.

Typical use::

    from spacr.anndata_export import export_anndata
    from spacr.selection import DataFilter, RangeFilter

    result = export_anndata(
        "/data/exp1/measurements/measurements.db",
        "/data/exp1/results/exp1.h5ad",
        data_filter=DataFilter().add(RangeFilter("cell_area", low=200)),
    )
    print(result.describe())

    import scanpy as sc
    adata = sc.read_h5ad("/data/exp1/results/exp1.h5ad")

What happens to NaN
-------------------
**Nothing, by default, and that is a decision rather than an omission.**

AnnData stores NaN happily. Most of what people reach for next does not:
``sc.pp.scale``, ``sc.pp.pca``, ``sc.pp.neighbors`` and essentially all of
scvi-tools either propagate NaN across the whole matrix or raise. So a
silent default matters, and there are only two honest candidates: keep the
NaN and say so loudly, or impute and say so loudly. Quietly filling with
zero is not one of them -- in spaCR a NaN is usually *meaningful*: a
``pathogen_*`` column is NaN for a cell with no pathogen in it, a Zernike
column is NaN when mahotas was not installed, and a correlation column is
NaN when one channel was flat. Zero-filling the first turns "no pathogen"
into "a pathogen of zero size" and puts it into the same distribution as a
measured one.

:data:`NAN_KEEP` (the default) therefore writes the NaN through, and pays
for it with *visibility*:

* ``var['n_missing']`` / ``var['frac_missing']`` -- per feature;
* ``obs['n_missing_features']`` -- per object;
* ``uns['spacr']['nan']`` -- totals, the policy that was applied, the shape
  those totals were counted over (``n_objects_counted`` x
  ``n_features_counted``, i.e. post-filter and pre-policy), and the ten
  worst columns by name;
* a printed warning naming those columns and the scanpy calls that will
  fail on them.

Every count in that record -- and :attr:`ExportResult.n_missing` with it --
is taken on the matrix **as the policy received it**, which for a dropping
policy is bigger than the one written. ``ExportResult.frac_missing`` divides
by that same matrix (:attr:`ExportResult.counted_shape`) rather than by the
written shape, so it stays a fraction.

The alternatives are explicit, and every one of them records what it did:

``NAN_DROP_FEATURES``
    drop any feature column containing a NaN. Cheap and safe; on a real
    database it often removes every ``pathogen_*`` column, which is the
    honest consequence of asking for a complete matrix.
``NAN_DROP_OBJECTS``
    drop any object row containing a NaN. Usually removes almost
    everything, for the same reason; offered because on a
    single-object-type export it is often exactly right.
``NAN_ZERO`` / ``NAN_MEAN``
    impute. Both write a ``layers['missing']`` boolean mask of the same
    shape by default, because an imputed matrix that cannot be told from a
    measured one is a trap, and 1 byte per cell against 4 is a fair price.

**Infinities are treated as missing under every policy.** spaCR produces
+/-inf from ratio features (a denominator of zero), and an inf survives
``dropna`` while destroying any scaling, PCA or distance computed from it.
They are converted to NaN before the policy runs and counted separately in
``uns['spacr']['nan']['n_infinite']`` so the substitution is never silent.

Where the cell -> nucleus -> pathogen links go, and why not ``obsp``
--------------------------------------------------------------------
They go in ``obs`` and ``uns``. **Never** ``obsp``.

``obsp`` is an ``n_obs x n_obs`` matrix: a relation among *the observations
of this AnnData*. The parent links are not that, in either export shape:

* In the joined, cell-anchored export (the default) the nuclei and pathogens
  are not observations at all -- the join in
  :func:`spacr.io._read_and_join_tables` has already averaged them onto
  their parent cell. There is no row for an ``obsp`` entry to point at.
* In a per-table export of ``nucleus``, the *cells* are not observations,
  for the mirror-image reason.
* An ``obsp`` link would only be well defined for one AnnData whose ``obs``
  is the disjoint union of cells, nuclei and pathogens -- and that matrix is
  useless downstream, because ``X`` would be block-structured by
  construction (a nucleus row has no ``cell_area``) and every tool would see
  it as ~60% missing data.

What is recorded instead is exact and survives everything ``obsp`` does not
(subsetting, concatenation, an ``h5ad`` round trip through a tool that
knows nothing about spaCR):

* joined export -- ``obs['count_nucleus']`` / ``obs['count_pathogen']``
  (how many children were averaged into this row) and
  ``var['is_aggregated']`` (which columns are such an average). Without the
  second one it is genuinely easy to report a per-nucleus statistic that is
  in fact a per-cell mean of nuclei.
* per-table export of a child -- ``obs['cell_id']``, the schema's own
  ``parent_column`` from :data:`spacr.schema.OBJECT_TABLE_SCHEMAS`. A plain
  foreign key, which is what ``adata.obs.groupby('cell_id')`` already
  expects.
* both -- ``uns['spacr']['relationships']``, stating the parent table, the
  key it is joined on, whether the child features were aggregated, and (for
  :func:`export_anndata_set`) the sibling ``.h5ad`` file holding the parent.

The labels the table join drops
-------------------------------
:func:`spacr.io._read_and_join_tables` takes six named columns off
``png_list`` -- the object id, ``png_path`` and the four field keys -- and
drops the rest. The rest is every annotation column the Annotate app added
and every score the classifier wrote, which for this export are among the
most valuable columns in the database: without them ``obs`` has no label to
train on, group by, or colour a UMAP with. They are re-attached by object
key (:func:`_attach_png_labels`) *before* filtering, so a filter can name
one -- "export the cells I annotated as infected" -- and they are separated
into ``uns['spacr']['annotation_columns']`` and
``uns['spacr']['prediction_columns']`` using :mod:`spacr.agreement`'s own
rule, which exists because a model's class column is indistinguishable *by
shape* from an annotation pass.

They land in ``obs``, never in ``X``: a label is not a measurement, and
scaling, PCA and clustering it alongside the features is how a classifier's
own output ends up as a "feature" that predicts it.

Provenance
----------
``uns['spacr']`` carries the spaCR version, the settings hash
(:func:`spacr.artifacts.settings_hash`), the run id, the absolute source
database path, the tables read, the filter that was applied and the counts
before and after it. The written file is also registered with
:mod:`spacr.artifacts` under kind :data:`ANNDATA_KIND`, with the project's
``measurements-db`` artifact as its input -- so ``spacr.artifacts.is_stale``
reports the export as stale the moment Measure is re-run, which is the
whole reason the registry exists.

The optional dependency
-----------------------
``anndata`` is an optional extra: ``pip install "spacr[anndata]"``. It is
imported inside the functions that need it, never at module scope, so
``import spacr.anndata_export`` works without it and the failure is one
actionable sentence (:exc:`AnnDataExtraMissing`) rather than a traceback
six frames deep.
"""
from __future__ import annotations

import os
import sqlite3
import sys
import tempfile
import warnings
from dataclasses import dataclass, replace as _dataclass_replace
from datetime import datetime, timezone
from typing import Any, Dict, List, Mapping, Optional, Sequence, Tuple, Union

import numpy as np
import pandas as pd

from .. import schema
from ..selection import (OBJECT_KEY_COLUMNS, OBJECT_TYPE_COLUMN, DataFilter,
                         FilterError, Selection, object_keys,
                         untyped_object_key, with_object_type)

__all__ = [
    "ANNDATA_EXTRA",
    "ANNDATA_KIND",
    "ANNDATA_MISSING_MESSAGE",
    "APP_KEY",
    "AnnDataExtraMissing",
    "CONDITION_FALLBACK",
    "DEFAULT_CONDITION_MAP",
    "DEFAULT_TABLES",
    "DuplicateObjectKeys",
    "ExportResult",
    "NAN_DROP_FEATURES",
    "NAN_DROP_OBJECTS",
    "NAN_KEEP",
    "NAN_MEAN",
    "NAN_POLICIES",
    "NAN_ZERO",
    "anndata_export_settings",
    "build_anndata",
    "default_out_path",
    "export_anndata",
    "export_anndata_set",
    "feature_columns",
    "register_anndata_settings",
    "require_anndata",
    "resolve_db_path",
    "run_anndata_export",
]



#: The ``setup.py`` extra that provides :mod:`anndata`.
ANNDATA_EXTRA = "anndata"

#: What the user is told when the extra is not installed. One sentence of
#: diagnosis and one command, following :data:`spacr.qt._QT_MISSING_MESSAGE`.
[docs] ANNDATA_MISSING_MESSAGE = """\ Exporting to AnnData (.h5ad) needs the optional `anndata` extra, which is not installed in this environment (missing module: {module}). Install it with: python -m pip install "spacr[anndata]" scanpy is a separate install and is NOT required to write the file: python -m pip install scanpy\ """
[docs] class AnnDataExtraMissing(ImportError): """:mod:`anndata` is not installed. An :class:`ImportError` subclass so a caller that already guards the export with ``except ImportError`` keeps working, and so the message -- not a traceback through ``anndata``'s own import machinery -- is what reaches the user. """
[docs] class DuplicateObjectKeys(ValueError): """Two rows claim the same object key. AnnData requires unique ``obs_names``, and spaCR's writers append: two rows *can* share all five key columns when a field was measured twice (see ``tests/test_db_contract.py``, which measures exactly that). This is raised rather than repaired, because deduplicating -- keeping the first, keeping the last, or averaging -- changes the numbers and is the caller's decision, not the exporter's. """
[docs] def require_anndata(): """Import and return :mod:`anndata`, or raise a message worth reading. :returns: the imported :mod:`anndata` module. :raises AnnDataExtraMissing: when the extra is not installed. """ try: import anndata except ImportError as exc: module = (getattr(exc, "name", None) or "anndata").split(".", 1)[0] raise AnnDataExtraMissing( ANNDATA_MISSING_MESSAGE.format(module=module)) from exc return anndata
#: The artifact kind an exported ``.h5ad`` registers under. #: :data:`spacr.ports.ALL_KINDS` is documented as the built-in vocabulary #: rather than a closed set, so a new kind is a declaration, not a violation. ANNDATA_KIND = "anndata" #: Settings / app key. The key :func:`register_anndata_settings` files #: the defaults under, and the key the folded page on Measure is built #: with, so the panel and the export it runs agree. APP_KEY = "anndata_export" #: Tables read for the default, cell-anchored export. Exactly the list #: :func:`spacr.core.generate_image_umap` reads, so an export and an #: embedding describe the same population. DEFAULT_TABLES: Tuple[str, ...] = ( "cell", "cytoplasm", "nucleus", "pathogen", "png_list") #: Keep NaN in ``X`` and report it. The default; see the module docstring. NAN_KEEP = "keep" #: Drop every feature column containing a NaN. NAN_DROP_FEATURES = "drop_features" #: Drop every object row containing a NaN. NAN_DROP_OBJECTS = "drop_objects" #: Replace NaN with ``0.0``, recording a ``layers['missing']`` mask. NAN_ZERO = "zero" #: Replace NaN with the feature's mean, recording a ``layers['missing']`` mask. NAN_MEAN = "mean" #: Every accepted ``nan_policy``. NAN_POLICIES: Tuple[str, ...] = ( NAN_KEEP, NAN_DROP_FEATURES, NAN_DROP_OBJECTS, NAN_ZERO, NAN_MEAN) #: The column -> condition mapping :func:`spacr.utils.map_condition` applies #: by default. Spelled out here rather than imported: ``spacr.utils`` pulls #: torch and cellpose, which is a 10-second import for a four-entry dict, and #: this module is deliberately importable on a machine that cannot segment. #: ``tests/test_anndata_export.py`` pins the two definitions together. DEFAULT_CONDITION_MAP: Dict[str, str] = { "c1": "neg", "c2": "pos", "c3": "mix"} #: What a column not named in the condition map maps to -- the same fallback #: :func:`spacr.utils.map_condition` uses. CONDITION_FALLBACK = "screen" #: Columns that are numeric and are nonetheless not measurements: the #: cluster label :func:`spacr.core.generate_image_umap` writes back, and the #: aggregation counts :func:`spacr.io._read_and_join_tables` adds. These land #: in ``obs``. ``count_*`` is already caught by #: :func:`spacr.schema.is_provenance_column`; ``cluster`` is not. _NON_FEATURE_NUMERIC = frozenset({"cluster"}) #: Prefixes of columns this module adds to ``obs`` itself, reserved so a #: database column of the same name cannot silently overwrite one. _RESERVED_OBS = ("n_missing_features",) @dataclass(frozen=True)
[docs] class ExportResult: """What one export produced, and what it left out. :param path: the ``.h5ad`` written, or ``""`` for an in-memory build. :param n_obs: objects in the exported matrix. :param n_vars: feature columns in ``X``. :param n_obs_before_filter: objects read from the database. :param obs_columns: the ``obs`` column names. :param obsm_keys: the ``obsm`` keys written. :param nan_policy: the policy that was applied. :param n_missing: NaN cells in ``X`` *before* the policy ran. :param n_infinite: non-finite cells converted to NaN before that count. :param dropped_features: feature columns removed by the policy. :param dropped_objects: objects removed by the policy. :param n_obs_counted: rows of the matrix ``n_missing`` was counted in -- i.e. after ``data_filter``/``selection``/``row_limit`` and *before* the ``nan_policy`` ran. ``0`` on a record that did not record it, in which case :attr:`counted_shape` reconstructs it from the drops. :param n_vars_counted: columns of that same pre-policy matrix. :param artifact_id: the :mod:`spacr.artifacts` id, or ``""`` when the file was not registered. :param warnings: everything the export decided the user must know. Three object counts live here and they are three different stages, which is the whole reason they are named apart: :attr:`n_obs_before_filter` is what the database held, :attr:`counted_shape` is what survived the filter and was handed to the NaN policy, and :attr:`n_obs` is what was written. :meth:`describe` attributes each loss to the stage that caused it rather than charging all of them to the filter. """ path: str n_obs: int n_vars: int n_obs_before_filter: int obs_columns: Tuple[str, ...] = () obsm_keys: Tuple[str, ...] = () nan_policy: str = NAN_KEEP n_missing: int = 0 n_infinite: int = 0 dropped_features: Tuple[str, ...] = () dropped_objects: int = 0 n_obs_counted: int = 0 n_vars_counted: int = 0 artifact_id: str = "" warnings: Tuple[str, ...] = () @property
[docs] def counted_shape(self) -> Tuple[int, int]: """``(rows, columns)`` of the matrix :attr:`n_missing` was counted in. That matrix is the post-filter, **pre**-``nan_policy`` one: the NaN were counted before anything was dropped or imputed. A record built by this module states it outright; one built by hand (or by an older spaCR) leaves the two fields at ``0``, and the shape is reconstructed from the drops instead -- ``drop_objects`` is the only policy that removes rows and ``drop_features`` the only one that removes columns, so adding them back recovers what the policy was given. """ rows = self.n_obs_counted or (self.n_obs + self.dropped_objects) columns = (self.n_vars_counted or (self.n_vars + len(self.dropped_features))) return int(rows), int(columns)
@property
[docs] def frac_missing(self) -> float: """:attr:`n_missing` over the cells of :attr:`counted_shape`. Numerator and denominator describe **the same matrix** -- the one the NaN were counted in, before the policy dropped or imputed anything -- so this is a fraction and can never exceed 1.0. Dividing the pre-policy count by the post-policy ``n_obs * n_vars`` is what made ``drop_objects`` report 114.3% missing. :returns: the fraction, or ``0.0`` for an empty matrix. """ rows, columns = self.counted_shape cells = rows * columns return (self.n_missing / cells) if cells else 0.0
[docs] def describe(self) -> str: """One human paragraph: shape, filtering, missingness, provenance. Each count is charged to the stage that caused it: the filter line counts only what the filter removed, and the ``dropped ...`` lines only what the NaN policy removed. Charging the policy's row drops to the filter as well made one loss of four objects read as eight. """ counted_obs, counted_vars = self.counted_shape lines = [ f"{self.n_obs} objects x {self.n_vars} features" + (f" -> {self.path}" if self.path else " (in memory)") ] removed_by_filter = self.n_obs_before_filter - counted_obs if removed_by_filter > 0: lines.append( f" filtered from {self.n_obs_before_filter} objects " f"({removed_by_filter} removed)") if self.n_missing: over = ("" if (counted_obs, counted_vars) == (self.n_obs, self.n_vars) else f" of the {counted_obs} x {counted_vars} matrix " f"the policy was given") lines.append( f" {self.n_missing} missing values " f"({self.frac_missing:.1%}{over}), " f"policy {self.nan_policy!r}") if self.n_infinite: lines.append( f" {self.n_infinite} non-finite values treated as missing") if self.dropped_features: shown = ", ".join(self.dropped_features[:5]) more = (f" +{len(self.dropped_features) - 5}" if len(self.dropped_features) > 5 else "") lines.append( f" dropped features (nan_policy {self.nan_policy!r}): " f"{shown}{more}") if self.dropped_objects: lines.append( f" dropped objects (nan_policy {self.nan_policy!r}): " f"{self.dropped_objects}") if self.obsm_keys: lines.append(f" obsm: {', '.join(self.obsm_keys)}") if self.artifact_id: lines.append(f" artifact {self.artifact_id}") return "\n".join(lines)
def _available_tables(db_path: Union[str, os.PathLike]) -> Tuple[str, ...]: """Return the tables present in ``db_path``, in ``sqlite_master`` order. :param db_path: a ``measurements.db``. :returns: table names; empty when the file does not exist. """ if not os.path.isfile(os.fspath(db_path)): return () connection = sqlite3.connect(os.fspath(db_path), timeout=30) try: return tuple( row[0] for row in connection.execute( "SELECT name FROM sqlite_master WHERE type='table' " "AND name NOT LIKE 'sqlite_%' ORDER BY name")) finally: connection.close() def _read_frame(db_path: str, tables: Sequence[str], single_table: Optional[str], *, collapse_duplicate_identity: bool = True, ) -> Tuple[pd.DataFrame, Tuple[str, ...]]: """Read the object frame this export will describe. :param db_path: a ``measurements.db``. :param tables: tables to join, for the default cell-anchored export. :param single_table: read exactly this one object table instead, one row per object of that type. :returns: ``(frame, tables_actually_read)``. :raises ValueError: when nothing usable is in the database. """ present = set(_available_tables(db_path)) if not present: raise ValueError( f"{db_path} holds no tables; it is not a spaCR measurements " f"database. Point at <project>/measurements/measurements.db.") if single_table is not None: if single_table not in present: raise ValueError( f"table {single_table!r} is not in {db_path}. Present: " f"{sorted(present & set(schema.OWNED_TABLES))}.") connection = sqlite3.connect(db_path, timeout=30) try: frame = pd.read_sql(f'SELECT * FROM "{single_table}"', connection) finally: connection.close() return with_object_type(frame, single_table), (single_table,) wanted = [t for t in tables if t in present] if not wanted: raise ValueError( f"none of {list(tables)} is in {db_path}. Present: " f"{sorted(present & set(schema.OWNED_TABLES))}.") if "cell" not in wanted: raise ValueError( f"the joined export is anchored on the 'cell' table, which " f"{db_path} does not have. Pass single_table= to export one of " f"{sorted(t for t in wanted if t in schema.OBJECT_TABLES)} " f"on its own.") from ..io import _read_and_join_tables frame = _read_and_join_tables( db_path, list(wanted), collapse_duplicate_identity=collapse_duplicate_identity) if frame is None: raise ValueError( f"could not join {wanted} from {db_path}; the join returned " f"nothing. Run `spacr doctor --db {db_path}` for the diagnosis.") return with_object_type(frame, "cell"), tuple(wanted) def _attach_png_labels(frame: pd.DataFrame, db_path: str, anchor: str, *, timelapse: bool) -> Tuple[pd.DataFrame, List[str]]: """Bring the annotation and prediction columns back out of ``png_list``. :func:`spacr.io._read_and_join_tables` takes six named columns off ``png_list`` -- the object id, ``png_path`` and the four field keys -- and drops everything else. Everything else is precisely the annotation columns the Annotate app added and the score columns the classifier wrote, which for an AnnData export are among the most valuable columns in the database: without them ``adata.obs`` has no label to train on, group by, or colour a UMAP with, and a user would have to re-join ``png_list`` by hand outside spaCR. The columns are recognised, not guessed: everything ``png_list`` gets from :func:`spacr.utils.filepaths_to_database` is listed in :data:`spacr.agreement._METADATA_COLUMNS`, so what is left over is what a human or a model put there. The join is on the anchor object's own crop id (``cell_id`` for a cell export, ``nucleus_id`` for a nucleus one), migrated to an integer label with :func:`spacr.utils.object_label_from_png_id` -- the one implementation that survives the ``'omulti'`` / ``'onone'`` / ``'error'`` / NULL values the real writers produce. :returns: ``(frame, attached)`` -- the frame and the column names added. A ``png_list`` that cannot be attached unambiguously adds nothing and is reported, rather than attaching a plausible wrong label. """ from ..agreement import _METADATA_COLUMNS from ..utils import PNG_OBJECT_ID_COLUMNS, object_label_from_png_id id_column = PNG_OBJECT_ID_COLUMNS.get(anchor) if not id_column: return frame, [] connection = sqlite3.connect(db_path, timeout=30) try: png = pd.read_sql('SELECT * FROM "png_list"', connection) except Exception: return frame, [] finally: connection.close() if id_column not in png.columns: return frame, [] extra = [c for c in png.columns if c not in _METADATA_COLUMNS and c not in frame.columns and c != id_column] if not extra: return frame, [] key_columns = list(OBJECT_KEY_COLUMNS) if timelapse: key_columns = list(schema.TIMEPOINT_KEY_COLUMNS) + [ schema.OBJECT_LABEL_KEY] if not all(c in png.columns for c in key_columns[:-1]): return frame, [] labels = object_label_from_png_id(png[id_column]) png = png.loc[labels.notna()].copy() if png.empty: return frame, [] png[schema.OBJECT_LABEL_KEY] = labels.loc[png.index].astype("int64") try: keys = object_keys(png, timelapse=timelapse, object_type=anchor) except FilterError: return frame, [] subset = png[extra].copy() subset.index = pd.Index(keys) duplicated = subset.index.duplicated(keep=False) if duplicated.any(): conflicting = subset.loc[duplicated] agreed = conflicting.groupby(level=0).nunique(dropna=False).le(1).all() if not bool(agreed.all()): return frame, [] subset = subset[~subset.index.duplicated(keep="first")] frame = frame.copy() frame_keys = object_keys(frame, timelapse=timelapse) aligned = subset.reindex(pd.Index(frame_keys)) for column in extra: frame[column] = aligned[column].to_numpy() return frame, list(extra) def _measurement_units(frame: pd.DataFrame) -> Optional[str]: """The single ``measurement_units`` value of ``frame``, or ``None``. ``None`` when the column is absent (a legacy database) or carries more than one value (a database measured twice under different calibration). A mixed frame deliberately gets ``None`` rather than a majority vote: :func:`spacr.feature_dict.parse_column` then states the condition instead of asserting a unit that is wrong for some of the rows. """ if "measurement_units" not in frame.columns: return None values = frame["measurement_units"].dropna().unique() if len(values) != 1: return None return str(values[0]) def _label_columns(frame: pd.DataFrame, db_path: Optional[str]) -> Tuple[Tuple[str, ...], Tuple[str, ...]]: """Split the label-ish columns into ``(annotation, prediction)``. Both are numeric and neither is a measurement, so both would otherwise land in ``X`` and be scaled, PCA'd and clustered as if they were features. :mod:`spacr.agreement` already knows the difference -- it has to, because offering a model's column as a third annotator gave a real database kappa = -0.004 -- so its answer is reused rather than a second list being written here. :param frame: the frame being exported. :param db_path: the database, when the human-annotation guess (which needs the column's distinct-value count) can be made against it. :returns: ``(annotation_columns, prediction_columns)``, both in frame order and disjoint. """ from ..agreement import PNG_TABLE, _is_model_column, annotation_columns columns = list(frame.columns) predictions = tuple(c for c in columns if _is_model_column(c, columns)) annotations: Tuple[str, ...] = () if db_path: try: guessed = annotation_columns(db_path, table=PNG_TABLE) except Exception: guessed = [] annotations = tuple(c for c in columns if c in set(guessed) and c not in predictions) return annotations, predictions
[docs] def feature_columns(frame: pd.DataFrame, *, exclude: Sequence[str] = (), db_path: Optional[str] = None) -> List[str]: """The columns of ``frame`` that belong in ``X``, in frame order. A column is a feature when it is numeric, is not provenance or identity by :func:`spacr.schema.is_provenance_column`, is not a human annotation or a model output, is not the cluster label, and was not excluded by the caller. :param frame: an object frame read from a measurements database. :param exclude: extra column names to keep out of ``X``. :param db_path: the source database, used to recognise human annotation columns; omit it and only the model columns are recognised. :returns: feature column names. """ annotations, predictions = _label_columns(frame, db_path) blocked = (set(exclude) | set(annotations) | set(predictions) | set(_NON_FEATURE_NUMERIC)) keep: List[str] = [] for column in frame.columns: if column in blocked: continue if schema.is_provenance_column(column): continue if not pd.api.types.is_numeric_dtype(frame[column]): continue keep.append(str(column)) return keep
def _source_table(column: str, entry, tables: Sequence[str]) -> str: """Which measurement table ``column`` came out of. Read off the object type :func:`spacr.feature_dict.parse_column` found, falling back to the join suffix (``..._nucleus``) and then to the anchor table. ``""`` when it cannot be attributed, which is honest -- the alternative is a confident wrong answer in ``var``. """ for obj in (entry.object_type_2, entry.object_type): if obj and obj in tables: return str(obj) for table in tables: if column.endswith(f"_{table}"): return table if column.startswith("count_"): return column[len("count_"):] return tables[0] if tables else "" def _build_var(features: Sequence[str], frame: pd.DataFrame, tables: Sequence[str], anchor: str, units: Optional[str]) -> pd.DataFrame: """Build the per-feature ``var`` frame. Every field comes from :func:`spacr.feature_dict.parse_column`, resolved under ``units`` so geometric columns carry a concrete unit rather than a conditional one. Nothing is guessed: an unrecognised column arrives with ``family='unknown'`` and an empty description, which is the dictionary's own answer and is more use than a fabricated one. """ from ..feature_dict import describe_columns entries = describe_columns(features, units) aggregated_tables = {t for t in tables if t in schema.CHILD_OBJECT_TABLES} if anchor == "cell" else set() rows = [] for column, entry in zip(features, entries): source = _source_table(column, entry, tables) values = pd.to_numeric(frame[column], errors="coerce") finite = np.isfinite(values.to_numpy(dtype=float, na_value=np.nan)) rows.append({ "object_type": entry.object_type or "", "object_type_2": entry.object_type_2 or "", "channel": (-1 if entry.channel is None else int(entry.channel)), "channel_2": (-1 if entry.channel_2 is None else int(entry.channel_2)), "channel_scope": entry.channel_scope, "family": entry.family, "feature_key": entry.key or "", "description": entry.description or "", "unit": entry.unit or "", "measurement_units": entry.measurement_units or (units or ""), "computed_by": entry.computed_by, "module": entry.module, "notes": entry.notes or "", "written_when": entry.written_when or "", "concepts": ";".join(entry.concepts), "source_table": source, "is_aggregated": bool(source in aggregated_tables), "n_missing": int((~finite).sum()), "frac_missing": (float((~finite).sum()) / len(frame)) if len(frame) else 0.0, "n_infinite": int(np.isinf( values.to_numpy(dtype=float, na_value=np.nan)).sum()), }) var = pd.DataFrame(rows, index=pd.Index(list(features), name=None)) _hdf5_metadata(var) for categorical in ("object_type", "channel_scope", "family", "source_table", "measurement_units"): var[categorical] = var[categorical].astype("category") return var def _redundant_identity_columns(columns: Sequence[str], tables: Sequence[str]) -> List[str]: """Join-suffixed copies of identity columns the anchor already carries. :func:`spacr.io._read_and_join_tables` suffixes every colliding column, so a four-table join hands back ``plateID`` and then ``plateID_nucleus``, ``plateID_pathogen`` and ``plateID_cytoplasm`` -- four spellings of one plate. Two of them are worse than redundant: ``object_label_nucleus`` and ``cell_id_pathogen`` come through the child aggregation, so their values are the *mean of the child labels*, a number with no referent at all. Only provenance columns are removed, and only when the unsuffixed column is present to stand for them. Feature columns keep their suffixes, because ``nucleus_area_nucleus`` is a real measurement. :param columns: the candidate ``obs`` columns. :param tables: the tables that were joined. :returns: the columns to drop, in input order. """ present = set(columns) drop: List[str] = [] for column in columns: for table in tables: suffix = f"_{table}" if not column.endswith(suffix): continue base = column[:-len(suffix)] if (base in present and base != column and schema.is_provenance_column(base)): drop.append(column) break return drop def _hdf5_metadata(frame: pd.DataFrame) -> None: """Normalize owned metadata in place for AnnData's HDF5 string encoding. Preserve values and missingness while avoiding nullable string storage, which AnnData requires callers to opt into globally. Entirely missing object columns use empty categoricals; numeric metadata is untouched. """ frame.index = frame.index.astype(object) for column in frame.columns: values = frame[column] if isinstance(values.dtype, pd.CategoricalDtype): if isinstance(values.cat.categories.dtype, pd.StringDtype): frame[column] = values.cat.rename_categories( values.cat.categories.astype(object)) elif isinstance(values.dtype, pd.StringDtype): frame[column] = values = values.astype(object) if values.dtype == object and values.isna().all(): frame[column] = pd.Categorical(values) def _build_obs(frame: pd.DataFrame, features: Sequence[str], annotations: Sequence[str], predictions: Sequence[str], *, timelapse: bool, condition_map: Optional[Mapping[str, str]], condition_column: str, drop_columns: Sequence[str] = ()) -> pd.DataFrame: """Build ``obs``: everything that is not a feature, plus the key columns. The index is :func:`spacr.selection.object_keys`, which is spaCR's own object identity -- not a new one invented for AnnData. Entirely missing object-typed metadata uses an empty categorical for HDF5 storage. Missingness is preserved without inventing a numeric value or a string label; populated calibration columns remain numeric. """ obs = frame.drop(columns=[c for c in features if c in frame.columns]) obs = obs.drop(columns=[c for c in drop_columns if c in obs.columns]) obs = obs.copy() keys = object_keys(frame, timelapse=timelapse) duplicated = pd.Index(keys).duplicated() if duplicated.any(): offenders = sorted(set(pd.Index(keys)[duplicated])) raise DuplicateObjectKeys( f"{int(duplicated.sum())} of {len(keys)} rows repeat an object " f"key, and AnnData requires unique obs_names. First few: " f"{offenders[:5]}. This means the database holds more than one " f"row for the same object -- usually a field measured twice. " f"Deduplicate it deliberately (spacr.resume can delete a field's " f"rows before a re-measure); this export will not guess which " f"row is the right one." + ("" if timelapse else " If this is a timelapse database, pass timelapse=True so each " "frame keys separately.")) obs.index = pd.Index(keys, name="object_key") if condition_map is not None and condition_column in frame.columns: mapping = dict(condition_map) obs["condition"] = [ mapping.get(str(value), CONDITION_FALLBACK) for value in frame[condition_column]] _hdf5_metadata(obs) categorical = list(OBJECT_KEY_COLUMNS[:-1]) + [ schema.PRC_KEY, schema.PRCF_KEY, "condition", "cluster", "measurement_units", *annotations, *predictions] if timelapse: categorical.append(schema.TIME_KEY) for column in categorical: if column in obs.columns and obs[column].nunique(dropna=False) <= 2048: obs[column] = obs[column].astype("category") return obs def _apply_nan_policy(matrix: np.ndarray, features: List[str], policy: str) -> Tuple[np.ndarray, List[str], np.ndarray, Optional[np.ndarray], Dict[str, Any]]: """Apply ``policy`` to ``matrix``, returning what it did. :param matrix: the float feature matrix, non-finite values already converted to NaN by the caller. :param features: column names of ``matrix``. :param policy: one of :data:`NAN_POLICIES`. :returns: ``(matrix, features, keep_rows, missing_mask, report)`` where ``keep_rows`` is a boolean mask over the original rows and ``missing_mask`` is the ``layers['missing']`` array or ``None``. :raises ValueError: on an unknown policy. """ if policy not in NAN_POLICIES: raise ValueError( f"nan_policy={policy!r} is not one of {list(NAN_POLICIES)}. See " f"the spacr.anndata_export module docstring for what each does " f"and why 'keep' is the default.") missing = np.isnan(matrix) report: Dict[str, Any] = { "policy": policy, "n_missing": int(missing.sum()), "n_objects_counted": int(matrix.shape[0]), "n_features_counted": int(matrix.shape[1]), "n_features_with_missing": int((missing.any(axis=0)).sum()), "n_objects_with_missing": int((missing.any(axis=1)).sum()), "imputed": False, "dropped_features": [], "dropped_objects": 0, } keep_rows = np.ones(matrix.shape[0], dtype=bool) mask: Optional[np.ndarray] = None if policy == NAN_KEEP or not missing.any(): report["worst_features"] = _worst_features(missing, features) return matrix, list(features), keep_rows, mask, report if policy == NAN_DROP_FEATURES: keep = ~missing.any(axis=0) report["dropped_features"] = [f for f, k in zip(features, keep) if not k] report["worst_features"] = _worst_features(missing, features) return (matrix[:, keep], [f for f, k in zip(features, keep) if k], keep_rows, mask, report) if policy == NAN_DROP_OBJECTS: keep_rows = ~missing.any(axis=1) report["dropped_objects"] = int((~keep_rows).sum()) report["worst_features"] = _worst_features(missing, features) return matrix[keep_rows], list(features), keep_rows, mask, report mask = missing report["imputed"] = True if policy == NAN_ZERO: matrix = np.where(missing, 0.0, matrix) else: with warnings.catch_warnings(): warnings.simplefilter("ignore", category=RuntimeWarning) means = np.nanmean(matrix, axis=0) means = np.where(np.isfinite(means), means, 0.0) matrix = np.where(missing, means[None, :], matrix) report["worst_features"] = _worst_features(missing, features) return matrix, list(features), keep_rows, mask, report def _worst_features(missing: np.ndarray, features: Sequence[str], limit: int = 10) -> List[List[Any]]: """The ``limit`` features with the most missing values, worst first. Returned as ``[[name, count], ...]`` rather than a dict because ``uns`` is written to HDF5 and a list of pairs round-trips through every ``h5ad`` reader, while a dict of arbitrary column names does not. """ if not len(features) or not missing.any(): return [] counts = missing.sum(axis=0) order = np.argsort(-counts)[:limit] return [[str(features[i]), int(counts[i])] for i in order if counts[i] > 0] def _obsm_name(name: str) -> str: """Normalise an embedding name to the scanpy ``X_*`` convention.""" text = str(name) return text if text.startswith("X_") else f"X_{text}" def _align_embedding(values: Any, keys: pd.Index, *, timelapse: bool, name: str) -> np.ndarray: """Coerce one embedding to an ``(n_obs, k)`` array aligned to ``keys``. A :class:`pandas.DataFrame` carrying the object key columns is aligned **by key**; anything else is taken positionally against the *unfiltered* frame and then narrowed. Aligning a keyed frame positionally is the bug this exists to prevent: the whole point of the filtered export is that the exported rows are a subset, and a positional take would then attach the wrong point to every object after the first gap. :raises ValueError: on a length or key mismatch. """ if isinstance(values, pd.DataFrame) and all( c in values.columns for c in OBJECT_KEY_COLUMNS): coordinate_columns = [c for c in values.columns if c not in set(OBJECT_KEY_COLUMNS) and pd.api.types.is_numeric_dtype(values[c]) and c != schema.OBJECT_LABEL_KEY] if not coordinate_columns: raise ValueError( f"embedding {name!r} carries the object key columns but no " f"numeric coordinate columns to go with them.") indexed = values.set_index( object_keys(values, timelapse=timelapse))[coordinate_columns] wanted = [k if k in indexed.index else untyped_object_key(k) for k in keys] missing = [k for k in wanted if k not in indexed.index] if missing: raise ValueError( f"embedding {name!r} has no coordinates for " f"{len(missing)} of the {len(keys)} exported objects " f"(first: {missing[:3]}). It was computed on a different " f"population; recompute it on the filtered frame or pass " f"the same filter to both.") return np.asarray(indexed.loc[wanted].to_numpy(), dtype=np.float32) array = np.asarray(values, dtype=np.float32) if array.ndim == 1: array = array.reshape(-1, 1) if array.shape[0] == len(keys): return array raise ValueError( f"embedding {name!r} has {array.shape[0]} rows but the export has " f"{len(keys)} objects. Pass a DataFrame carrying " f"{list(OBJECT_KEY_COLUMNS)} to have it aligned by object key " f"instead of by position.") def _compute_umap(matrix: np.ndarray, features: Sequence[str], settings: Mapping[str, Any]) -> np.ndarray: """Compute a 2-D UMAP through the same path the UMAP app uses. :func:`spacr.utils.reduction_and_clustering` is the function :func:`spacr.core.generate_image_umap` calls, with the seed it uses, so an exported ``X_umap`` and the app's plot are the same embedding rather than two that merely look alike. NaN is imputed to the feature mean *for the reducer only* -- UMAP cannot consume NaN and would otherwise return an all-NaN embedding -- and ``X`` itself is untouched. The substitution is recorded in ``uns['spacr']['umap']``. """ from ..utils import reduction_and_clustering values = np.array(matrix, dtype=float, copy=True) finite = np.isfinite(values) if not finite.all(): with warnings.catch_warnings(): warnings.simplefilter("ignore", category=RuntimeWarning) means = np.nanmean(np.where(finite, values, np.nan), axis=0) means = np.where(np.isfinite(means), means, 0.0) values = np.where(finite, values, means[None, :]) embedding, _labels, _reducer = reduction_and_clustering( values, n_neighbors=settings.get("n_neighbors", 15), min_dist=settings.get("min_dist", 0.1), metric=settings.get("metric", "euclidean"), eps=settings.get("eps", 0.5), min_samples=settings.get("min_samples", 5), clustering=settings.get("clustering", "dbscan"), reduction_method=settings.get("reduction_method", "umap"), verbose=False, n_jobs=settings.get("n_jobs", 1)) return np.asarray(embedding, dtype=np.float32) def _h5ad_safe(value: Any) -> Any: """Coerce ``value`` into something ``h5ad`` can store. ``None`` becomes ``""``, tuples become lists, everything unrecognised becomes its ``repr``. AnnData's HDF5 writer raises on a ``None`` or an arbitrary object buried three dicts deep, and losing a finished export to a provenance field would be absurd. """ if value is None: return "" if isinstance(value, (str, bool, int, float, np.integer, np.floating)): return value if isinstance(value, Mapping): return {str(k): _h5ad_safe(v) for k, v in value.items()} if isinstance(value, np.ndarray): return value if isinstance(value, (list, tuple, set)): return [_h5ad_safe(v) for v in value] return repr(value) def _filter_record(data_filter: Optional[DataFilter], selection: Optional[Selection]) -> Dict[str, Any]: """Describe the filtering applied, in prose and in structure.""" record: Dict[str, Any] = { "description": data_filter.describe() if data_filter else "no filter", "clauses": [], "selection_source": "", "selection_size": 0, } for clause in (data_filter.clauses if data_filter else []): entry = {"column": clause.column, "type": type(clause).__name__} if hasattr(clause, "low"): entry["low"] = "" if clause.low is None else float(clause.low) entry["high"] = "" if clause.high is None else float(clause.high) else: entry["values"] = [str(v) for v in clause.values] record["clauses"].append(entry) if selection is not None and selection.is_active: record["selection_source"] = selection.source record["selection_size"] = len(selection) return record def _relationships(tables: Sequence[str], anchor: str, joined: bool) -> Dict[str, Any]: """Describe the cell -> nucleus / pathogen links this export carries. See the module docstring for why this is ``uns`` and ``obs`` rather than ``obsp``. """ record: Dict[str, Any] = { "storage": "obs and uns, never obsp", "why_not_obsp": ( "obsp is an n_obs x n_obs relation among THIS AnnData's " "observations. In a cell-anchored export the nuclei and " "pathogens have been aggregated onto their parent and are not " "observations at all, so there is no row to point at; in a " "child-anchored export the cells are not observations, for the " "mirror-image reason. The link is a foreign key, and it is " "stored as one."), "anchor": anchor, "children": {}, } if joined: for table in tables: if table in schema.CHILD_OBJECT_TABLES and table != anchor: record["children"][table] = { "aggregated": True, "aggregation": "mean over the children of each parent", "count_column": f"count_{table}", "column_suffix": f"_{table}", "see": "var['is_aggregated']", } else: contract = schema.OBJECT_TABLE_SCHEMAS.get(anchor) parent = getattr(contract, "parent_column", None) if contract else None if parent: record["parent"] = { "table": "cell", "obs_column": parent, "aggregated": False, "note": ("a plain foreign key: adata.obs.groupby('cell_id') " "gives this export's objects per parent cell."), } return record def _assemble_tables(db_path: Union[str, os.PathLike], *, tables: Sequence[str] = DEFAULT_TABLES, single_table: Optional[str] = None, data_filter: Optional[DataFilter] = None, selection: Optional[Selection] = None, row_limit: Optional[int] = None, timelapse: bool = False, nan_policy: str = NAN_KEEP, exclude: Sequence[str] = (), condition_map: Optional[Mapping[str, str]] = None, condition_column: str = schema.COLUMN_KEY, attach_labels: bool = True, drop_redundant_identity: bool = True, ) -> Dict[str, Any]: """Read, filter and split a measurements database into export tables. The part of an export that does not depend on the file format: the per-object metadata frame, the per-feature description, the float64 feature matrix after the NaN policy, and the counts and notes every format records. :func:`build_anndata` wraps the result in an AnnData; the Parquet and R exports write it as tables. Arguments are those of :func:`build_anndata`. :returns: a dict with ``db_path``, ``frame``, ``obs``, ``var``, ``matrix``, ``features``, ``missing_mask``, ``nan_report``, ``notes``, ``n_before``, ``n_missing_before``, ``n_infinite``, ``read_tables``, ``anchor``, ``joined``, ``units``, ``annotations`` and ``predictions``. :raises DuplicateObjectKeys: when two rows claim one object key. :raises ValueError: on an unknown ``nan_policy`` or an unusable database. """ db_path = os.path.abspath(os.path.expanduser(os.fspath(db_path))) frame, read_tables = _read_frame( db_path, tables, single_table, collapse_duplicate_identity=drop_redundant_identity) n_before = len(frame) joined = single_table is None anchor = "cell" if joined else str(single_table) notes: List[str] = [] if attach_labels and "png_list" in _available_tables(db_path): frame, attached = _attach_png_labels( frame, db_path, anchor, timelapse=timelapse) if attached: notes.append( f"attached {len(attached)} label column(s) from png_list " f"that the table join drops: {', '.join(attached[:6])}" + (" ..." if len(attached) > 6 else "")) if data_filter is not None and not data_filter.is_empty: frame = frame.loc[data_filter.mask(frame)] if selection is not None and selection.is_active: frame = frame.loc[selection.mask_for(frame, timelapse=timelapse)] if row_limit is not None and len(frame) > int(row_limit): notes.append( f"row_limit={int(row_limit)} truncated the export from " f"{len(frame)} objects; it is a cap, not a filter, so the " f"exported objects are simply the first {int(row_limit)} in " f"table order.") frame = frame.iloc[:int(row_limit)] frame = frame.reset_index(drop=True) units = _measurement_units(frame) annotations, predictions = _label_columns(frame, db_path) features = feature_columns(frame, exclude=exclude, db_path=db_path) if not features: raise ValueError( f"no feature columns found in {db_path}: every numeric column " f"was identity, provenance, an annotation or a model output. " f"This is usually a database whose object tables were never " f"written -- check `spacr doctor --db {db_path}`.") redundant: List[str] = [] if joined and drop_redundant_identity: non_features = [c for c in frame.columns if c not in set(features)] redundant = _redundant_identity_columns(non_features, read_tables) if redundant: notes.append( f"{len(redundant)} join-suffixed copies of identity columns " f"were dropped from obs (e.g. " f"{', '.join(redundant[:3])}); the anchor table's own " f"columns carry the same values, and the ones that came " f"through the child aggregation carried the mean of a label " f"rather than a label. Pass " f"drop_redundant_identity=False to keep them.") obs = _build_obs(frame, features, annotations, predictions, timelapse=timelapse, condition_map=condition_map, condition_column=condition_column, drop_columns=redundant) var = _build_var(features, frame, read_tables, anchor, units) matrix = frame[features].apply( pd.to_numeric, errors="coerce").to_numpy(dtype=np.float64) infinite = np.isinf(matrix) n_infinite = int(infinite.sum()) if n_infinite: matrix = np.where(infinite, np.nan, matrix) notes.append( f"{n_infinite} non-finite (+/-inf) values were converted to NaN " f"before the nan_policy ran: an inf survives dropna() and then " f"destroys any scaling, PCA or distance computed from it.") matrix, features, keep_rows, missing_mask, nan_report = _apply_nan_policy( matrix, features, nan_policy) if not keep_rows.all(): obs = obs.loc[keep_rows] frame = frame.loc[keep_rows].reset_index(drop=True) if nan_report["dropped_features"]: var = var.loc[features] var = var.copy() var["n_missing_raw"] = var["n_missing"].to_numpy() var["frac_missing_raw"] = var["frac_missing"].to_numpy() final_missing = np.isnan(matrix).sum(axis=0) var["n_missing"] = final_missing.astype(np.int64) var["frac_missing"] = (final_missing / matrix.shape[0] if matrix.shape[0] else np.zeros(matrix.shape[1])) n_missing_before = int(nan_report["n_missing"]) per_object_missing = ( np.isnan(matrix).sum(axis=1) if nan_policy in (NAN_KEEP, NAN_DROP_FEATURES, NAN_DROP_OBJECTS) else (missing_mask.sum(axis=1) if missing_mask is not None else np.zeros(matrix.shape[0], dtype=int))) obs = obs.copy() obs["n_missing_features"] = np.asarray(per_object_missing, dtype=np.int32) return { "db_path": db_path, "frame": frame, "obs": obs, "var": var, "matrix": matrix, "features": list(features), "missing_mask": missing_mask, "nan_report": nan_report, "notes": notes, "n_before": int(n_before), "n_missing_before": n_missing_before, "n_infinite": n_infinite, "read_tables": tuple(read_tables), "anchor": anchor, "joined": bool(joined), "units": units, "annotations": list(annotations), "predictions": list(predictions), } def _provenance_record(parts: Mapping[str, Any], *, n_objects: int, n_features: int, data_filter: Optional[DataFilter], selection: Optional[Selection], settings: Optional[Mapping[str, Any]], run_id: str, timelapse: bool) -> Dict[str, Any]: """The provenance every export format records. :param parts: the result of :func:`_assemble_tables`. :param n_objects: objects written. :param n_features: feature columns written. :param data_filter: the filter applied, or ``None``. :param selection: the selection applied, or ``None``. :param settings: the run settings, hashed into the record. :param run_id: the run this export belongs to; read from the database when empty. :param timelapse: whether each frame was keyed separately. :returns: a dict of plain values: version, settings hash, run id, source, filter, missing-value report, label columns, relationships, notes and feature-dictionary coverage. """ from ..artifacts import settings_hash from ..feature_dict import coverage as feature_coverage from ..version import get_version db_path = parts["db_path"] record: Dict[str, Any] = { "spacr_version": get_version(), "settings_hash": settings_hash(settings), "run_id": str(run_id or _run_id_from_db(db_path)), "source_database": db_path, "source_tables": list(parts["read_tables"]), "anchor_object": parts["anchor"], "joined": bool(parts["joined"]), "exported_utc": datetime.now(timezone.utc).isoformat(), "object_key_columns": list(OBJECT_KEY_COLUMNS), "timelapse": bool(timelapse), "measurement_units": parts["units"] or "", "n_objects": int(n_objects), "n_objects_before_filter": int(parts["n_before"]), "n_features": int(n_features), "filter": _filter_record(data_filter, selection), "nan": parts["nan_report"], "annotation_columns": list(parts["annotations"]), "prediction_columns": list(parts["predictions"]), "relationships": _relationships( parts["read_tables"], parts["anchor"], parts["joined"]), "notes": list(parts["notes"]), } explained = feature_coverage(parts["features"], parts["units"]) record["feature_dictionary"] = { "total": int(explained.total), "explained": int(explained.explained), "unknown": list(explained.unknown[:50]), } return record
[docs] def build_anndata(db_path: Union[str, os.PathLike], *, tables: Sequence[str] = DEFAULT_TABLES, single_table: Optional[str] = None, data_filter: Optional[DataFilter] = None, selection: Optional[Selection] = None, row_limit: Optional[int] = None, timelapse: bool = False, nan_policy: str = NAN_KEEP, missing_layer: Optional[bool] = None, dtype: str = "float32", exclude: Sequence[str] = (), embeddings: Optional[Mapping[str, Any]] = None, compute_umap: bool = False, umap_settings: Optional[Mapping[str, Any]] = None, condition_map: Optional[Mapping[str, str]] = None, condition_column: str = schema.COLUMN_KEY, attach_labels: bool = True, drop_redundant_identity: bool = True, settings: Optional[Mapping[str, Any]] = None, run_id: str = "", verbose: bool = True): """Build an :class:`anndata.AnnData` from a spaCR measurements database. The mapping is described in full in the module docstring; in short, ``X`` is the numeric measurements, ``obs`` is everything else about the object, ``var`` is :mod:`spacr.feature_dict`'s description of each feature, ``obsm`` holds embeddings and ``uns`` holds provenance. :param db_path: a ``measurements.db``. :param tables: tables to join for the default cell-anchored export. :param single_table: export exactly this object table instead, one row per object of that type -- which is the only way to get a nucleus-level or pathogen-level matrix, since the join averages children onto their parent. :param data_filter: a :class:`spacr.selection.DataFilter`. Declarative and re-appliable; recorded in ``uns``. :param selection: a :class:`spacr.selection.Selection` -- the keys a view pointed at. Applied after ``data_filter``. :param row_limit: hard cap on exported objects, applied last. A blunt instrument on purpose, for "give me something I can open" without inventing a filter that means something it does not. :param timelapse: key each frame of an object separately. :param nan_policy: one of :data:`NAN_POLICIES`; see the module docstring. The default keeps NaN and reports it. :param missing_layer: write ``layers['missing']``. Defaults to True for the imputing policies and False otherwise. :param dtype: ``X`` dtype. ``float32`` by default -- the scanpy convention, half the memory, and far more precision than any microscope measurement carries. :param exclude: feature columns to keep out of ``X``. :param embeddings: ``{name: array or DataFrame}``. Names are normalised to the scanpy ``X_*`` convention, so ``'umap'`` becomes ``X_umap``. A DataFrame carrying :data:`spacr.selection.OBJECT_KEY_COLUMNS` is aligned **by key**, which is what makes an embedding computed on the whole plate usable with a filtered export. :param compute_umap: compute ``X_umap`` here, through the same :func:`spacr.utils.reduction_and_clustering` call :func:`spacr.core.generate_image_umap` makes. Off by default: it imports the segmentation stack and costs minutes on a large table. :param umap_settings: overrides for that computation. :param condition_map: ``{column value: label}`` written to ``obs['condition']``; :data:`DEFAULT_CONDITION_MAP` is the mapping :func:`spacr.utils.map_condition` applies. :param condition_column: which column ``condition_map`` reads. :param attach_labels: bring the annotation and prediction columns back out of ``png_list``, which the table join drops. On by default -- see :func:`_attach_png_labels`. :param drop_redundant_identity: drop the join's suffixed copies of identity columns (``plateID_nucleus``, ``object_label_pathogen``) from ``obs``. On by default; see :func:`_redundant_identity_columns` for why two of them are worse than merely duplicated. :param settings: the run settings, hashed into ``uns`` provenance. :param run_id: the run this export belongs to. :param verbose: print the summary and the missing-data warning. :returns: ``(adata, result)`` -- the AnnData and an :class:`ExportResult`. :raises AnnDataExtraMissing: when ``anndata`` is not installed. :raises DuplicateObjectKeys: when two rows claim one object key. :raises ValueError: on an unknown ``nan_policy``, an unusable database, or an embedding that cannot be aligned. """ anndata = require_anndata() parts = _assemble_tables( db_path, tables=tables, single_table=single_table, data_filter=data_filter, selection=selection, row_limit=row_limit, timelapse=timelapse, nan_policy=nan_policy, exclude=exclude, condition_map=condition_map, condition_column=condition_column, attach_labels=attach_labels, drop_redundant_identity=drop_redundant_identity) obs, var, matrix = parts["obs"], parts["var"], parts["matrix"] features = parts["features"] missing_mask, nan_report = parts["missing_mask"], parts["nan_report"] notes = parts["notes"] n_before, n_infinite = parts["n_before"], parts["n_infinite"] n_missing_before = parts["n_missing_before"] matrix = np.ascontiguousarray(matrix, dtype=np.dtype(dtype)) adata = anndata.AnnData(X=matrix, obs=obs, var=var) if missing_layer is None: missing_layer = bool(nan_report["imputed"]) if missing_layer and missing_mask is not None: adata.layers["missing"] = missing_mask elif missing_layer: adata.layers["missing"] = np.isnan(np.asarray(matrix, dtype=float)) keys = pd.Index(adata.obs_names) obsm_notes: Dict[str, Any] = {} for name, values in dict(embeddings or {}).items(): adata.obsm[_obsm_name(name)] = _align_embedding( values, keys, timelapse=timelapse, name=str(name)) if compute_umap and "X_umap" not in adata.obsm: if len(adata) < 3: notes.append( f"compute_umap was asked for but the export has " f"{len(adata)} objects; UMAP needs at least 3. No X_umap " f"was written.") else: adata.obsm["X_umap"] = _compute_umap( np.asarray(adata.X, dtype=float), features, dict(umap_settings or {})) obsm_notes["X_umap"] = { "computed_by": "spacr.utils.reduction_and_clustering", "nan_handling": ("feature-mean imputed for the reducer only; " "X itself is untouched"), } from ..artifacts import material_settings provenance = _provenance_record( parts, n_objects=int(adata.n_obs), n_features=int(adata.n_vars), data_filter=data_filter, selection=selection, settings=settings, run_id=run_id, timelapse=timelapse) provenance["artifact"] = { "module": APP_KEY, "kind": ANNDATA_KIND, "role": "h5ad", "note": ("the artifact id is derived from this file's content " "fingerprint and so cannot live inside it; find the " "record with spacr.artifacts.by_kind('anndata', " "project=<project root>)"), } if obsm_notes: provenance["umap"] = obsm_notes.get("X_umap", {}) adata.uns["spacr"] = _h5ad_safe(provenance) adata.uns["spacr_settings"] = _h5ad_safe(material_settings(settings)) result = ExportResult( path="", n_obs=int(adata.n_obs), n_vars=int(adata.n_vars), n_obs_before_filter=int(n_before), obs_columns=tuple(str(c) for c in adata.obs.columns), obsm_keys=tuple(str(k) for k in adata.obsm.keys()), nan_policy=nan_policy, n_missing=n_missing_before, n_infinite=n_infinite, dropped_features=tuple(nan_report["dropped_features"]), dropped_objects=int(nan_report["dropped_objects"]), n_obs_counted=int(nan_report["n_objects_counted"]), n_vars_counted=int(nan_report["n_features_counted"]), warnings=tuple(notes), ) if verbose: print(result.describe()) _warn_about_missing(nan_report, nan_policy) for note in notes: print(f" note: {note}") return adata, result
def _warn_about_missing(report: Mapping[str, Any], policy: str) -> None: """Print the missing-data warning, naming the columns and the risk.""" if policy != NAN_KEEP or not report.get("n_missing"): return worst = report.get("worst_features") or [] names = ", ".join(f"{name} ({count})" for name, count in worst[:5]) print( f" {report['n_missing']} missing values kept in X across " f"{report['n_features_with_missing']} features. AnnData stores " f"them; sc.pp.scale, sc.pp.pca and sc.pp.neighbors do not -- they " f"will propagate NaN across the whole matrix. Worst: {names}. " f"Pass nan_policy='drop_features' or 'mean' if you need a complete " f"matrix; see var['n_missing'] for the full picture.") def _run_id_from_db(db_path: str) -> str: """The most recent run id recorded in ``db_path``, or ``""``. Read from ``settings_history`` through :func:`spacr.io.read_settings_history` rather than from :mod:`spacr.runctx`, so an export run days after the measurement still attributes the data to the run that produced it. """ try: from ..io import read_settings_history history = read_settings_history(db_path) except Exception: return "" for entry in reversed(history or []): run = str(entry.get("run_id") or "") if run: return run return "" def _write_h5ad_atomic(adata: Any, path: Union[str, os.PathLike], **kwargs: Any) -> None: """Publish a complete HDF5 file while preserving any previous export. Write in a temporary directory beside the destination, flush the completed file, and replace the destination only on success. Failed writes leave no partial export or scratch files; the original exception reaches the caller. """ path = os.path.abspath(os.fspath(path)) with tempfile.TemporaryDirectory(prefix=".anndata-", dir=os.path.dirname(path)) as folder: pending = os.path.join(folder, "pending.h5ad") adata.write_h5ad(pending, **kwargs) with open(pending, "rb") as handle: os.fsync(handle.fileno()) os.replace(pending, path)
[docs] def export_anndata(db_path: Union[str, os.PathLike], out_path: Union[str, os.PathLike], *, compression: Optional[str] = "gzip", register: bool = True, project: Union[str, os.PathLike, None] = None, settings: Optional[Mapping[str, Any]] = None, **kwargs: Any) -> ExportResult: """Write a spaCR measurements database to ``out_path`` as ``.h5ad``. Everything :func:`build_anndata` accepts is accepted here and passed through; this adds the write and the artifact registration. A failed write preserves any existing destination; only a completed file replaces it and is registered as an artifact. :param db_path: a ``measurements.db``. :param out_path: the ``.h5ad`` to write. Parent directories are created. :param compression: ``h5py`` compression for ``X`` and the layers; ``None`` for an uncompressed file. ``gzip`` typically halves a feature matrix and costs a few seconds. :param register: record the file with :mod:`spacr.artifacts`, with the project's ``measurements-db`` as its input so a re-run of Measure marks this export stale. :param project: the project root the artifact belongs to. Inferred from ``db_path`` (``<project>/measurements/measurements.db``) when omitted. :param settings: the run settings, hashed into the artifact and ``uns``. :param kwargs: passed to :func:`build_anndata`. :returns: the :class:`ExportResult`, with ``path`` and ``artifact_id`` filled in. :raises AnnDataExtraMissing: when ``anndata`` is not installed. """ out_path = os.path.abspath(os.path.expanduser(os.fspath(out_path))) adata, result = build_anndata(db_path, settings=settings, **kwargs) os.makedirs(os.path.dirname(out_path), exist_ok=True) _write_h5ad_atomic(adata, out_path, compression=compression) artifact_id = "" if register: artifact_id = _register(out_path, db_path, project, settings, result, adata) return _dataclass_replace(result, path=out_path, artifact_id=artifact_id)
def _project_root(db_path: str, project: Union[str, os.PathLike, None]) -> str: """The project root for the artifact registry. ``<project>/measurements/measurements.db`` is the layout every spaCR writer uses, so the root is two directories up. """ if project is not None: return os.path.abspath(os.path.expanduser(os.fspath(project))) measurements = os.path.dirname(os.path.abspath(db_path)) if os.path.basename(measurements) == "measurements": return os.path.dirname(measurements) return measurements def _register(out_path: str, db_path: str, project: Union[str, os.PathLike, None], settings: Optional[Mapping[str, Any]], result: ExportResult, adata) -> str: """Register the written file, returning its artifact id or ``""``. Never raises: a registry that cannot be opened (a read-only project, a network filesystem that refuses the lock) must not lose a finished export. The failure is warned about and the file stands on its own, because ``uns['spacr']`` already carries the same provenance. """ try: from .. import artifacts, ports root = _project_root(db_path, project) inputs: List[str] = [] try: upstream = artifacts.latest(ports.MEASUREMENTS_DB, project=root) if upstream is not None: inputs.append(upstream.artifact_id) except Exception: pass record = artifacts.register( project=root, module=APP_KEY, kind=ANNDATA_KIND, role="h5ad", path=out_path, settings=settings, inputs=inputs, run_id=str(adata.uns["spacr"].get("run_id", "")), extra={ "n_obs": int(result.n_obs), "n_vars": int(result.n_vars), "n_obs_before_filter": int(result.n_obs_before_filter), "nan_policy": result.nan_policy, "n_missing": int(result.n_missing), "source_database": db_path, "obsm": list(result.obsm_keys), }) return record.artifact_id except Exception as exc: warnings.warn( f"the AnnData export at {out_path} was written but could not be " f"registered with spacr.artifacts ({exc}). Its provenance is " f"still in uns['spacr'].", RuntimeWarning, stacklevel=2) return ""
[docs] def export_anndata_set(db_path: Union[str, os.PathLike], out_dir: Union[str, os.PathLike], *, object_tables: Sequence[str] = ("cell", "nucleus", "pathogen", "cytoplasm"), prefix: str = "", **kwargs: Any) -> Dict[str, ExportResult]: """Write one ``.h5ad`` per object table, cross-referenced by ``uns``. The honest shape for a multi-compartment experiment: a nucleus is not a cell and averaging it onto one loses the distribution. Each file holds one object type at its own granularity, and each child file records its parent in ``uns['spacr']['relationships']`` -- table, key column and the sibling file -- so the set can be reassembled (or handed to ``MuData``) without guessing. :param db_path: a ``measurements.db``. :param out_dir: folder to write into; created if missing. :param object_tables: which object tables to export. Tables absent from the database are skipped, not an error. :param prefix: prepended to each file name. :param kwargs: passed to :func:`export_anndata`. :returns: ``{table: ExportResult}`` for the tables actually written. """ db_path = os.path.abspath(os.path.expanduser(os.fspath(db_path))) out_dir = os.path.abspath(os.path.expanduser(os.fspath(out_dir))) os.makedirs(out_dir, exist_ok=True) present = set(_available_tables(db_path)) results: Dict[str, ExportResult] = {} files = {table: os.path.join(out_dir, f"{prefix}{table}.h5ad") for table in object_tables if table in present} for table, path in files.items(): contract = schema.OBJECT_TABLE_SCHEMAS.get(table) parent_column = getattr(contract, "parent_column", None) result = export_anndata(db_path, path, single_table=table, **kwargs) results[table] = result if parent_column and "cell" in files: _stamp_parent_file(path, files["cell"], parent_column) return results
def _stamp_parent_file(child_path: str, parent_path: str, parent_column: str) -> None: """Record the sibling file holding this child's parents. Done as a small re-open rather than in :func:`build_anndata` because only the caller writing the whole set knows where the parent landed. """ try: anndata = require_anndata() adata = anndata.read_h5ad(child_path) relationships = dict(adata.uns["spacr"].get("relationships", {})) parent = dict(relationships.get("parent", {})) parent["file"] = os.path.basename(parent_path) parent["obs_column"] = parent_column relationships["parent"] = parent provenance = dict(adata.uns["spacr"]) provenance["relationships"] = relationships adata.uns["spacr"] = provenance _hdf5_metadata(adata.obs) _hdf5_metadata(adata.var) _write_h5ad_atomic(adata, child_path) except Exception as exc: warnings.warn( f"could not record the parent file in {child_path}: {exc}", RuntimeWarning, stacklevel=2) _FORMAT_H5AD = "h5ad" _FORMAT_PARQUET = "parquet" _FORMAT_R = "r" _FORMAT_ALL = "all" #: Every accepted ``anndata_format``: the AnnData file alone, the Parquet #: tables alone, the Parquet tables with the R loader (and ``.rds`` data #: frames when pyreadr is installed), or all of them. _EXPORT_FORMATS: Tuple[str, ...] = ( _FORMAT_H5AD, _FORMAT_PARQUET, _FORMAT_R, _FORMAT_ALL) #: Name and version written into every table's Parquet schema metadata. _TABLES_FORMAT = "spacr-tables" _TABLES_FORMAT_VERSION = 1 #: The artifact kind the objects table registers under. _TABLES_KIND = "parquet_tables" #: File name of the generated R loader. _R_LOADER_NAME = "load_spacr_export.R" #: How the per-well table summarises each feature. _WELL_STATISTICS: Tuple[str, ...] = ("mean", "median") #: The column every table joins on: :func:`spacr.selection.object_keys`. _OBJECT_KEY_COLUMN = "object_key" #: Column holding each well's object count in the per-well table. _WELL_COUNT_COLUMN = "n_objects" @dataclass(frozen=True) class _TablesResult: """What one Parquet/R table export wrote. :param directory: the folder written. :param files: every file written, absolute paths, tables first. :param n_objects: rows of the per-object table. :param n_features: feature columns. :param n_wells: rows of the per-well table; ``0`` when not written. :param r_loader: the R loader script, or ``""``. :param rds_files: ``.rds`` data frames written, if any. :param artifact_id: the :mod:`spacr.artifacts` id of the per-object table, or ``""``. :param warnings: everything the export decided the user must know. """ directory: str files: Tuple[str, ...] n_objects: int n_features: int n_wells: int = 0 r_loader: str = "" rds_files: Tuple[str, ...] = () artifact_id: str = "" warnings: Tuple[str, ...] = () def describe(self) -> str: """One paragraph: shape, folder, files and what R needs.""" lines = [f"{self.n_objects} objects x {self.n_features} features " f"-> {self.directory}"] if self.n_wells: lines.append(f" {self.n_wells} wells in wells.parquet") lines.append(" files: " + ", ".join( os.path.basename(path) for path in self.files)) if self.r_loader: lines.append( f" in R: source('{os.path.basename(self.r_loader)}'); " f"sce <- load_spacr_export()") if self.artifact_id: lines.append(f" artifact {self.artifact_id}") for note in self.warnings: lines.append(f" note: {note}") return "\n".join(lines) def _well_names(rows: pd.Series, columns: pd.Series) -> pd.Series: """Plate-map well names (``'C07'``) for row and column ids. A pair with no well name, such as a positional passthrough, gets a missing value rather than an invented name. """ names: Dict[Tuple[str, str], Any] = {} for row, column in set(zip(rows.astype(str), columns.astype(str))): try: names[(row, column)] = schema.well_id(row, column) except (schema.KeyParseError, ValueError, TypeError): names[(row, column)] = None values = [names[(row, column)] for row, column in zip(rows.astype(str), columns.astype(str))] return pd.Series(pd.Categorical(values), index=rows.index) def _add_well_keys(frame: pd.DataFrame) -> pd.DataFrame: """Add the ``prc`` well key and the plate-map well name when derivable. Both are categoricals, as are the plate, row, column and field ids. """ keys = (schema.PLATE_KEY, schema.ROW_KEY, schema.COLUMN_KEY) if not all(key in frame.columns for key in keys): return frame frame = frame.copy() if schema.PRC_KEY not in frame.columns: frame[schema.PRC_KEY] = schema.compose_prc_column(frame) if schema.WELL_KEY not in frame.columns: frame[schema.WELL_KEY] = _well_names( frame[schema.ROW_KEY], frame[schema.COLUMN_KEY]) for column in (*keys, schema.FIELD_KEY, schema.PRC_KEY, schema.WELL_KEY): if column in frame.columns and not isinstance( frame[column].dtype, pd.CategoricalDtype): frame[column] = frame[column].astype("category") return frame def _numeric_categoricals_as_values(frame: pd.DataFrame) -> pd.DataFrame: """Store categoricals of numbers or booleans as their plain values. Parquet reads a dictionary of numbers back as plain numbers, so a numeric label such as a 0/1 annotation is written as a number in the first place, and reads back with the type it was written with: integers as ``int64`` (``Int64`` when a value is missing), floats as ``float64``, booleans as ``bool`` (``boolean`` when a value is missing). Text categoricals are left as they are. """ frame = frame.copy() for column in frame.columns: values = frame[column] if not isinstance(values.dtype, pd.CategoricalDtype): continue kind = values.cat.categories.dtype.kind if kind not in "iufb": continue target: Any = values.cat.categories.dtype if values.isna().any() and kind in "iu": target = pd.Int64Dtype() elif values.isna().any() and kind == "b": target = pd.BooleanDtype() frame[column] = values.astype(target) return frame def _objects_table(parts: Mapping[str, Any], dtype: str) -> pd.DataFrame: """One row per object: key, metadata, labels and every feature.""" obs = parts["obs"] objects = obs.reset_index() objects = objects.rename( columns={objects.columns[0]: _OBJECT_KEY_COLUMN}) objects[_OBJECT_KEY_COLUMN] = objects[_OBJECT_KEY_COLUMN].astype(str) for column in objects.columns: values = objects[column] if values.dtype == object and pd.api.types.infer_dtype( values, skipna=True) not in ("string", "empty", "boolean"): objects[column] = values.map( lambda value: value if value is None or ( isinstance(value, float) and np.isnan(value)) else str(value)) objects = _add_well_keys(_numeric_categoricals_as_values(objects)) if OBJECT_TYPE_COLUMN in objects.columns: objects[OBJECT_TYPE_COLUMN] = objects[OBJECT_TYPE_COLUMN].astype( "category") features = pd.DataFrame( np.asarray(parts["matrix"], dtype=np.dtype(dtype)), columns=list(parts["features"])) return pd.concat([objects.reset_index(drop=True), features], axis=1) def _well_key_columns(objects: pd.DataFrame, timelapse: bool) -> List[str]: """The columns that identify a well, as present in ``objects``.""" keys = [schema.PLATE_KEY, schema.ROW_KEY, schema.COLUMN_KEY] if not all(key in objects.columns for key in keys): return [] if timelapse and schema.TIME_KEY in objects.columns: keys.append(schema.TIME_KEY) return keys def _wells_table(objects: pd.DataFrame, features: Sequence[str], keys: Sequence[str], statistic: str) -> pd.DataFrame: """One row per well: its keys, object count and each feature's summary. Missing values are skipped, so a well's summary is over the objects that have the feature. ``condition`` is kept when it is the same for every object in the well. """ grouped = objects.groupby(list(keys), observed=True, sort=True) wells = grouped[list(features)].agg(statistic) wells.insert(0, _WELL_COUNT_COLUMN, grouped.size().astype(np.int64)) if "condition" in objects.columns: conditions = grouped["condition"].agg( lambda values: values.iloc[0] if values.nunique(dropna=False) == 1 else None) wells.insert(1, "condition", pd.Categorical(conditions)) wells = wells.reset_index() return _add_well_keys(wells) def _embeddings_table(keys: pd.Index, parts: Mapping[str, Any], embeddings: Optional[Mapping[str, Any]], compute_umap: bool, umap_settings: Optional[Mapping[str, Any]], timelapse: bool, notes: List[str] ) -> Tuple[Optional[pd.DataFrame], Dict[str, List[str]]]: """Object key plus one ``<name>_<k>`` column per embedding dimension.""" arrays: Dict[str, np.ndarray] = {} for name, values in dict(embeddings or {}).items(): arrays[_obsm_name(name)] = _align_embedding( values, keys, timelapse=timelapse, name=str(name)) if compute_umap and "X_umap" not in arrays: if len(keys) < 3: notes.append( f"compute_umap was asked for but the export has {len(keys)} " f"objects; UMAP needs at least 3. No X_umap was written.") else: arrays["X_umap"] = _compute_umap( np.asarray(parts["matrix"], dtype=float), parts["features"], dict(umap_settings or {})) if not arrays: return None, {} table = pd.DataFrame({_OBJECT_KEY_COLUMN: [str(key) for key in keys]}) columns: Dict[str, List[str]] = {} for name, array in arrays.items(): names = [f"{name}_{i + 1}" for i in range(array.shape[1])] columns[name] = names for i, column in enumerate(names): table[column] = np.asarray(array[:, i], dtype=np.float64) return table, columns def _provenance_table(provenance: Mapping[str, Any], settings: Optional[Mapping[str, Any]]) -> pd.DataFrame: """Provenance and material settings as ``section``, ``key``, ``value``. Every value is JSON text, so nested records keep their structure. """ import json from ..artifacts import material_settings rows = [("export", str(key), json.dumps(value, default=str, sort_keys=True)) for key, value in provenance.items()] rows += [("settings", str(key), json.dumps(value, default=str, sort_keys=True)) for key, value in sorted(material_settings(settings).items(), key=lambda item: str(item[0]))] frame = pd.DataFrame(rows, columns=["section", "key", "value"]) frame["section"] = frame["section"].astype("category") return frame _R_LOADER_TEMPLATE = r'''## Load a spaCR table export into R. ## ## Written by spaCR {version} on {date} beside the Parquet tables it reads. ## ## In R: ## source("{loader}") ## sce <- load_spacr_export() # a SingleCellExperiment ## tables <- read_spacr_tables() # a list of data frames ## ## From a shell, to save the SingleCellExperiment as an .rds file: ## Rscript {loader} [export folder] [output.rds] ## ## Reading the tables needs the R package arrow (install.packages("arrow")) ## or nanoparquet (install.packages("nanoparquet")). load_spacr_export() ## also needs Bioconductor's SingleCellExperiment: ## install.packages("BiocManager"); BiocManager::install("SingleCellExperiment") ## jsonlite, when installed, turns the provenance values into R lists. .spacr_export_dir <- local({{ sourced <- NULL for (frame in rev(sys.frames())) {{ candidate <- frame$ofile if (is.character(candidate) && length(candidate) == 1L) {{ sourced <- candidate break }} }} if (!is.null(sourced)) {{ dirname(normalizePath(sourced)) }} else {{ args <- commandArgs(trailingOnly = FALSE) file_arg <- sub("^--file=", "", args[grep("^--file=", args)]) if (length(file_arg) == 1L) dirname(normalizePath(file_arg)) else getwd() }} }}) .spacr_read_parquet <- function(path) {{ if (requireNamespace("arrow", quietly = TRUE)) {{ return(as.data.frame(arrow::read_parquet(path))) }} if (requireNamespace("nanoparquet", quietly = TRUE)) {{ return(as.data.frame(nanoparquet::read_parquet(path))) }} stop("Reading a spaCR export needs the R package arrow ", "(install.packages(\"arrow\")) or nanoparquet ", "(install.packages(\"nanoparquet\")).", call. = FALSE) }} read_spacr_tables <- function(dir = .spacr_export_dir) {{ files <- c({files}) tables <- list() for (name in names(files)) {{ path <- file.path(dir, files[[name]]) if (file.exists(path)) tables[[name]] <- .spacr_read_parquet(path) }} for (key in c({factors})) {{ for (name in c("objects", "wells")) {{ table <- tables[[name]] if (!is.null(table) && key %in% names(table) && !is.factor(table[[key]])) {{ tables[[name]][[key]] <- factor(table[[key]]) }} }} }} tables }} read_spacr_provenance <- function(dir = .spacr_export_dir) {{ table <- .spacr_read_parquet(file.path(dir, "{provenance}")) parse <- if (requireNamespace("jsonlite", quietly = TRUE)) {{ function(text) jsonlite::fromJSON(text, simplifyVector = TRUE) }} else {{ identity }} out <- list() for (section in unique(as.character(table$section))) {{ rows <- table[as.character(table$section) == section, , drop = FALSE] out[[section]] <- stats::setNames( lapply(as.character(rows$value), parse), as.character(rows$key)) }} out }} load_spacr_export <- function(dir = .spacr_export_dir) {{ if (!requireNamespace("SingleCellExperiment", quietly = TRUE)) {{ stop("load_spacr_export() needs Bioconductor's SingleCellExperiment: ", "install.packages(\"BiocManager\"); ", "BiocManager::install(\"SingleCellExperiment\"). ", "read_spacr_tables() works without it.", call. = FALSE) }} tables <- read_spacr_tables(dir) objects <- tables$objects features <- as.character(tables$features${feature}) keys <- as.character(objects${key}) values <- as.matrix(objects[, features, drop = FALSE]) storage.mode(values) <- "double" values <- t(values) dimnames(values) <- list(features, keys) col_data <- objects[, setdiff(names(objects), features), drop = FALSE] rownames(col_data) <- keys row_data <- tables$features rownames(row_data) <- features sce <- SingleCellExperiment::SingleCellExperiment( assays = list(measurements = values), colData = S4Vectors::DataFrame(col_data, check.names = FALSE), rowData = S4Vectors::DataFrame(row_data, check.names = FALSE)) embeddings <- tables$embeddings if (!is.null(embeddings)) {{ columns <- setdiff(names(embeddings), "{key}") rows <- match(keys, as.character(embeddings${key})) for (name in unique(sub("_[0-9]+$", "", columns))) {{ picked <- columns[sub("_[0-9]+$", "", columns) == name] coords <- as.matrix(embeddings[rows, picked, drop = FALSE]) rownames(coords) <- keys SingleCellExperiment::reducedDim(sce, name) <- coords }} }} S4Vectors::metadata(sce) <- list( spacr = read_spacr_provenance(dir), wells = tables$wells) sce }} if (!interactive() && sys.nframe() == 0L) {{ args <- commandArgs(trailingOnly = TRUE) dir <- if (length(args) >= 1L) args[[1]] else .spacr_export_dir out <- if (length(args) >= 2L) args[[2]] else file.path(dir, "{rds}") sce <- load_spacr_export(dir) saveRDS(sce, out) cat("Wrote", out, "with", ncol(sce), "objects and", nrow(sce), "features\n") }} ''' def _r_loader_script(files: Mapping[str, str], version: str) -> str: """The R loader for a table export. :param files: ``{table name: file name}`` of the Parquet tables. :param version: the spaCR version recorded in the header. :returns: the script text. """ listed = ", ".join(f'{name} = "{path}"' for name, path in files.items()) factors = ", ".join(f'"{key}"' for key in ( schema.PLATE_KEY, schema.ROW_KEY, schema.COLUMN_KEY, schema.FIELD_KEY, schema.PRC_KEY, schema.WELL_KEY)) return _R_LOADER_TEMPLATE.format( version=version, date=datetime.now(timezone.utc).strftime("%Y-%m-%d"), loader=_R_LOADER_NAME, files=listed, factors=factors, provenance=files.get("provenance", "provenance.parquet"), feature="feature", key=_OBJECT_KEY_COLUMN, rds="spacr_export_sce.rds") def _register_tables(objects_path: str, db_path: str, project: Union[str, os.PathLike, None], settings: Optional[Mapping[str, Any]], run_id: str, extra: Mapping[str, Any]) -> str: """Register the per-object table, returning its artifact id or ``""``. Never raises: a registry that cannot be opened must not lose a finished export, whose provenance is already in every file it wrote. """ try: from .. import artifacts, ports root = _project_root(db_path, project) inputs: List[str] = [] try: upstream = artifacts.latest(ports.MEASUREMENTS_DB, project=root) if upstream is not None: inputs.append(upstream.artifact_id) except Exception: pass record = artifacts.register( project=root, module=APP_KEY, kind=_TABLES_KIND, role="objects", path=objects_path, settings=settings, inputs=inputs, run_id=run_id, extra=dict(extra)) return record.artifact_id except Exception as exc: warnings.warn( f"the table export at {objects_path} was written but could not " f"be registered with spacr.artifacts ({exc}). Its provenance is " f"still in provenance.parquet.", RuntimeWarning, stacklevel=2) return "" def _export_tables(db_path: Union[str, os.PathLike], out_dir: Union[str, os.PathLike], *, r_loader: bool = False, rds: Optional[bool] = None, well_statistic: str = "mean", dtype: str = "float64", embeddings: Optional[Mapping[str, Any]] = None, compute_umap: bool = False, umap_settings: Optional[Mapping[str, Any]] = None, register: bool = True, project: Union[str, os.PathLike, None] = None, settings: Optional[Mapping[str, Any]] = None, run_id: str = "", timelapse: bool = False, data_filter: Optional[DataFilter] = None, selection: Optional[Selection] = None, verbose: bool = True, **kwargs: Any) -> _TablesResult: """Write a measurements database as tidy Parquet tables, optionally for R. The same objects, features, labels, filter and missing-value policy as :func:`build_anndata`, written as tables into ``out_dir``: ``objects.parquet`` one row per object: ``object_key`` (the key the other tables join on), the plate, row, column and field ids, ``prc`` and ``wellID`` well keys, the metadata, annotation and prediction columns, and one column per feature. Plate and well keys are categoricals. ``features.parquet`` one row per feature column, keyed by ``feature``: the feature dictionary's description, unit, source table and missing counts. ``wells.parquet`` one row per well: its keys, ``n_objects`` and each feature's ``well_statistic`` over the well's objects. Not written when the objects carry no plate, row and column. ``embeddings.parquet`` ``object_key`` plus ``<name>_<k>`` columns, when embeddings were given or ``compute_umap`` is on. ``provenance.parquet`` ``section``, ``key`` and JSON ``value`` rows: the version, run id, source, filter, missing-value report and material settings. Every file also carries a JSON description (table role, key columns, feature columns and the provenance) in its Parquet schema metadata, readable with :func:`spacr.tabular._parquet_metadata` or ``pyarrow.parquet.read_schema``. With ``r_loader`` the folder also gets ``load_spacr_export.R``, which reads the tables with arrow or nanoparquet and builds a SingleCellExperiment (``measurements`` assay, ``colData`` from the objects, ``rowData`` from the features, embeddings as ``reducedDims``, provenance and wells in ``metadata``). :param db_path: a ``measurements.db``. :param out_dir: folder to write into; created if missing. Existing files of the same names are replaced. :param r_loader: write the R loader script. :param rds: also write each table as an ``.rds`` data frame through pyreadr. ``None`` (default) writes them when ``r_loader`` is on and pyreadr is installed, and notes when it is not; ``True`` requires pyreadr; ``False`` never writes them. :param well_statistic: ``'mean'`` or ``'median'``. :param dtype: dtype of the feature columns. ``float64`` keeps the database's values exactly. :param embeddings: as for :func:`build_anndata`. :param compute_umap: as for :func:`build_anndata`. :param umap_settings: as for :func:`build_anndata`. :param register: record ``objects.parquet`` with :mod:`spacr.artifacts`. :param project: the project root the artifact belongs to. :param settings: the run settings, hashed and stored in the provenance. :param run_id: the run this export belongs to. :param timelapse: key each frame of an object separately; the per-well table then has one row per well and time point. :param data_filter: as for :func:`build_anndata`. :param selection: as for :func:`build_anndata`. :param verbose: print the summary. :param kwargs: passed to the table assembly: ``tables``, ``single_table``, ``row_limit``, ``nan_policy``, ``exclude``, ``condition_map``, ``condition_column``, ``attach_labels`` and ``drop_redundant_identity``. :returns: a :class:`_TablesResult`. :raises ImportError: with install instructions when pyarrow is missing, or pyreadr when ``rds=True``. :raises ValueError: on an unknown ``well_statistic`` and everything :func:`build_anndata` raises it for. """ from .. import tabular from ..version import get_version if well_statistic not in _WELL_STATISTICS: raise ValueError( f"well_statistic={well_statistic!r} is not one of " f"{list(_WELL_STATISTICS)}.") tabular._require_optional("pyarrow", tabular._PYARROW_MISSING_MESSAGE) if rds: tabular._require_optional("pyreadr", tabular._PYREADR_MISSING_MESSAGE) out_dir = os.path.abspath(os.path.expanduser(os.fspath(out_dir))) parts = _assemble_tables(db_path, timelapse=timelapse, data_filter=data_filter, selection=selection, **kwargs) db_path = parts["db_path"] notes: List[str] = list(parts["notes"]) features = list(parts["features"]) objects = _objects_table(parts, dtype) keys = pd.Index(objects[_OBJECT_KEY_COLUMN]) embedding_table, embedding_columns = _embeddings_table( keys, parts, embeddings, compute_umap, umap_settings, timelapse, notes) well_keys = _well_key_columns(objects, timelapse) wells = (_wells_table(objects, features, well_keys, well_statistic) if well_keys else None) if wells is None: notes.append( "no wells.parquet: the objects carry no plateID, rowID and " "columnID to group by.") provenance = _provenance_record( parts, n_objects=len(objects), n_features=len(features), data_filter=data_filter, selection=selection, settings=settings, run_id=run_id, timelapse=timelapse) provenance["notes"] = list(notes) names = {"objects": "objects.parquet", "features": "features.parquet"} if wells is not None: names["wells"] = "wells.parquet" if embedding_table is not None: names["embeddings"] = "embeddings.parquet" names["provenance"] = "provenance.parquet" provenance["tables"] = { "format": _TABLES_FORMAT, "format_version": _TABLES_FORMAT_VERSION, "files": dict(names), "object_key": _OBJECT_KEY_COLUMN, "well_key_columns": list(well_keys), "well_statistic": well_statistic, "embeddings": embedding_columns, } provenance = _h5ad_safe(provenance) features_table = parts["var"].copy() features_table.index = pd.Index(features, name="feature") features_table = features_table.reset_index() for column in features_table.columns: if features_table[column].dtype == object: features_table[column] = features_table[column].astype(str) metadata_columns = [c for c in objects.columns if c not in set(features)] frames = { "objects": (objects, { "feature_columns": features, "metadata_columns": metadata_columns, "annotation_columns": list(parts["annotations"]), "prediction_columns": list(parts["predictions"]), }), "features": (features_table, {"key_column": "feature"}), } if wells is not None: frames["wells"] = (wells, { "feature_columns": features, "statistic": well_statistic, "count_column": _WELL_COUNT_COLUMN}) if embedding_table is not None: frames["embeddings"] = (embedding_table, {"embeddings": embedding_columns}) frames["provenance"] = (_provenance_table(provenance, settings), { "columns": {"section": "'export' or 'settings'", "key": "the provenance field or settings key", "value": "JSON text"}}) written: List[str] = [] for name, (frame, extra) in frames.items(): description = { "format": _TABLES_FORMAT, "format_version": _TABLES_FORMAT_VERSION, "table": name, "object_key": _OBJECT_KEY_COLUMN, "well_key_columns": list(well_keys), **extra, "provenance": provenance, } written.append(tabular._write_parquet( frame, os.path.join(out_dir, names[name]), metadata=description, canonicalise=name in ("objects", "wells"))) loader = "" if r_loader: loader = os.path.join(out_dir, _R_LOADER_NAME) with open(loader, "w", encoding="utf-8", newline="\n") as handle: handle.write(_r_loader_script(names, get_version())) written.append(loader) rds_files: List[str] = [] want_rds = r_loader if rds is None else bool(rds) if want_rds: try: tabular._require_optional( "pyreadr", tabular._PYREADR_MISSING_MESSAGE) except ImportError as exc: notes.append(str(exc).splitlines()[0] + " No .rds data frames " "were written; load_spacr_export.R reads the " "Parquet tables instead.") else: for name, (frame, _extra) in frames.items(): rds_files.append(tabular._write_rds( frame, os.path.join(out_dir, f"{name}.rds"), canonicalise=name in ("objects", "wells"))) written.extend(rds_files) artifact_id = "" if register: artifact_id = _register_tables( written[0], db_path, project, settings, str(provenance.get("run_id", "")), {"n_objects": len(objects), "n_features": len(features), "n_wells": 0 if wells is None else len(wells), "files": [os.path.basename(path) for path in written], "source_database": db_path}) result = _TablesResult( directory=out_dir, files=tuple(written), n_objects=len(objects), n_features=len(features), n_wells=0 if wells is None else len(wells), r_loader=loader, rds_files=tuple(rds_files), artifact_id=artifact_id, warnings=tuple(notes)) if verbose: print(result.describe()) return result def _default_tables_dir(src: Union[str, os.PathLike], single_table: str = "") -> str: """Where the tables land when ``anndata_tidy_dir`` is empty. ``<project>/results/<project name>_tables``, with the object table appended to the name for a single-table export. """ root = _project_root(resolve_db_path(src), None) name = os.path.basename(root.rstrip(os.sep)) or "spacr" if single_table: name = f"{name}_{single_table}" return os.path.join(root, "results", f"{name}_tables")
[docs] def anndata_export_settings(settings: Optional[Mapping[str, Any]] = None ) -> Dict[str, Any]: """Return this module's settings, filling in anything absent. :param settings: an existing settings dict to complete. :returns: a new dict; the caller's is not mutated. """ resolved = dict(settings or {}) resolved.setdefault("src", "") resolved.setdefault("anndata_out", "") resolved.setdefault("anndata_tables", list(DEFAULT_TABLES)) resolved.setdefault("anndata_single_table", "") resolved.setdefault("anndata_nan_policy", NAN_KEEP) resolved.setdefault("anndata_dtype", "float32") resolved.setdefault("anndata_row_limit", 0) resolved.setdefault("anndata_compute_umap", False) resolved.setdefault("anndata_compression", "gzip") resolved.setdefault("anndata_register_artifact", True) resolved.setdefault("anndata_format", _FORMAT_H5AD) resolved.setdefault("anndata_tidy_dir", "") return resolved
[docs] def resolve_db_path(src: Union[str, os.PathLike]) -> str: """The ``measurements.db`` a settings ``src`` means. ``src`` is a project root everywhere else in spaCR, and this module's argument is a database, so one of the two has to give. A path that ends in a database file is taken as one, and anything else is read as a project root laid out the way every spaCR writer leaves it. :param src: project root, or the database itself. :returns: an absolute path, which is NOT checked for existence -- an absent file is :func:`spacr.validate.validate_settings`'s report to make, with the path in it, rather than an exception from here. """ src = os.path.abspath(os.path.expanduser(os.fspath(src))) if os.path.splitext(src)[1].lower() in (".db", ".sqlite", ".sqlite3"): return src return os.path.join(src, "measurements", "measurements.db")
[docs] def default_out_path(src: Union[str, os.PathLike], single_table: str = "") -> str: """Where the export lands when ``anndata_out`` is empty. ``<project>/results/<project name>.h5ad`` -- beside the other things a finished run produced, named after the project, because a folder of ``export.h5ad`` files is a folder nobody can tell apart. :param src: project root, or the database. :param single_table: the one object table being exported, if any; it joins the file name, since a nucleus-level export and a cell-level one of the same project are different files. :returns: an absolute ``.h5ad`` path. """ root = _project_root(resolve_db_path(src), None) name = os.path.basename(root.rstrip(os.sep)) or "spacr" if single_table: name = f"{name}_{single_table}" return os.path.join(root, "results", f"{name}.h5ad")
[docs] def run_anndata_export(settings: Optional[Mapping[str, Any]] = None ) -> Union[ExportResult, "_TablesResult"]: """Run the export from a settings dict. The headless entry point. The ``fn(settings)`` shape ``spacr-run``, the Qt Run button and :mod:`spacr.validate` all dispatch to, wrapped around :func:`export_anndata` -- which keeps its own explicit keyword signature, because a function whose arguments are a dict is a function nobody can call from a notebook. Every key it reads is one :func:`anndata_export_settings` declares and :func:`register_anndata_settings` gave a type and a tooltip, so the form the GUI draws and the keys honoured here are the same list. ``anndata_format`` chooses what is written: ``'h5ad'`` the AnnData file, ``'parquet'`` the tidy Parquet tables of :func:`_export_tables` in ``anndata_tidy_dir``, ``'r'`` those tables with the R loader script (and ``.rds`` data frames when pyreadr is installed), and ``'all'`` every one of them. The tables keep float64 features whatever ``anndata_dtype`` says, which sets the AnnData matrix only. :param settings: the run settings. ``src`` is the project root (or the database); everything else falls back to :func:`anndata_export_settings`. :returns: the :class:`ExportResult` when an ``.h5ad`` was written, and otherwise the tables result; ``describe()`` on either is what the console prints. :raises ValueError: when ``src`` is empty -- there is nothing to export and no path to name in the message otherwise -- or ``anndata_format`` is not a known format. :raises AnnDataExtraMissing: when ``anndata`` is not installed and the format includes ``.h5ad``. :raises ImportError: when pyarrow is not installed and the format includes the tables. """ resolved = anndata_export_settings(settings) src = str(resolved.get("src") or "").strip() if not src: raise ValueError( "anndata_export needs src: the spaCR project whose " "measurements/measurements.db is exported.") db_path = resolve_db_path(src) single_table = str(resolved.get("anndata_single_table") or "").strip() out_path = str(resolved.get("anndata_out") or "").strip() if not out_path: out_path = default_out_path(src, single_table) tables = resolved.get("anndata_tables") or list(DEFAULT_TABLES) if isinstance(tables, str): tables = [part.strip() for part in tables.split(",") if part.strip()] row_limit = int(resolved.get("anndata_row_limit") or 0) compression = str(resolved.get("anndata_compression") or "").strip() export_format = str( resolved.get("anndata_format") or _FORMAT_H5AD).strip().lower() if export_format not in _EXPORT_FORMATS: raise ValueError( f"anndata_format={export_format!r} is not one of " f"{list(_EXPORT_FORMATS)}.") writes_h5ad = export_format in (_FORMAT_H5AD, _FORMAT_ALL) if writes_h5ad: require_anndata() register = bool(resolved.get("anndata_register_artifact", True)) common = dict( tables=tuple(tables), single_table=single_table or None, row_limit=row_limit or None, nan_policy=str(resolved.get("anndata_nan_policy") or NAN_KEEP), compute_umap=bool(resolved.get("anndata_compute_umap", False)), ) tables_result = None if export_format != _FORMAT_H5AD: tables_dir = str(resolved.get("anndata_tidy_dir") or "").strip() tables_result = _export_tables( db_path, tables_dir or _default_tables_dir(src, single_table), r_loader=export_format in (_FORMAT_R, _FORMAT_ALL), register=register, settings=resolved, **common) if not writes_h5ad: return tables_result return export_anndata( db_path, out_path, compression=compression or None, register=register, settings=resolved, dtype=str(resolved.get("anndata_dtype") or "float32"), **common, )
_TYPES = { "anndata_out": str, "anndata_tables": list, "anndata_single_table": str, "anndata_nan_policy": str, "anndata_dtype": str, "anndata_row_limit": int, "anndata_compute_umap": bool, "anndata_compression": str, "anndata_register_artifact": bool, "anndata_format": str, "anndata_tidy_dir": str, } _TOOLTIPS = { "anndata_out": ( "(str) - Path of the .h5ad file to write. An existing file at this " "path is overwritten, and the export is registered against it in " "artifacts.db, so writing twice to one path replaces the earlier " "record rather than adding a second. Default " "<src>/results/<project>.h5ad."), "anndata_tables": ( "(list) - Object tables joined into the cell-anchored export. " "Dropping a table drops its features from X: without 'pathogen' " "there are no pathogen columns and no count_pathogen, and without " "'png_list' the crop paths, annotations and model scores are all " "absent from obs. Default ['cell', 'cytoplasm', 'nucleus', " "'pathogen', 'png_list']."), "anndata_single_table": ( "(str) - Export this one object table instead of the join, one row per object of that type. The only way to get a nucleus-level or pathogen-level matrix: the join averages children onto their parent cell. Empty means the joined export. Default ''."), "anndata_nan_policy": ( "(str) - What happens to missing values in X: 'keep' (default; " "AnnData stores them, scanpy's scale/pca/neighbors do not), " "'drop_features', 'drop_objects', 'zero' or 'mean'. The two " "imputing policies also write a layers['missing'] mask."), "anndata_dtype": ( "(str) - dtype of X. Default 'float32' - the scanpy convention and " "half the memory of float64."), "anndata_row_limit": ( '(int) - Maximum number of objects exported after filtering. 0 disables the limit. When a limit is set, the first N objects in table order are retained. Default 0.'), "anndata_compute_umap": ( "(bool) - Compute obsm['X_umap'] during the export, through the " "same reducer the UMAP app uses. Off by default: it costs minutes " "on a large table."), "anndata_compression": ( "(str) - HDF5 compression for X and the layers. 'gzip' (default) " "typically halves the file; '' writes it uncompressed."), "anndata_register_artifact": ( "(bool) - Record the written file with spacr.artifacts, so a re-run " "of Measure marks the export stale. Default True."), "anndata_format": ( "(str) - What the export writes. 'h5ad' writes the AnnData file; " "'parquet' writes tidy Parquet tables (objects, features, wells, " "embeddings, provenance) that pandas, arrow and R read with their " "types kept; 'r' adds load_spacr_export.R, which builds a " "SingleCellExperiment in R, and .rds data frames when pyreadr is " "installed; 'all' writes every one. The tables need pyarrow. " "Default 'h5ad'."), "anndata_tidy_dir": ( "(str) - Folder for the Parquet tables and the R loader when " "anndata_format is 'parquet', 'r' or 'all'. Files of the same " "names in it are replaced. Default <src>/results/<project>_tables."), } _DESCRIPTION = ( "Export the measurement tables as AnnData (.h5ad) - N objects x M " "features with per-object metadata, feature definitions, embeddings and " "provenance - so scanpy, scvi-tools and squidpy can read a spaCR run " "directly." )
[docs] def register_anndata_settings(replace: bool = False) -> bool: """Register this module's settings through the defaults seam. Uses :func:`spacr.settings.register_defaults` rather than appending to ``spacr/settings.py``, so this module owns its own knobs and adding one is not a merge conflict in a file nobody owns. :param replace: re-register over an existing registration. :returns: True if it registered, False if it already was. """ from ..settings import has_registered_defaults, register_defaults if has_registered_defaults(APP_KEY) and not replace: return False register_defaults( APP_KEY, anndata_export_settings, replace=replace, expected_types=_TYPES, tooltips=_TOOLTIPS, categories={ "General": ["anndata_out", "anndata_single_table", "anndata_nan_policy", "anndata_format"], "Advanced": ["anndata_tables", "anndata_dtype", "anndata_row_limit", "anndata_compute_umap", "anndata_compression", "anndata_register_artifact", "anndata_tidy_dir"], }, description=_DESCRIPTION) return True
register_anndata_settings()