spacr.anndata_export

Workflow inputs and outputs

AnnData Export

Export an AnnData .h5ad file with object metadata and measured features for downstream single-cell analysis.

Open: Measure → AnnData Export.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route: cell, nucleus, pathogen, cytoplasm. Relevant columns, depending on the route: plateID, rowID, columnID, fieldID.

Outputs

  • AnnData export — An exported .h5ad file containing measured features and object metadata.

Before this module

  • Measure: Export compatible feature and metadata columns.

API reference.

Module tutorial.

Export a spaCR measurements.db as AnnData (.h5ad).

spaCR already produces the exact shape AnnData was designed for: N objects x M features, plus per-object metadata and an embedding. Writing that out as .h5ad puts scanpy, scvi-tools and squidpy within reach of a spaCR user for the cost of one function call, instead of a bespoke join script per lab.

Nothing here invents an identity, a filter or a feature/metadata boundary. All four already exist in spaCR and are reused verbatim:

obs_names

spacr.selection.object_keys() over spacr.selection.OBJECT_KEY_COLUMNS – the schema’s own row key (plateID, rowID, columnID, fieldID, object_label), joined on spacr.schema.KEY_SEPARATOR. A key from a UMAP lasso names the same object in an exported .h5ad.

X / obs split

spacr.schema.is_provenance_column(), which is the boundary spacr.feature_dict and every model path already use, plus spacr.agreement’s knowledge of which columns a model wrote and which a human did.

var

spacr.feature_dict.parse_column(), resolved against the row’s own measurement_units so a 3-D run’s cell_area is documented as the volume it is rather than the area it is named after.

Filtering

spacr.selection.DataFilter and spacr.selection.Selection – the same objects the GUI’s linked views publish, so “export what I am looking at” is one call and not a second filter language.

Typical use:

from spacr.anndata_export import export_anndata
from spacr.selection import DataFilter, RangeFilter

result = export_anndata(
    "/data/exp1/measurements/measurements.db",
    "/data/exp1/results/exp1.h5ad",
    data_filter=DataFilter().add(RangeFilter("cell_area", low=200)),
)
print(result.describe())

import scanpy as sc
adata = sc.read_h5ad("/data/exp1/results/exp1.h5ad")

What happens to NaN

Nothing, by default, and that is a decision rather than an omission.

AnnData stores NaN happily. Most of what people reach for next does not: sc.pp.scale, sc.pp.pca, sc.pp.neighbors and essentially all of scvi-tools either propagate NaN across the whole matrix or raise. So a silent default matters, and there are only two honest candidates: keep the NaN and say so loudly, or impute and say so loudly. Quietly filling with zero is not one of them – in spaCR a NaN is usually meaningful: a pathogen_* column is NaN for a cell with no pathogen in it, a Zernike column is NaN when mahotas was not installed, and a correlation column is NaN when one channel was flat. Zero-filling the first turns “no pathogen” into “a pathogen of zero size” and puts it into the same distribution as a measured one.

NAN_KEEP (the default) therefore writes the NaN through, and pays for it with visibility:

  • var['n_missing'] / var['frac_missing'] – per feature;

  • obs['n_missing_features'] – per object;

  • uns['spacr']['nan'] – totals, the policy that was applied, the shape those totals were counted over (n_objects_counted x n_features_counted, i.e. post-filter and pre-policy), and the ten worst columns by name;

  • a printed warning naming those columns and the scanpy calls that will fail on them.

Every count in that record – and ExportResult.n_missing with it – is taken on the matrix as the policy received it, which for a dropping policy is bigger than the one written. ExportResult.frac_missing divides by that same matrix (ExportResult.counted_shape) rather than by the written shape, so it stays a fraction.

The alternatives are explicit, and every one of them records what it did:

NAN_DROP_FEATURES

drop any feature column containing a NaN. Cheap and safe; on a real database it often removes every pathogen_* column, which is the honest consequence of asking for a complete matrix.

NAN_DROP_OBJECTS

drop any object row containing a NaN. Usually removes almost everything, for the same reason; offered because on a single-object-type export it is often exactly right.

NAN_ZERO / NAN_MEAN

impute. Both write a layers['missing'] boolean mask of the same shape by default, because an imputed matrix that cannot be told from a measured one is a trap, and 1 byte per cell against 4 is a fair price.

Infinities are treated as missing under every policy. spaCR produces +/-inf from ratio features (a denominator of zero), and an inf survives dropna while destroying any scaling, PCA or distance computed from it. They are converted to NaN before the policy runs and counted separately in uns['spacr']['nan']['n_infinite'] so the substitution is never silent.

The labels the table join drops

spacr.io._read_and_join_tables() takes six named columns off png_list – the object id, png_path and the four field keys – and drops the rest. The rest is every annotation column the Annotate app added and every score the classifier wrote, which for this export are among the most valuable columns in the database: without them obs has no label to train on, group by, or colour a UMAP with. They are re-attached by object key (_attach_png_labels()) before filtering, so a filter can name one – “export the cells I annotated as infected” – and they are separated into uns['spacr']['annotation_columns'] and uns['spacr']['prediction_columns'] using spacr.agreement’s own rule, which exists because a model’s class column is indistinguishable by shape from an annotation pass.

They land in obs, never in X: a label is not a measurement, and scaling, PCA and clustering it alongside the features is how a classifier’s own output ends up as a “feature” that predicts it.

Provenance

uns['spacr'] carries the spaCR version, the settings hash (spacr.artifacts.settings_hash()), the run id, the absolute source database path, the tables read, the filter that was applied and the counts before and after it. The written file is also registered with spacr.artifacts under kind ANNDATA_KIND, with the project’s measurements-db artifact as its input – so spacr.artifacts.is_stale reports the export as stale the moment Measure is re-run, which is the whole reason the registry exists.

The optional dependency

anndata is an optional extra: pip install "spacr[anndata]". It is imported inside the functions that need it, never at module scope, so import spacr.anndata_export works without it and the failure is one actionable sentence (AnnDataExtraMissing) rather than a traceback six frames deep.

Attributes

Exceptions

AnnDataExtraMissing

anndata is not installed.

DuplicateObjectKeys

Two rows claim the same object key.

Classes

ExportResult

What one export produced, and what it left out.

Functions

anndata_export_settings(→ Dict[str, Any])

Return this module's settings, filling in anything absent.

build_anndata(db_path, *[, tables, single_table, ...])

Build an anndata.AnnData from a spaCR measurements database.

default_out_path(→ str)

Where the export lands when anndata_out is empty.

export_anndata(→ ExportResult)

Write a spaCR measurements database to out_path as .h5ad.

export_anndata_set(, prefix, **kwargs, ExportResult])

Write one .h5ad per object table, cross-referenced by uns.

feature_columns(, db_path)

The columns of frame that belong in X, in frame order.

register_anndata_settings(→ bool)

Register this module's settings through the defaults seam.

require_anndata()

Import and return anndata, or raise a message worth reading.

resolve_db_path(→ str)

The measurements.db a settings src means.

run_anndata_export(→ Union[ExportResult, _TablesResult])

Run the export from a settings dict. The headless entry point.

Package Contents

exception spacr.anndata_export.AnnDataExtraMissing[source]

Bases: ImportError

anndata is not installed.

An ImportError subclass so a caller that already guards the export with except ImportError keeps working, and so the message – not a traceback through anndata’s own import machinery – is what reaches the user.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.anndata_export.DuplicateObjectKeys[source]

Bases: ValueError

Two rows claim the same object key.

AnnData requires unique obs_names, and spaCR’s writers append: two rows can share all five key columns when a field was measured twice (see tests/test_db_contract.py, which measures exactly that). This is raised rather than repaired, because deduplicating – keeping the first, keeping the last, or averaging – changes the numbers and is the caller’s decision, not the exporter’s.

Initialize self. See help(type(self)) for accurate signature.

class spacr.anndata_export.ExportResult[source]

What one export produced, and what it left out.

Parameters:
  • path – the .h5ad written, or "" for an in-memory build.

  • n_obs – objects in the exported matrix.

  • n_vars – feature columns in X.

  • n_obs_before_filter – objects read from the database.

  • obs_columns – the obs column names.

  • obsm_keys – the obsm keys written.

  • nan_policy – the policy that was applied.

  • n_missing – NaN cells in X before the policy ran.

  • n_infinite – non-finite cells converted to NaN before that count.

  • dropped_features – feature columns removed by the policy.

  • dropped_objects – objects removed by the policy.

  • n_obs_counted – rows of the matrix n_missing was counted in – i.e. after data_filter/selection/row_limit and before the nan_policy ran. 0 on a record that did not record it, in which case counted_shape reconstructs it from the drops.

  • n_vars_counted – columns of that same pre-policy matrix.

  • artifact_id – the spacr.artifacts id, or "" when the file was not registered.

  • warnings – everything the export decided the user must know.

Three object counts live here and they are three different stages, which is the whole reason they are named apart: n_obs_before_filter is what the database held, counted_shape is what survived the filter and was handed to the NaN policy, and n_obs is what was written. describe() attributes each loss to the stage that caused it rather than charging all of them to the filter.

describe() → str[source]

One human paragraph: shape, filtering, missingness, provenance.

Each count is charged to the stage that caused it: the filter line counts only what the filter removed, and the dropped ... lines only what the NaN policy removed. Charging the policy’s row drops to the filter as well made one loss of four objects read as eight.

property counted_shape: Tuple[int, int][source]

(rows, columns) of the matrix n_missing was counted in.

That matrix is the post-filter, pre-nan_policy one: the NaN were counted before anything was dropped or imputed. A record built by this module states it outright; one built by hand (or by an older spaCR) leaves the two fields at 0, and the shape is reconstructed from the drops instead – drop_objects is the only policy that removes rows and drop_features the only one that removes columns, so adding them back recovers what the policy was given.

property frac_missing: float[source]

n_missing over the cells of counted_shape.

Numerator and denominator describe the same matrix – the one the NaN were counted in, before the policy dropped or imputed anything – so this is a fraction and can never exceed 1.0. Dividing the pre-policy count by the post-policy n_obs * n_vars is what made drop_objects report 114.3% missing.

Returns:

the fraction, or 0.0 for an empty matrix.

spacr.anndata_export.anndata_export_settings(settings: Mapping[str, Any] | None = None) → Dict[str, Any][source]

Return this module’s settings, filling in anything absent.

Parameters:

settings – an existing settings dict to complete.

Returns:

a new dict; the caller’s is not mutated.

spacr.anndata_export.build_anndata(db_path: str | os.PathLike, *, tables: Sequence[str] = DEFAULT_TABLES, single_table: str | None = None, data_filter: spacr.selection.DataFilter | None = None, selection: spacr.selection.Selection | None = None, row_limit: int | None = None, timelapse: bool = False, nan_policy: str = NAN_KEEP, missing_layer: bool | None = None, dtype: str = 'float32', exclude: Sequence[str] = (), embeddings: Mapping[str, Any] | None = None, compute_umap: bool = False, umap_settings: Mapping[str, Any] | None = None, condition_map: Mapping[str, str] | None = None, condition_column: str = schema.COLUMN_KEY, attach_labels: bool = True, drop_redundant_identity: bool = True, settings: Mapping[str, Any] | None = None, run_id: str = '', verbose: bool = True)[source]

Build an anndata.AnnData from a spaCR measurements database.

The mapping is described in full in the module docstring; in short, X is the numeric measurements, obs is everything else about the object, var is spacr.feature_dict’s description of each feature, obsm holds embeddings and uns holds provenance.

Parameters:
  • db_path – a measurements.db.

  • tables – tables to join for the default cell-anchored export.

  • single_table – export exactly this object table instead, one row per object of that type – which is the only way to get a nucleus-level or pathogen-level matrix, since the join averages children onto their parent.

  • data_filter – a spacr.selection.DataFilter. Declarative and re-appliable; recorded in uns.

  • selection – a spacr.selection.Selection – the keys a view pointed at. Applied after data_filter.

  • row_limit – hard cap on exported objects, applied last. A blunt instrument on purpose, for “give me something I can open” without inventing a filter that means something it does not.

  • timelapse – key each frame of an object separately.

  • nan_policy – one of NAN_POLICIES; see the module docstring. The default keeps NaN and reports it.

  • missing_layer – write layers['missing']. Defaults to True for the imputing policies and False otherwise.

  • dtype – X dtype. float32 by default – the scanpy convention, half the memory, and far more precision than any microscope measurement carries.

  • exclude – feature columns to keep out of X.

  • embeddings – {name: array or DataFrame}. Names are normalised to the scanpy X_* convention, so 'umap' becomes X_umap. A DataFrame carrying spacr.selection.OBJECT_KEY_COLUMNS is aligned by key, which is what makes an embedding computed on the whole plate usable with a filtered export.

  • compute_umap – compute X_umap here, through the same spacr.utils.reduction_and_clustering() call spacr.core.generate_image_umap() makes. Off by default: it imports the segmentation stack and costs minutes on a large table.

  • umap_settings – overrides for that computation.

  • condition_map – {column value: label} written to obs['condition']; DEFAULT_CONDITION_MAP is the mapping spacr.utils.map_condition() applies.

  • condition_column – which column condition_map reads.

  • attach_labels – bring the annotation and prediction columns back out of png_list, which the table join drops. On by default – see _attach_png_labels().

  • drop_redundant_identity – drop the join’s suffixed copies of identity columns (plateID_nucleus, object_label_pathogen) from obs. On by default; see _redundant_identity_columns() for why two of them are worse than merely duplicated.

  • settings – the run settings, hashed into uns provenance.

  • run_id – the run this export belongs to.

  • verbose – print the summary and the missing-data warning.

Returns:

(adata, result) – the AnnData and an ExportResult.

Raises:
spacr.anndata_export.default_out_path(src: str | os.PathLike, single_table: str = '') → str[source]

Where the export lands when anndata_out is empty.

<project>/results/<project name>.h5ad – beside the other things a finished run produced, named after the project, because a folder of export.h5ad files is a folder nobody can tell apart.

Parameters:
  • src – project root, or the database.

  • single_table – the one object table being exported, if any; it joins the file name, since a nucleus-level export and a cell-level one of the same project are different files.

Returns:

an absolute .h5ad path.

spacr.anndata_export.export_anndata(db_path: str | os.PathLike, out_path: str | os.PathLike, *, compression: str | None = 'gzip', register: bool = True, project: str | os.PathLike | None = None, settings: Mapping[str, Any] | None = None, **kwargs: Any) → ExportResult[source]

Write a spaCR measurements database to out_path as .h5ad.

Everything build_anndata() accepts is accepted here and passed through; this adds the write and the artifact registration. A failed write preserves any existing destination; only a completed file replaces it and is registered as an artifact.

Parameters:
  • db_path – a measurements.db.

  • out_path – the .h5ad to write. Parent directories are created.

  • compression – h5py compression for X and the layers; None for an uncompressed file. gzip typically halves a feature matrix and costs a few seconds.

  • register – record the file with spacr.artifacts, with the project’s measurements-db as its input so a re-run of Measure marks this export stale.

  • project – the project root the artifact belongs to. Inferred from db_path (<project>/measurements/measurements.db) when omitted.

  • settings – the run settings, hashed into the artifact and uns.

  • kwargs – passed to build_anndata().

Returns:

the ExportResult, with path and artifact_id filled in.

Raises:

AnnDataExtraMissing – when anndata is not installed.

spacr.anndata_export.export_anndata_set(db_path: str | os.PathLike, out_dir: str | os.PathLike, *, object_tables: Sequence[str] = ('cell', 'nucleus', 'pathogen', 'cytoplasm'), prefix: str = '', **kwargs: Any) → Dict[str, ExportResult][source]

Write one .h5ad per object table, cross-referenced by uns.

The honest shape for a multi-compartment experiment: a nucleus is not a cell and averaging it onto one loses the distribution. Each file holds one object type at its own granularity, and each child file records its parent in uns['spacr']['relationships'] – table, key column and the sibling file – so the set can be reassembled (or handed to MuData) without guessing.

Parameters:
  • db_path – a measurements.db.

  • out_dir – folder to write into; created if missing.

  • object_tables – which object tables to export. Tables absent from the database are skipped, not an error.

  • prefix – prepended to each file name.

  • kwargs – passed to export_anndata().

Returns:

{table: ExportResult} for the tables actually written.

spacr.anndata_export.feature_columns(frame: pandas.DataFrame, *, exclude: Sequence[str] = (), db_path: str | None = None) → List[str][source]

The columns of frame that belong in X, in frame order.

A column is a feature when it is numeric, is not provenance or identity by spacr.schema.is_provenance_column(), is not a human annotation or a model output, is not the cluster label, and was not excluded by the caller.

Parameters:
  • frame – an object frame read from a measurements database.

  • exclude – extra column names to keep out of X.

  • db_path – the source database, used to recognise human annotation columns; omit it and only the model columns are recognised.

Returns:

feature column names.

spacr.anndata_export.register_anndata_settings(replace: bool = False) → bool[source]

Register this module’s settings through the defaults seam.

Uses spacr.settings.register_defaults() rather than appending to spacr/settings.py, so this module owns its own knobs and adding one is not a merge conflict in a file nobody owns.

Parameters:

replace – re-register over an existing registration.

Returns:

True if it registered, False if it already was.

spacr.anndata_export.require_anndata()[source]

Import and return anndata, or raise a message worth reading.

Returns:

the imported anndata module.

Raises:

AnnDataExtraMissing – when the extra is not installed.

spacr.anndata_export.resolve_db_path(src: str | os.PathLike) → str[source]

The measurements.db a settings src means.

src is a project root everywhere else in spaCR, and this module’s argument is a database, so one of the two has to give. A path that ends in a database file is taken as one, and anything else is read as a project root laid out the way every spaCR writer leaves it.

Parameters:

src – project root, or the database itself.

Returns:

an absolute path, which is NOT checked for existence – an absent file is spacr.validate.validate_settings()’s report to make, with the path in it, rather than an exception from here.

spacr.anndata_export.run_anndata_export(settings: Mapping[str, Any] | None = None) → ExportResult | _TablesResult[source]

Run the export from a settings dict. The headless entry point.

The fn(settings) shape spacr-run, the Qt Run button and spacr.validate all dispatch to, wrapped around export_anndata() – which keeps its own explicit keyword signature, because a function whose arguments are a dict is a function nobody can call from a notebook.

Every key it reads is one anndata_export_settings() declares and register_anndata_settings() gave a type and a tooltip, so the form the GUI draws and the keys honoured here are the same list.

anndata_format chooses what is written: 'h5ad' the AnnData file, 'parquet' the tidy Parquet tables of _export_tables() in anndata_tidy_dir, 'r' those tables with the R loader script (and .rds data frames when pyreadr is installed), and 'all' every one of them. The tables keep float64 features whatever anndata_dtype says, which sets the AnnData matrix only.

Parameters:

settings – the run settings. src is the project root (or the database); everything else falls back to anndata_export_settings().

Returns:

the ExportResult when an .h5ad was written, and otherwise the tables result; describe() on either is what the console prints.

Raises:
  • ValueError – when src is empty – there is nothing to export and no path to name in the message otherwise – or anndata_format is not a known format.

  • AnnDataExtraMissing – when anndata is not installed and the format includes .h5ad.

  • ImportError – when pyarrow is not installed and the format includes the tables.

spacr.anndata_export.ANNDATA_MISSING_MESSAGE = Multiline-String[source]
Show Value
"""Exporting to AnnData (.h5ad) needs the optional `anndata` extra, which is
not installed in this environment (missing module: {module}).

Install it with:

    python -m pip install "spacr[anndata]"

scanpy is a separate install and is NOT required to write the file:

    python -m pip install scanpy"""