spacr.anndata_export¶
Workflow inputs and outputs¶
AnnData Export¶
Export an AnnData .h5ad file with object metadata and measured features for downstream single-cell analysis.
Open: Measure → AnnData Export.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route:
cell,nucleus,pathogen,cytoplasm. Relevant columns, depending on the route:plateID,rowID,columnID,fieldID.
Outputs
AnnData export — An exported .h5ad file containing measured features and object metadata.
Before this module
Measure: Export compatible feature and metadata columns.
Export a spaCR measurements.db as AnnData (.h5ad).
spaCR already produces the exact shape AnnData was designed for: N objects
x M features, plus per-object metadata and an embedding. Writing that out as
.h5ad puts scanpy, scvi-tools and squidpy within reach of a spaCR user
for the cost of one function call, instead of a bespoke join script per lab.
Nothing here invents an identity, a filter or a feature/metadata boundary. All four already exist in spaCR and are reused verbatim:
obs_namesspacr.selection.object_keys()overspacr.selection.OBJECT_KEY_COLUMNS– the schema’s own row key (plateID,rowID,columnID,fieldID,object_label), joined onspacr.schema.KEY_SEPARATOR. A key from a UMAP lasso names the same object in an exported.h5ad.X/obssplitspacr.schema.is_provenance_column(), which is the boundaryspacr.feature_dictand every model path already use, plusspacr.agreement’s knowledge of which columns a model wrote and which a human did.varspacr.feature_dict.parse_column(), resolved against the row’s ownmeasurement_unitsso a 3-D run’scell_areais documented as the volume it is rather than the area it is named after.- Filtering
spacr.selection.DataFilterandspacr.selection.Selection– the same objects the GUI’s linked views publish, so “export what I am looking at” is one call and not a second filter language.
Typical use:
from spacr.anndata_export import export_anndata
from spacr.selection import DataFilter, RangeFilter
result = export_anndata(
"/data/exp1/measurements/measurements.db",
"/data/exp1/results/exp1.h5ad",
data_filter=DataFilter().add(RangeFilter("cell_area", low=200)),
)
print(result.describe())
import scanpy as sc
adata = sc.read_h5ad("/data/exp1/results/exp1.h5ad")
What happens to NaN¶
Nothing, by default, and that is a decision rather than an omission.
AnnData stores NaN happily. Most of what people reach for next does not:
sc.pp.scale, sc.pp.pca, sc.pp.neighbors and essentially all of
scvi-tools either propagate NaN across the whole matrix or raise. So a
silent default matters, and there are only two honest candidates: keep the
NaN and say so loudly, or impute and say so loudly. Quietly filling with
zero is not one of them – in spaCR a NaN is usually meaningful: a
pathogen_* column is NaN for a cell with no pathogen in it, a Zernike
column is NaN when mahotas was not installed, and a correlation column is
NaN when one channel was flat. Zero-filling the first turns “no pathogen”
into “a pathogen of zero size” and puts it into the same distribution as a
measured one.
NAN_KEEP (the default) therefore writes the NaN through, and pays
for it with visibility:
var['n_missing']/var['frac_missing']– per feature;obs['n_missing_features']– per object;uns['spacr']['nan']– totals, the policy that was applied, the shape those totals were counted over (n_objects_countedxn_features_counted, i.e. post-filter and pre-policy), and the ten worst columns by name;a printed warning naming those columns and the scanpy calls that will fail on them.
Every count in that record – and ExportResult.n_missing with it –
is taken on the matrix as the policy received it, which for a dropping
policy is bigger than the one written. ExportResult.frac_missing divides
by that same matrix (ExportResult.counted_shape) rather than by the
written shape, so it stays a fraction.
The alternatives are explicit, and every one of them records what it did:
NAN_DROP_FEATURESdrop any feature column containing a NaN. Cheap and safe; on a real database it often removes every
pathogen_*column, which is the honest consequence of asking for a complete matrix.NAN_DROP_OBJECTSdrop any object row containing a NaN. Usually removes almost everything, for the same reason; offered because on a single-object-type export it is often exactly right.
NAN_ZERO/NAN_MEANimpute. Both write a
layers['missing']boolean mask of the same shape by default, because an imputed matrix that cannot be told from a measured one is a trap, and 1 byte per cell against 4 is a fair price.
Infinities are treated as missing under every policy. spaCR produces
+/-inf from ratio features (a denominator of zero), and an inf survives
dropna while destroying any scaling, PCA or distance computed from it.
They are converted to NaN before the policy runs and counted separately in
uns['spacr']['nan']['n_infinite'] so the substitution is never silent.
Where the cell -> nucleus -> pathogen links go, and why not obsp¶
They go in obs and uns. Never obsp.
obsp is an n_obs x n_obs matrix: a relation among the observations
of this AnnData. The parent links are not that, in either export shape:
In the joined, cell-anchored export (the default) the nuclei and pathogens are not observations at all – the join in
spacr.io._read_and_join_tables()has already averaged them onto their parent cell. There is no row for anobspentry to point at.In a per-table export of
nucleus, the cells are not observations, for the mirror-image reason.An
obsplink would only be well defined for one AnnData whoseobsis the disjoint union of cells, nuclei and pathogens – and that matrix is useless downstream, becauseXwould be block-structured by construction (a nucleus row has nocell_area) and every tool would see it as ~60% missing data.
What is recorded instead is exact and survives everything obsp does not
(subsetting, concatenation, an h5ad round trip through a tool that
knows nothing about spaCR):
joined export –
obs['count_nucleus']/obs['count_pathogen'](how many children were averaged into this row) andvar['is_aggregated'](which columns are such an average). Without the second one it is genuinely easy to report a per-nucleus statistic that is in fact a per-cell mean of nuclei.per-table export of a child –
obs['cell_id'], the schema’s ownparent_columnfromspacr.schema.OBJECT_TABLE_SCHEMAS. A plain foreign key, which is whatadata.obs.groupby('cell_id')already expects.both –
uns['spacr']['relationships'], stating the parent table, the key it is joined on, whether the child features were aggregated, and (forexport_anndata_set()) the sibling.h5adfile holding the parent.
The labels the table join drops¶
spacr.io._read_and_join_tables() takes six named columns off
png_list – the object id, png_path and the four field keys – and
drops the rest. The rest is every annotation column the Annotate app added
and every score the classifier wrote, which for this export are among the
most valuable columns in the database: without them obs has no label to
train on, group by, or colour a UMAP with. They are re-attached by object
key (_attach_png_labels()) before filtering, so a filter can name
one – “export the cells I annotated as infected” – and they are separated
into uns['spacr']['annotation_columns'] and
uns['spacr']['prediction_columns'] using spacr.agreement’s own
rule, which exists because a model’s class column is indistinguishable by
shape from an annotation pass.
They land in obs, never in X: a label is not a measurement, and
scaling, PCA and clustering it alongside the features is how a classifier’s
own output ends up as a “feature” that predicts it.
Provenance¶
uns['spacr'] carries the spaCR version, the settings hash
(spacr.artifacts.settings_hash()), the run id, the absolute source
database path, the tables read, the filter that was applied and the counts
before and after it. The written file is also registered with
spacr.artifacts under kind ANNDATA_KIND, with the project’s
measurements-db artifact as its input – so spacr.artifacts.is_stale
reports the export as stale the moment Measure is re-run, which is the
whole reason the registry exists.
The optional dependency¶
anndata is an optional extra: pip install "spacr[anndata]". It is
imported inside the functions that need it, never at module scope, so
import spacr.anndata_export works without it and the failure is one
actionable sentence (AnnDataExtraMissing) rather than a traceback
six frames deep.
Attributes¶
Exceptions¶
|
|
Two rows claim the same object key. |
Classes¶
What one export produced, and what it left out. |
Functions¶
|
Return this module's settings, filling in anything absent. |
|
Build an |
|
Where the export lands when |
|
Write a spaCR measurements database to |
|
Write one |
|
The columns of |
|
Register this module's settings through the defaults seam. |
Import and return |
|
|
The |
|
Run the export from a settings dict. The headless entry point. |
Package Contents¶
- exception spacr.anndata_export.AnnDataExtraMissing[source]¶
Bases:
ImportErroranndatais not installed.An
ImportErrorsubclass so a caller that already guards the export withexcept ImportErrorkeeps working, and so the message – not a traceback throughanndata’s own import machinery – is what reaches the user.Initialize self. See help(type(self)) for accurate signature.
- exception spacr.anndata_export.DuplicateObjectKeys[source]¶
Bases:
ValueErrorTwo rows claim the same object key.
AnnData requires unique
obs_names, and spaCR’s writers append: two rows can share all five key columns when a field was measured twice (seetests/test_db_contract.py, which measures exactly that). This is raised rather than repaired, because deduplicating – keeping the first, keeping the last, or averaging – changes the numbers and is the caller’s decision, not the exporter’s.Initialize self. See help(type(self)) for accurate signature.
- class spacr.anndata_export.ExportResult[source]¶
What one export produced, and what it left out.
- Parameters:
path – the
.h5adwritten, or""for an in-memory build.n_obs – objects in the exported matrix.
n_vars – feature columns in
X.n_obs_before_filter – objects read from the database.
obs_columns – the
obscolumn names.obsm_keys – the
obsmkeys written.nan_policy – the policy that was applied.
n_missing – NaN cells in
Xbefore the policy ran.n_infinite – non-finite cells converted to NaN before that count.
dropped_features – feature columns removed by the policy.
dropped_objects – objects removed by the policy.
n_obs_counted – rows of the matrix
n_missingwas counted in – i.e. afterdata_filter/selection/row_limitand before thenan_policyran.0on a record that did not record it, in which casecounted_shapereconstructs it from the drops.n_vars_counted – columns of that same pre-policy matrix.
artifact_id – the
spacr.artifactsid, or""when the file was not registered.warnings – everything the export decided the user must know.
Three object counts live here and they are three different stages, which is the whole reason they are named apart:
n_obs_before_filteris what the database held,counted_shapeis what survived the filter and was handed to the NaN policy, andn_obsis what was written.describe()attributes each loss to the stage that caused it rather than charging all of them to the filter.- describe() str[source]¶
One human paragraph: shape, filtering, missingness, provenance.
Each count is charged to the stage that caused it: the filter line counts only what the filter removed, and the
dropped ...lines only what the NaN policy removed. Charging the policy’s row drops to the filter as well made one loss of four objects read as eight.
- property counted_shape: Tuple[int, int][source]¶
(rows, columns)of the matrixn_missingwas counted in.That matrix is the post-filter, pre-
nan_policyone: the NaN were counted before anything was dropped or imputed. A record built by this module states it outright; one built by hand (or by an older spaCR) leaves the two fields at0, and the shape is reconstructed from the drops instead –drop_objectsis the only policy that removes rows anddrop_featuresthe only one that removes columns, so adding them back recovers what the policy was given.
- property frac_missing: float[source]¶
n_missingover the cells ofcounted_shape.Numerator and denominator describe the same matrix – the one the NaN were counted in, before the policy dropped or imputed anything – so this is a fraction and can never exceed 1.0. Dividing the pre-policy count by the post-policy
n_obs * n_varsis what madedrop_objectsreport 114.3% missing.- Returns:
the fraction, or
0.0for an empty matrix.
- spacr.anndata_export.anndata_export_settings(settings: Mapping[str, Any] | None = None) Dict[str, Any][source]¶
Return this module’s settings, filling in anything absent.
- Parameters:
settings – an existing settings dict to complete.
- Returns:
a new dict; the caller’s is not mutated.
- spacr.anndata_export.build_anndata(db_path: str | os.PathLike, *, tables: Sequence[str] = DEFAULT_TABLES, single_table: str | None = None, data_filter: spacr.selection.DataFilter | None = None, selection: spacr.selection.Selection | None = None, row_limit: int | None = None, timelapse: bool = False, nan_policy: str = NAN_KEEP, missing_layer: bool | None = None, dtype: str = 'float32', exclude: Sequence[str] = (), embeddings: Mapping[str, Any] | None = None, compute_umap: bool = False, umap_settings: Mapping[str, Any] | None = None, condition_map: Mapping[str, str] | None = None, condition_column: str = schema.COLUMN_KEY, attach_labels: bool = True, drop_redundant_identity: bool = True, settings: Mapping[str, Any] | None = None, run_id: str = '', verbose: bool = True)[source]¶
Build an
anndata.AnnDatafrom a spaCR measurements database.The mapping is described in full in the module docstring; in short,
Xis the numeric measurements,obsis everything else about the object,varisspacr.feature_dict’s description of each feature,obsmholds embeddings andunsholds provenance.- Parameters:
db_path – a
measurements.db.tables – tables to join for the default cell-anchored export.
single_table – export exactly this object table instead, one row per object of that type – which is the only way to get a nucleus-level or pathogen-level matrix, since the join averages children onto their parent.
data_filter – a
spacr.selection.DataFilter. Declarative and re-appliable; recorded inuns.selection – a
spacr.selection.Selection– the keys a view pointed at. Applied afterdata_filter.row_limit – hard cap on exported objects, applied last. A blunt instrument on purpose, for “give me something I can open” without inventing a filter that means something it does not.
timelapse – key each frame of an object separately.
nan_policy – one of
NAN_POLICIES; see the module docstring. The default keeps NaN and reports it.missing_layer – write
layers['missing']. Defaults to True for the imputing policies and False otherwise.dtype –
Xdtype.float32by default – the scanpy convention, half the memory, and far more precision than any microscope measurement carries.exclude – feature columns to keep out of
X.embeddings –
{name: array or DataFrame}. Names are normalised to the scanpyX_*convention, so'umap'becomesX_umap. A DataFrame carryingspacr.selection.OBJECT_KEY_COLUMNSis aligned by key, which is what makes an embedding computed on the whole plate usable with a filtered export.compute_umap – compute
X_umaphere, through the samespacr.utils.reduction_and_clustering()callspacr.core.generate_image_umap()makes. Off by default: it imports the segmentation stack and costs minutes on a large table.umap_settings – overrides for that computation.
condition_map –
{column value: label}written toobs['condition'];DEFAULT_CONDITION_MAPis the mappingspacr.utils.map_condition()applies.condition_column – which column
condition_mapreads.attach_labels – bring the annotation and prediction columns back out of
png_list, which the table join drops. On by default – see_attach_png_labels().drop_redundant_identity – drop the join’s suffixed copies of identity columns (
plateID_nucleus,object_label_pathogen) fromobs. On by default; see_redundant_identity_columns()for why two of them are worse than merely duplicated.settings – the run settings, hashed into
unsprovenance.run_id – the run this export belongs to.
verbose – print the summary and the missing-data warning.
- Returns:
(adata, result)– the AnnData and anExportResult.- Raises:
AnnDataExtraMissing – when
anndatais not installed.DuplicateObjectKeys – when two rows claim one object key.
ValueError – on an unknown
nan_policy, an unusable database, or an embedding that cannot be aligned.
- spacr.anndata_export.default_out_path(src: str | os.PathLike, single_table: str = '') str[source]¶
Where the export lands when
anndata_outis empty.<project>/results/<project name>.h5ad– beside the other things a finished run produced, named after the project, because a folder ofexport.h5adfiles is a folder nobody can tell apart.- Parameters:
src – project root, or the database.
single_table – the one object table being exported, if any; it joins the file name, since a nucleus-level export and a cell-level one of the same project are different files.
- Returns:
an absolute
.h5adpath.
- spacr.anndata_export.export_anndata(db_path: str | os.PathLike, out_path: str | os.PathLike, *, compression: str | None = 'gzip', register: bool = True, project: str | os.PathLike | None = None, settings: Mapping[str, Any] | None = None, **kwargs: Any) ExportResult[source]¶
Write a spaCR measurements database to
out_pathas.h5ad.Everything
build_anndata()accepts is accepted here and passed through; this adds the write and the artifact registration. A failed write preserves any existing destination; only a completed file replaces it and is registered as an artifact.- Parameters:
db_path – a
measurements.db.out_path – the
.h5adto write. Parent directories are created.compression –
h5pycompression forXand the layers;Nonefor an uncompressed file.gziptypically halves a feature matrix and costs a few seconds.register – record the file with
spacr.artifacts, with the project’smeasurements-dbas its input so a re-run of Measure marks this export stale.project – the project root the artifact belongs to. Inferred from
db_path(<project>/measurements/measurements.db) when omitted.settings – the run settings, hashed into the artifact and
uns.kwargs – passed to
build_anndata().
- Returns:
the
ExportResult, withpathandartifact_idfilled in.- Raises:
AnnDataExtraMissing – when
anndatais not installed.
- spacr.anndata_export.export_anndata_set(db_path: str | os.PathLike, out_dir: str | os.PathLike, *, object_tables: Sequence[str] = ('cell', 'nucleus', 'pathogen', 'cytoplasm'), prefix: str = '', **kwargs: Any) Dict[str, ExportResult][source]¶
Write one
.h5adper object table, cross-referenced byuns.The honest shape for a multi-compartment experiment: a nucleus is not a cell and averaging it onto one loses the distribution. Each file holds one object type at its own granularity, and each child file records its parent in
uns['spacr']['relationships']– table, key column and the sibling file – so the set can be reassembled (or handed toMuData) without guessing.- Parameters:
db_path – a
measurements.db.out_dir – folder to write into; created if missing.
object_tables – which object tables to export. Tables absent from the database are skipped, not an error.
prefix – prepended to each file name.
kwargs – passed to
export_anndata().
- Returns:
{table: ExportResult}for the tables actually written.
- spacr.anndata_export.feature_columns(frame: pandas.DataFrame, *, exclude: Sequence[str] = (), db_path: str | None = None) List[str][source]¶
The columns of
framethat belong inX, in frame order.A column is a feature when it is numeric, is not provenance or identity by
spacr.schema.is_provenance_column(), is not a human annotation or a model output, is not the cluster label, and was not excluded by the caller.- Parameters:
frame – an object frame read from a measurements database.
exclude – extra column names to keep out of
X.db_path – the source database, used to recognise human annotation columns; omit it and only the model columns are recognised.
- Returns:
feature column names.
- spacr.anndata_export.register_anndata_settings(replace: bool = False) bool[source]¶
Register this module’s settings through the defaults seam.
Uses
spacr.settings.register_defaults()rather than appending tospacr/settings.py, so this module owns its own knobs and adding one is not a merge conflict in a file nobody owns.- Parameters:
replace – re-register over an existing registration.
- Returns:
True if it registered, False if it already was.
- spacr.anndata_export.require_anndata()[source]¶
Import and return
anndata, or raise a message worth reading.- Returns:
the imported
anndatamodule.- Raises:
AnnDataExtraMissing – when the extra is not installed.
- spacr.anndata_export.resolve_db_path(src: str | os.PathLike) str[source]¶
The
measurements.dba settingssrcmeans.srcis a project root everywhere else in spaCR, and this module’s argument is a database, so one of the two has to give. A path that ends in a database file is taken as one, and anything else is read as a project root laid out the way every spaCR writer leaves it.- Parameters:
src – project root, or the database itself.
- Returns:
an absolute path, which is NOT checked for existence – an absent file is
spacr.validate.validate_settings()’s report to make, with the path in it, rather than an exception from here.
- spacr.anndata_export.run_anndata_export(settings: Mapping[str, Any] | None = None) ExportResult | _TablesResult[source]¶
Run the export from a settings dict. The headless entry point.
The
fn(settings)shapespacr-run, the Qt Run button andspacr.validateall dispatch to, wrapped aroundexport_anndata()– which keeps its own explicit keyword signature, because a function whose arguments are a dict is a function nobody can call from a notebook.Every key it reads is one
anndata_export_settings()declares andregister_anndata_settings()gave a type and a tooltip, so the form the GUI draws and the keys honoured here are the same list.anndata_formatchooses what is written:'h5ad'the AnnData file,'parquet'the tidy Parquet tables of_export_tables()inanndata_tidy_dir,'r'those tables with the R loader script (and.rdsdata frames when pyreadr is installed), and'all'every one of them. The tables keep float64 features whateveranndata_dtypesays, which sets the AnnData matrix only.- Parameters:
settings – the run settings.
srcis the project root (or the database); everything else falls back toanndata_export_settings().- Returns:
the
ExportResultwhen an.h5adwas written, and otherwise the tables result;describe()on either is what the console prints.- Raises:
ValueError – when
srcis empty – there is nothing to export and no path to name in the message otherwise – oranndata_formatis not a known format.AnnDataExtraMissing – when
anndatais not installed and the format includes.h5ad.ImportError – when pyarrow is not installed and the format includes the tables.
- spacr.anndata_export.ANNDATA_MISSING_MESSAGE = Multiline-String[source]¶
Show Value
"""Exporting to AnnData (.h5ad) needs the optional `anndata` extra, which is not installed in this environment (missing module: {module}). Install it with: python -m pip install "spacr[anndata]" scanpy is a separate install and is NOT required to write the file: python -m pip install scanpy"""