spacr.feature_dict

Workflow inputs and outputs

Feature Dictionary

Look up feature definitions and units before choosing measurement columns.

Open: the application’s Help/tools menus.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Feature definitions — The installed feature dictionary; definitions describe existing measurements and do not compute them.

Outputs

  • Feature definitions — The installed feature dictionary; definitions describe existing measurements and do not compute them.

API reference.

Module tutorial.

Human-readable data dictionary for the columns of a spaCR measurements.db.

A finished spaCR run writes hundreds of columns per object table with names like cell_channel_1_percentile_75, nucleus_zernike_12 or pathogen_channel_0_channel_2_M1_correlation_85. This module turns those names back into prose: what the number means, what unit it is in, which object and which channel(s) it came from, and which line of code produced it.

Everything in KNOWN_PROPERTIES was derived by reading the emitters in spacr.measure and spacr.utils — not from guessing at the names. Where the code does something surprising (a doubled name prefix, a feature that is always NaN, a radial bin that covers the background) the entry says so in notes rather than describing the intent.

Geometric units are not fixed any more. A 2-D run measures in pixels, but a 3-D run measures a volume, and with voxel_size_z_um / voxel_size_xy_um set it measures in micrometres — under the same column names, because spacr.measure deliberately does not rename <object>_area (renaming would break every downstream selector). Which one a row is in is recorded on the row itself, in measurement_units. So the unit of a geometric column is a ConditionalUnit: describe_database() reads measurement_units out of the database it is documenting and resolves it, and a caller who has no database says so and gets the condition spelled out instead of a confident guess. See MEASUREMENT_UNITS.

Typical use:

from spacr.feature_dict import describe_database, export_dictionary

df = describe_database("/data/exp1/measurements/measurements.db")
export_dictionary("/data/exp1/measurements/measurements.db",
                  "/data/exp1/feature_dictionary.md", fmt="md")

The module is deliberately dependency-light: standard library plus pandas. It imports no torch, no cellpose and no scikit-image, so it can be used to explain a database on a machine that cannot run the pipeline.

Classes

Concept

One searchable idea, and the curated keys that answer to it.

ConditionalUnit

A unit that depends on how the row was measured.

Coverage

How much of a set of column names the dictionary can explain.

FeatureDoc

One feature, as opposed to one column.

FeatureEntry

A single database column, decomposed and explained.

FeatureScope

Where a curated feature exists, and what has to be on for it to.

PropertyInfo

One curated definition, shared by every column that instantiates it.

SearchHit

One search result: a feature, how well it matched, and why.

Functions

concept_of(→ str | None)

Resolve a user's word to a concept name, or None.

concepts_for(→ tuple[str, ...])

Which CONCEPTS a curated key answers to.

coverage(→ Coverage)

Measure what share of columns this dictionary explains.

describe_columns(→ list[FeatureEntry])

Describe every column name given, in order and without dropping any.

describe_database(→ Any)

Describe every column of a spaCR measurements database.

doc_for(→ FeatureDoc | None)

The FeatureDoc for a curated key, or None.

export_dictionary(→ pathlib.Path)

Write a data dictionary for db_path to out_path.

feature_docs(→ tuple[FeatureDoc, ...])

Every documented feature, metadata column and link column.

parse_column(→ FeatureEntry)

Decompose one measurements.db column name into a described feature.

scope_for(→ FeatureScope | None)

The FeatureScope of a curated key, or None if untabulated.

search_features(→ list[SearchHit])

Find features by name, by substring, or by concept.

Module Contents

class spacr.feature_dict.Concept[source]

One searchable idea, and the curated keys that answer to it.

Parameters:
  • name – canonical concept identifier accepted by the search filter.

  • gloss – short human explanation of the scientific idea.

  • synonyms – alternative query phrases resolved to name.

  • keys – curated feature keys associated with the concept, ordered from most to least characteristic for search ranking.

class spacr.feature_dict.ConditionalUnit[source]

A unit that depends on how the row was measured.

This module used to state a single unit for every geometric column and say in the string itself that “spaCR never applies a physical pixel size”. That was true until spacr.measure learned to measure a (Z, Y, X) mask in 3-D: a run with voxel_size_z_um and voxel_size_xy_um set reports micrometres, and one with anisotropy alone reports xy-pixel units. The column name is identical in all three cases — measure.py records the unit on the row instead of renaming the column — which is exactly why the dictionary has to read it from the data rather than assert it.

Parameters:
  • px – unit text for rows stamped measurement_units='px' (the 2-D pixel mode), or None when the column is not written in that mode.

  • px_xy – unit text for rows stamped measurement_units='px_xy' (3-D geometry scaled by anisotropy in xy-pixel units), or None when the column is not written in that mode.

  • um – unit text for rows stamped measurement_units='um' (3-D geometry measured with physical voxel sizes), or None when the column is not written in that mode.

A field is None when the column is not written at all in that mode — the _z/_y/_x centroid axes exist only in 3-D, for instance.

by_units() → dict[str, str | None][source]

{measurement_units value: unit} for all three modes.

conditional_text() → str[source]

Every possibility, each with the condition it holds under.

Used when the measurement_units of the data being described is not known. Long, but a wrong unit is worse than a long one.

resolve(measurement_units: str | None = None) → str[source]

Return the concrete unit for a row stamped measurement_units.

Parameters:

measurement_units – one of MEASUREMENT_UNITS. None, or any value this module does not recognise, returns conditional_text() — the condition, not a guess.

class spacr.feature_dict.Coverage[source]

How much of a set of column names the dictionary can explain.

Parameters:
  • total – number of input column names examined, including repeats.

  • explained – number of inputs that resolved to a known feature.

  • unknown – unresolved names in input order, with duplicates removed.

property fraction: float[source]

Explained share, in [0, 1]. An empty input is 1.0.

class spacr.feature_dict.FeatureDoc[source]

One feature, as opposed to one column.

cell_channel_0_percentile_75 and nucleus_channel_2_percentile_5 are two columns and one feature. The panel lists features and resolves columns onto them, because a user reading a results table wants “what is a percentile here” answered once, not four hundred times.

Parameters:
  • key – the curated key — the feature’s identity.

  • title – a human title for the key.

  • kind – "feature", "metadata", or "link", identifying how the column participates in a measurement table.

  • family – feature family used for browsing and search filtering.

  • concepts – searchable scientific concepts associated with the key.

  • description – scientific meaning of the value, or None when the curated dictionary has no definition.

  • unit – the unit, with a conditional unit spelled out in full.

  • computed_by – function or algorithm responsible for producing the value.

  • module – spaCR module that owns that computation or metadata field.

  • object_types – object types the feature is written for. Empty means “not per-object” (metadata) or “never written” (dead code).

  • channel_scope – CHANNEL_NONE / CHANNEL_SINGLE / CHANNEL_PAIR.

  • written_when – condition under which the column is emitted, or None when it is unconditional.

  • notes – caveats needed to interpret the value, or None when no additional warning applies.

  • examples – concrete column names that resolve back to this key.

to_dict() → dict[str, Any][source]

Return the doc as a plain dict.

class spacr.feature_dict.FeatureEntry[source]

A single database column, decomposed and explained.

Parameters:
  • column – Column name exactly as it appears in the database.

  • object_type – Object-type prefix parsed from the column, or None when the column has no object prefix.

  • channel – Zero-based first channel index parsed from the column, or None when no channel enters it.

  • family – Feature-family key from FEATURE_FAMILIES.

  • description – Prose meaning of the value, or None when the meaning is undetermined.

  • unit – Unit text for the value, or None for identifiers and unrecognised columns.

  • computed_by – Provenance string naming the producer, or "unknown" for an unrecognised column.

  • notes – Optional caveats about the column, or None when none apply.

  • channel_2 – Zero-based second channel index for two-channel colocalisation columns, otherwise None.

  • object_type_2 – Second object type carried by a pandas merge suffix such as ..._nucleus, otherwise None.

  • measurement_units – Measurement-unit stamp supplied while resolving a conditional unit; None for fixed-unit columns or when no stamp was supplied. An unrecognised supplied stamp is retained while unit states the unresolved conditions.

  • key – Curated KNOWN_PROPERTIES or META_COLUMNS key through which the column resolved, or None for an unrecognised column.

  • object_types – Object types for which the feature is written; empty for features that are not per-object.

  • channel_scope – CHANNEL_NONE, CHANNEL_SINGLE, or CHANNEL_PAIR, describing how channels enter the feature.

  • module – spaCR module that produces the value, or "unknown" when it cannot be identified.

  • written_when – Condition under which the column exists, or None when no condition is recorded.

  • concepts – Names from CONCEPTS associated with the feature.

to_dict() → dict[str, Any][source]

Return the entry as a plain dict, ready for a DataFrame row or JSON.

class spacr.feature_dict.FeatureScope[source]

Where a curated feature exists, and what has to be on for it to.

Parameters:
  • objects – object types whose per-object tables receive the feature. An empty tuple denotes a non-per-object or never-written feature.

  • channels – channel arity, one of CHANNEL_NONE, CHANNEL_SINGLE, or CHANNEL_PAIR, indicating whether the feature depends on zero, one, or two intensity channels.

  • module – dotted spaCR module name that computes or emits the feature.

  • written_when – human-readable condition under which the column is emitted, or None when every run writes it.

class spacr.feature_dict.PropertyInfo[source]

One curated definition, shared by every column that instantiates it.

Parameters:
  • family – feature-family key from FEATURE_FAMILIES used to group and filter the measurement.

  • description – prose explaining what the value means, or None only when the meaning could not be determined from the spaCR source.

  • unit – physical or derived unit, None for identifiers, or a ConditionalUnit resolved from the row’s measurement units.

  • computed_by – function or library call that produces the value; curated properties require non-empty provenance.

  • notes – optional caveats, known defects, or comparability warnings.

class spacr.feature_dict.SearchHit[source]

One search result: a feature, how well it matched, and why.

Parameters:
  • doc – feature definition selected by the search.

  • score – accumulated relevance score used for descending result order.

  • reason – semicolon-separated explanation of the rules that matched.

spacr.feature_dict.concept_of(word: str) → str | None[source]

Resolve a user’s word to a concept name, or None.

Matches a concept name or any of its synonyms, case-insensitively.

Parameters:

word – the user’s search word or phrase; converted to str, stripped of surrounding whitespace and lower-cased before lookup.

spacr.feature_dict.concepts_for(key: str) → tuple[str, ...][source]

Which CONCEPTS a curated key answers to.

Parameters:

key – curated feature key, matched exactly; a key in no concept gives an empty tuple.

spacr.feature_dict.coverage(columns: Iterable[str], measurement_units: str | None = None) → Coverage[source]

Measure what share of columns this dictionary explains.

The number a user should be able to check for themselves, and the number the test suite pins: a lookup panel that answers “no entry” for a third of a real table is not a lookup panel.

Parameters:
  • columns – Iterable of column names.

  • measurement_units – passed through to parse_column().

Returns:

A Coverage, whose unknown lists the names that were not explained — in input order, without duplicates.

spacr.feature_dict.describe_columns(columns: Iterable[str], measurement_units: str | None = None) → list[FeatureEntry][source]

Describe every column name given, in order and without dropping any.

Parameters:
  • columns – Iterable of column names.

  • measurement_units – The measurement_units value these columns were measured under; see parse_column().

Returns:

One FeatureEntry per input name, same order.

spacr.feature_dict.describe_database(db_path: str | pathlib.Path, table: str | None = None, measurement_units: str | None = None) → Any[source]

Describe every column of a spaCR measurements database.

Every column of every table is returned — a column that this dictionary cannot explain appears with family='unknown' and a null description rather than being omitted.

The units are read from the database, not assumed: each table’s own measurement_units column decides whether <object>_area is reported as a px^2 area, a cubic-xy-pixel volume or a um^3 volume. A table that cannot be pinned to one value gets units that state the condition instead.

Parameters:
  • db_path – Path to a SQLite database, typically measurements.db.

  • table – Restrict to a single table. None (default) covers all user tables.

  • measurement_units – Force the unit basis (one of MEASUREMENT_UNITS) instead of reading it from each table. For describing a frame that has been detached from its database.

Returns:

DataFrame with one row per (table, column) and the columns table, column, object_type, object_type_2, channel, channel_2, family, description, unit, measurement_units, computed_by, notes, key, object_types, channel_scope, module, written_when, concepts. object_types and concepts are comma-joined strings, and both channel columns are nullable Int64.

Raises:
spacr.feature_dict.doc_for(key: str) → FeatureDoc | None[source]

The FeatureDoc for a curated key, or None.

Parameters:

key – curated feature key, compared exactly against each FeatureDoc.key in feature_docs().

spacr.feature_dict.export_dictionary(db_path: str | pathlib.Path, out_path: str | pathlib.Path, fmt: str = 'csv') → pathlib.Path[source]

Write a data dictionary for db_path to out_path.

Parameters:
  • db_path – Path to the SQLite database to describe.

  • out_path – Destination file. Parent directories are created.

  • fmt – "csv", "md" or "json".

Returns:

The path written, as a Path.

Raises:
spacr.feature_dict.feature_docs() → tuple[FeatureDoc, ...][source]

Every documented feature, metadata column and link column.

Built once and cached. Order is KNOWN_PROPERTIES order, then META_COLUMNS, then the link columns.

spacr.feature_dict.parse_column(name: str, measurement_units: str | None = None) → FeatureEntry[source]

Decompose one measurements.db column name into a described feature.

The grammar this implements, derived from the f-strings in spacr.measure:

<object>_<stat>                                  morphology / moments
<object>_channel_<i>_<stat>                      single-channel intensity/texture
<object>_channel_<i>_channel_<j>_<stat>          two-channel colocalisation
<object>_channel_<i>_periphery_<stat>            inner boundary rim
<object>_channel_<i>_outside_<stat>              5 px surrounding ring
<object>_rad_dist_channel_<c>_bin_<b>            radial intensity profile
organelle_summary_<stat>                         per-parent organelle summary
<object>_volume_voxels / _volume_um3             3-D volumes, named by unit
<object>_channel_<i>_centroid_weighted_<z|y|x>   3-D centroid, named by axis
measurement_ndim / measurement_units / n_z /     per-row provenance stamp
voxel_size_z_um / voxel_size_xy_um

An unrecognised name never raises and is never dropped: it comes back as a FeatureEntry with family="unknown" and description=None.

Parameters:
  • name – Column name exactly as stored in the database.

  • measurement_units – The measurement_units value of the rows being described — one of MEASUREMENT_UNITS. Geometric columns have no single unit any more (a 3-D run measures a volume, in micrometres when it was given a voxel size), so pass this when you know it and the returned unit is concrete. Left None, unit states every possibility with the condition attached rather than guessing one; describe_database() fills it in from the database itself.

Returns:

A FeatureEntry describing the column.

spacr.feature_dict.scope_for(key: str) → FeatureScope | None[source]

The FeatureScope of a curated key, or None if untabulated.

Parameters:

key – curated feature key, e.g. "area"; looked up exactly (case-sensitive) in FEATURE_SCOPE.

spacr.feature_dict.search_features(query: str, *, object_type: str | None = None, concept: str | None = None, family: str | None = None, limit: int | None = None) → list[SearchHit][source]

Find features by name, by substring, or by concept.

Four ways in, because a user who does not know spaCR’s naming scheme has to be able to find things anyway:

  • a column name — cell_channel_1_percentile_75 resolves through parse_column() and its feature is the first hit, even though no such literal string appears anywhere in this module;

  • a curated key or part of one — percentile, zernike;

  • a concept — intensity, texture, shape, distance, and every synonym in CONCEPTS (how big, roundness, colocalisation, blurry…);

  • free text — matched against the definitions and the notes.

Parameters:
  • query – the search text. Empty returns everything, filtered.

  • object_type – restrict to features written for this object type. Filters on FeatureDoc.object_types, so asking for cell correctly excludes periphery_mean.

  • concept – restrict to one CONCEPTS name (or synonym).

  • family – restrict to one FEATURE_FAMILIES name.

  • limit – keep only the best limit hits.

Returns:

SearchHit list, best first; ties keep dictionary order.