spacr.schema

Canonical database keys, filename identities, and table schemas.

spaCR measurement tables share the plateID, rowID, columnID, and fieldID columns, with timeID added for time-lapse data. This module defines those names, their composed prc/prcf/prcfo forms, and the parsers used by database, image, and GUI code.

Numeric tokens may include common vendor prefixes such as s3 or T0003. Unparseable but non-empty tokens remain distinct in permissive mode and raise KeyParseError in strict mode; missing identity components always raise. Legacy helpers are retained for reading and testing older data.

The module has no third-party import-time dependencies. Data-frame helpers import pandas only when called.

Exceptions

KeyParseError

A key token was absent, or was rejected under strict=True.

ModelFeatureSchemaError

A declared model feature cannot be represented as numeric input.

ObjectTableSchemaError

An object-table frame violates its declared column or row contract.

SchemaError

Base for every failure to build or read a spaCR key.

WellParseError

A well identifier could not be turned into a row and a column.

Classes

ColumnCollision

What happened when several columns claimed one metadata key.

FieldID

One imaging field's identity — the key of every measurement row.

ObjectID

One segmented object's identity — the prcfo a merged row is keyed on.

ObjectTableSchema

Declarative contract for one object measurement table.

Functions

add_identity_columns(df[, source, timelapse, objects, ...])

Parse a name column into the canonical key columns.

add_screen_column(df[, screen, overwrite])

Return df carrying a filled-in SCREEN_KEY column.

canonical_column_name(→ str)

Return the canonical spelling of a metadata or feature column name.

canonical_plate_id(→ str)

The plate id in the form the rest of spaCR keys on.

canonical_rename_plan(columns[, requested])

{old: canonical} for the columns that can safely be renamed.

canonicalise_columns(df)

Return df with every legacy column name renamed canonically.

canonicalise_frame(frame, *[, report, warn, ...])

The whole vocabulary applied to one frame: the reader's normaliser.

coerce_model_feature_types(frame, *[, extra_features, ...])

Repair the two representations of a numeric measurement pandas fumbles.

column_id(→ str)

Return the canonical 'c<N>' column id.

column_index(→ Optional[int])

'c12' → 12; an unparseable id → None.

comparable_key_value(→ str)

One metadata value, reduced to the form two spellings are compared in.

comparable_key_values(→ Tuple[str, ...])

comparable_key_value() over a column.

compose_prc(→ str)

Return the prc well key: 'plate1_r1_c1'.

compose_prc_column(df[, columns])

Compose escaped plate-row-column identifiers for a frame.

compose_prcf(→ str)

Return the prcf field key.

compose_prcfo(→ str)

Return the prcfo object key: 'plate1_r1_c1_f2_o7'.

correct_metadata_column_names(df)

Rename legacy metadata columns to the canonical spaCR names.

escape_field_stem_plate(→ str)

Escape the plate component of a merged-stack field stem.

escape_filename_component(→ str)

Escape one free-text component for a separator-delimited filename.

field_id(→ str)

Return the canonical 'f<N>' field id.

field_index(→ Optional[int])

'f2' → 2; 'fxy' → None.

fold_column_name(→ str)

The form a column name is looked up by: lower case, no punctuation.

is_object_type(→ bool)

Whether object_type can be written into an object key.

is_positional_pair(→ bool)

True when (rowID, columnID) came from the positional passthrough.

is_positional_well(→ bool)

True when well is a bare number rather than <letters><digits>.

is_provenance_column(→ bool)

Return whether name is identity, annotation, or run provenance.

is_row_column_pair(→ bool)

True when (row, column) is recognisably a well's row and column.

is_within_plate_format(→ bool)

True when (row, column) lies inside an n_wells plate.

legacy_map_wells(→ Tuple[str, ...])

utils._map_wells reproduced bit for bit, 'error' tuple and all.

legacy_safe_int_convert(→ Any)

utils._safe_int_convert exactly, including the 0 default.

legacy_well_ids(→ Tuple[str, str])

utils._map_wells' well handling exactly, raising where it raises.

letters_from_row_index(→ str)

Inverse of row_index_from_letters(). 27 → 'AA'.

model_feature_columns([exclude])

Select numeric model inputs by schema role, never by dtype alone.

model_feature_frame(frame, **kwargs)

Return frame restricted to model_feature_columns().

normalise_plate_columns(frame)

Collapse a doubled p prefix in every column that carries a plate.

object_id(→ str)

Return the canonical object id used in prcfo.

object_index(→ Optional[int])

'o41' and 'nucleus41' → 41; an unparseable id → None.

object_table_schema(→ ObjectTableSchema)

Return the canonical schema for table.

object_type_prefix(→ str)

Return the canonical prefix an object of object_type is keyed with.

object_type_summary(→ str)

Describe object types once, collapsing organelle slots into a regex.

parse_field_stem(→ FieldID)

Parse a merged-stack file name into a FieldID.

parse_int_token(→ Optional[int])

Return the integer token denotes, or None — never 0.

parse_object_stem(→ ObjectID)

Parse a crop-PNG file name into an ObjectID.

parse_prcf(→ FieldID)

Parse a prcf string back into a FieldID.

parse_prcfo(→ ObjectID)

Parse a prcfo string back into an ObjectID.

parse_well(→ Tuple[str, str])

Return (rowID, columnID) for a well identifier.

plate_format_for(→ Optional[int])

Return the smallest standard plate format containing (row, column).

resolve_metadata_collisions(frame, *[, report, warn])

Collapse every group of columns that mean the same metadata key.

row_id(→ str)

Return the canonical 'r<N>' row id.

row_index(→ Optional[int])

'r3' → 3; 'C' → 3; an unparseable id → None.

row_index_from_letters(→ Optional[int])

'A' → 1, 'Z' → 26, 'AA' → 27, 'AF' → 32.

screen_id(→ str)

Return the canonical screen id, defaulting an absent one.

split_object_id(→ Tuple[Optional[str], str])

Split an object id into (object type, label).

strip_prefix(→ str)

Remove one leading prefix from value if it is there.

table_key_columns(→ Tuple[str, ...])

Return the columns that identify a row of table.

time_id(→ str)

Return the canonical 't<N>' timepoint id.

time_index(→ Optional[int])

't7' → 7; an unparseable id → None.

unescape_filename_component(→ str)

Invert escape_filename_component().

validate_object_table_frame(frame, table, *[, ...])

Validate an object-table frame against its canonical contract.

well_id(→ str)

Return the canonical well name: ('r3', 'c7') → 'C07'.

Module Contents

exception spacr.schema.KeyParseError[source]

Bases: SchemaError

A key token was absent, or was rejected under strict=True.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.schema.ModelFeatureSchemaError[source]

Bases: SchemaError

A declared model feature cannot be represented as numeric input.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.schema.ObjectTableSchemaError[source]

Bases: SchemaError

An object-table frame violates its declared column or row contract.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.schema.SchemaError[source]

Bases: ValueError

Base for every failure to build or read a spaCR key.

A subclass of ValueError so that call sites which already guard a parse with except ValueError keep working.

Initialize self. See help(type(self)) for accurate signature.

exception spacr.schema.WellParseError[source]

Bases: SchemaError

A well identifier could not be turned into a row and a column.

Initialize self. See help(type(self)) for accurate signature.

class spacr.schema.ColumnCollision[source]

What happened when several columns claimed one metadata key.

Parameters:
  • canonical – canonical metadata key claimed by all source columns.

  • sources – source columns in their original frame order.

  • chosen – source column retained under the canonical name.

  • dropped – redundant source columns removed from the frame.

  • disagreeing_rows – number of rows whose source values disagreed.

  • rows – total number of rows compared.

property agreed: bool[source]

Whether every row said the same thing.

property message: str[source]

The sentence a user is shown – printed, or warned.

class spacr.schema.FieldID[source]

One imaging field’s identity — the key of every measurement row.

Parameters:
  • plateID – plate identifier, escaped when composed into a field key.

  • rowID – stored row component, normally r<N> but potentially a legacy positional-well passthrough.

  • columnID – stored column component, normally c<N> but potentially a legacy positional-well passthrough.

  • fieldID – canonical f<label> imaging-field component.

  • timeID – canonical t<label> timepoint component inserted after the field, or None outside a timelapse.

classmethod build(plate: Any, well: Any = None, field: Any = None, time: Any = None, *, row: Any = None, column: Any = None, strict: bool = False) → FieldID[source]

Construct from a well string, or from a row and column.

Parameters:
  • plate – plate id.

  • well – well identifier, e.g. 'A01'. Mutually exclusive with row/column.

  • field – field token.

  • time – timepoint token, or None.

  • row – row index or 'r<N>', when there is no well string.

  • column – column index or 'c<N>'.

  • strict – reject unparseable field/time tokens and odd wells.

Returns:

the FieldID.

Raises:

KeyParseError – when neither a well nor a row/column pair is given.

to_dict(*, include_prcf: bool = False) → Dict[str, str][source]

Return the identity as the dict the tables carry.

Parameters:

include_prcf – also emit prc and prcf.

Returns:

{plateID, rowID, columnID, fieldID[, timeID][, prc, prcf]}.

with_object(obj: Any, object_type: Any = None) → ObjectID[source]

Return the ObjectID for one object in this field.

Parameters:
  • obj – object label, bare or already prefixed.

  • object_type – which object table it came from, or None for “not stated”. A nucleus and a pathogen with the same label in the same field are two objects, and without this they are one key.

property positional: bool[source]

True when this identity came from the positional passthrough.

property prc: str[source]

The prc well key.

property prcf: str[source]

The prcf field key, with the timepoint when there is one.

property well: str | None[source]

The well name ('A01'), or None for a positional well.

class spacr.schema.ObjectID[source]

One segmented object’s identity — the prcfo a merged row is keyed on.

Parameters:
  • plateID – plate identifier carried by the containing field and escaped when composed into a field or object key.

  • rowID – stored row component, normally r<N> but potentially a legacy positional-well passthrough.

  • columnID – stored column component, normally c<N> but potentially a legacy positional-well passthrough.

  • fieldID – canonical f<label> imaging-field component.

  • objectID – o<label> when the object’s type is not stated, or <type><label> when it is. The type lives inside this field rather than beside it so that two ObjectID values compare equal exactly when they name the same object — a separate objectType field would let ('o7', 'nucleus') and ('nucleus7', None) describe one object and compare unequal.

  • timeID – canonical t<label> timepoint component inserted between field and object, or None outside a timelapse.

to_dict(*, include_prcf: bool = False) → Dict[str, str][source]

Return the identity as a dict, including prcfo.

object_type appears only when the object has one. An untyped id emits the same dict it always did, so a reader that never learned about types sees no new column on data that has no type to report.

property field: FieldID[source]

The field this object sits in.

property objectLabel: str[source]

'7'.

Type:

The label with the type (or 'o') taken back off

property objectType: str | None[source]

Which object table this object came from, or None.

None is “not stated”, which is what every key written before object types existed carries. It is emphatically not “a cell”.

property prcf: str[source]

The prcf of the containing field.

property prcfo: str[source]

The prcfo object key.

class spacr.schema.ObjectTableSchema[source]

Declarative contract for one object measurement table.

Feature columns are open-ended because channel counts and enabled measurements vary per run, but they are not untyped: a feature written by a table starts with <object_type>_ and is numeric. Unknown annotation or provenance columns remain permitted so older databases and user-added labels are not destroyed by validation.

Parameters:
  • table – SQLite table name.

  • object_type – required feature-column prefix.

  • parent_column – optional link to a parent cell.

feature_column(name: Any) → bool[source]

Return whether name belongs to this table’s feature namespace.

Parameters:

name – candidate column name.

row_key_columns(*, timelapse: bool = False) → Tuple[str, ...][source]

Return the columns that must be unique within one write batch.

validate(frame, *, timelapse: bool | None = None)[source]

Validate and return a canonical-column copy of frame.

Parameters:

frame – pandas frame to validate against this table contract.

property identifier_columns: Tuple[str, ...][source]

Prefixed link/label columns emitted by morphology measurement.

property optional_columns: Tuple[str, ...][source]

Stable optional metadata, including the shared provenance stamp.

property required_columns: Tuple[str, ...][source]

Columns every row set of this table must expose.

spacr.schema.add_identity_columns(df, source: str = 'file_name', *, timelapse: bool = False, objects: bool = False, strict: bool = False, include_prcf: bool = True)[source]

Parse a name column into the canonical key columns.

The vectorised form of parse_field_stem() / parse_object_stem(), for the writers that today do df[[...]] = df[col].apply(lambda x: pd.Series(_map_wells(x))) — a line that positionally unpacks a tuple whose length changes with timelapse, so a mismatched flag misaligns every column.

Parameters:
  • df – pandas.DataFrame with a column of file names.

  • source – name of that column. Default 'file_name'.

  • timelapse – names carry a timepoint.

  • objects – names are crop PNGs, so also emit prcfo.

  • strict – reject unparseable tokens.

  • include_prcf – also emit prc and prcf.

Returns:

a new frame with the key columns added.

Raises:

KeyParseError – when source is not a column of df.

spacr.schema.add_screen_column(df, screen: Any = None, *, overwrite: bool = False)[source]

Return df carrying a filled-in SCREEN_KEY column.

The three cases, and why each behaves as it does:

  • The frame has no screen column. It gains one, holding screen or DEFAULT_SCREEN. That is a single-screen project, which is every project written before and it must keep working.

  • The frame has one, and screen is None. Its labels are kept. Relabelling a frame that already knows which experiment it came from would move rows between screens with nothing on screen to say so.

  • The frame has one and screen was given. The caller is looking at the files and has said which screen this is, so it wins — but only because they said so. overwrite=False (the default) still fills blank values only; pass overwrite=True to restamp every row.

Blanks are never left blank: None, NaN and '' become DEFAULT_SCREEN, because an empty screen id is not an identity and every row carrying one would group with every other.

Parameters:
  • df – pandas.DataFrame.

  • screen – the screen label for rows that do not have one.

  • overwrite – replace existing labels instead of filling blanks.

Returns:

a new frame; df is not modified.

spacr.schema.canonical_column_name(name: Any) → str[source]

Return the canonical spelling of a metadata or feature column name.

Metadata lookup folds case and punctuation (fold_column_name()), so 'RowID', 'rowid', 'Row Name' and 'row_name' all resolve to 'rowID'. Feature rewrites (LEGACY_COLUMN_PATTERNS) are case-sensitive. A name matching neither is returned unchanged.

Note

What is deliberately not in the vocabulary

Aliases are whole words. col is here because spaCR itself wrote it; a one-letter c is not, and never will be, because it would capture a measurement column called c and rename a real variable into a plate key. A name that is not normalised is visible the moment a user looks at the picker; a measurement silently renamed to columnID is not, and it corrupts the join it lands in.

This is the only implementation. spacr.database_schema and spacr.utils re-export this function rather than defining their own.

Note

Migration note (one function, was two)

database_schema.canonical_column_name used to be a second, narrower implementation: 11 aliases, matched case-sensitively. Which one a caller got depended on whether it had imported spacr.schema or spacr.utils, and the two disagreed. Adopting this one widens what a database migration renames:

  • plate_id, row_id, column_id, col_name, field_id, rowid, time, timepoint, channel_name, chan_id and slice_id are now renamed to their canonical spellings;

  • any case variant is now renamed too, so a database column spelled Row or RowID is repaired instead of being left for a reader to trip over. SQLite compares identifiers case-insensitively, so RowID -> rowID is a pure respelling of one column, but pandas reports whatever spelling is stored — which is how a frame ends up with no rowID column on a database that has one.

The non-destructive rule is unchanged: a table already carrying the canonical name keeps both columns.

Parameters:

name – column name as it appears in a table or CSV.

Returns:

the canonical name, or name unchanged.

Example

>>> canonical_column_name('column_name')
'columnID'
>>> canonical_column_name('cell_periphery_25_percentile')
'cell_periphery_percentile_25'
>>> canonical_column_name('cell_area')
'cell_area'
spacr.schema.canonical_plate_id(plate: Any) → str[source]

The plate id in the form the rest of spaCR keys on.

A legacy score CSV stamps its plate pplate1 while the sequencing counts stamp it plate1. The two then do not join, the merge returns zero rows, and the run dies several steps later inside a plot with KeyError: 0 – nowhere near the mismatch and with nothing on screen naming a plate.

This lives here, and not in spacr.multi_database where it was written, because there must be one normaliser: utils.correct_metadata had the rule for frames, multi_database had it for scalars and for database reads, and the pair had to be pinned against each other by test precisely because they were two. Both now call this. spacr.multi_database re-exports the name, so no caller moved.

Parameters:

plate – a plate id, from anywhere.

Returns:

the id with a doubled p prefix collapsed; everything else unchanged.

spacr.schema.canonical_rename_plan(columns, requested=None)[source]

{old: canonical} for the columns that can safely be renamed.

The one definition of the “target already exists, keep both” rule for frames, shared by canonicalise_columns() and utils.canonicalize_measurement_columns so the two cannot drift apart again — they have already disagreed once about which spellings they fix.

The test folds case, because SQLite compares identifiers case-insensitively and these frames are written with to_sql. A frame holding row and rowid already has the canonical column: renaming row to rowID produced ['rowID', 'rowid'], which looks fine in pandas and makes to_sql raise duplicate column name: rowid. This is the same rule, and the same reasoning, as the database-level rename in spacr.database_schema — see the comment there about a plate whose database could not be opened at all.

A column is excluded from its own comparison, so a pure respelling (rowid -> rowID, one column, one identifier as far as SQLite is concerned) is still made rather than being read as a collision with itself.

Parameters:
  • columns – the frame’s column names, in order.

  • requested – optional explicit {source: canonical_target} choices supplied by the metadata resolver. They pass through the same case-folded collision guard as built-in aliases.

Returns:

dict mapping each renameable name to its canonical form; empty when there is nothing to do.

spacr.schema.canonicalise_columns(df)[source]

Return df with every legacy column name renamed canonically.

Metadata aliases (LEGACY_COLUMN_NAMES) and the legacy feature spellings (LEGACY_COLUMN_PATTERNS) — this used to fix only the former while utils.canonicalize_measurement_columns, the other frame-level canonicaliser, fixed both. Both now call one canonical_column_name(), so a frame gets the same columns whichever of the two a caller reached for.

A rename is skipped when the canonical name is already present, which is the same rule utils.rename_columns_in_db follows: a frame carrying both spellings keeps both untouched rather than having one silently overwrite the other. Dropping data to tidy a name is never the right trade — a human can decide which column is authoritative, and until then both stay reachable.

“Already present” is decided case-insensitively by canonical_rename_plan(), because these frames get written with to_sql and SQLite compares identifiers case-insensitively.

Parameters:

df – pandas.DataFrame whose columns may use legacy names.

Returns:

a new frame with canonical column names.

spacr.schema.canonicalise_frame(frame, *, report=None, warn=None, repair_plate_ids: bool = True)[source]

The whole vocabulary applied to one frame: the reader’s normaliser.

Three steps, in this order and for this reason:

  1. resolve_metadata_collisions() – collapse duplicate opinions before renaming, because renaming first is what produces two columns called rowID and a to_sql that refuses the table.

  2. canonical_rename_plan() – the surviving legacy spellings, plus the legacy feature spellings, renamed.

  3. normalise_plate_columns() – the pplate1 value repair, on every column that carries a plate.

Every collision is recorded on frame.attrs['column_collisions'] as well as reported, so a GUI can show what a read decided without having intercepted the callbacks.

Parameters:
Returns:

a new frame.

spacr.schema.coerce_model_feature_types(frame, *, extra_features=(), exclude=(), allow_unknown: bool = False)[source]

Repair the two representations of a numeric measurement pandas fumbles.

A column with no values at all comes back as object. This is the common one and it is not a data problem: pandas.read_sql builds the frame from the rows it gets, so a column that is NULL in every row arrives as a column of None and pandas types that object – even though SQLite declares it REAL. spaCR writes NULL for an honest NaN, and whole measurements are legitimately NaN for a whole database: skew_intensity/kurtosis_intensity are NaN for every uniform object, and mode_intensity was NaN for every object in every database written before the SciPy shim in measure._extended_regionprops_table. Such a column is converted to float64 NaN, which is what it always meant. Nothing is lost, so nothing is warned about – and the caller’s own all-NaN filter then drops it and says so.

Numeric text is coerced, loudly. Values inserted as text make pandas use object for the whole column even when every one of them is a valid number; '12.0' is recoverable and is recovered, with a UserWarning naming the columns, because a measurement stored as text is a database that wants looking at. 'n/a' is not recoverable and is never quietly turned into NaN and fitted on: it raises ModelFeatureSchemaError naming every offending column at once.

The input frame is returned unchanged when no conversion is needed. A shallow copy is made on the first conversion, so callers do not have their source data mutated and wide database joins do not get copied needlessly.

Parameters:
  • frame – the measurements frame. Anything else — a Series included — raises ModelFeatureSchemaError, not TypeError.

  • extra_features – names to treat as declared features whatever the feature dictionary makes of them. The only way to get an unrecognised text column repaired, and it opts that column into the error above too.

  • exclude – names to leave alone entirely — never repaired, never reported. Tested first, so it overrides extra_features. It is iterated, so a bare string excludes its letters and hence nothing.

  • allow_unknown – widen what counts as a feature to unrecognised columns — but an unrecognised column is then skipped rather than repaired, so this only ever converts fewer columns. It reaches DERIVED_MODEL_FEATURES as well: with it set, a text or all-NULL recruitment stays object and is then silently dropped by model_feature_columns() instead of being read as numbers.

spacr.schema.column_id(column: Any, *, strict: bool = False) → str[source]

Return the canonical 'c<N>' column id.

Parameters:
  • column – column index, or an already-prefixed 'c<N>'.

  • strict – raise instead of preserving an unparseable token.

Returns:

'c<N>'.

spacr.schema.column_index(value: Any) → int | None[source]

'c12' → 12; an unparseable id → None.

Parameters:

value – prefixed column id or numeric column token.

spacr.schema.comparable_key_value(value: Any) → str[source]

One metadata value, reduced to the form two spellings are compared in.

1, 1.0, '01' and ' 1 ' are the same well. A dtype difference is not a disagreement, and this is the whole reason the comparison is not Series.equals: a naive equality warns on every file that stored one copy of the well as text and the other as a number, and a warning that fires every time teaches the user to ignore the one that matters.

Missing is its own value: None, NaN and '' all reduce to '', so two columns that are both blank on a row agree there.

Parameters:

value – a single cell.

Returns:

the comparison string.

spacr.schema.comparable_key_values(values) → Tuple[str, ...][source]

comparable_key_value() over a column.

Parameters:

values – iterable of metadata cells to reduce for comparison.

spacr.schema.compose_prc(plate: Any, row: Any, column: Any) → str[source]

Return the prc well key: 'plate1_r1_c1'.

Parameters:
  • plate – plate id.

  • row – row index or 'r<N>'.

  • column – column index or 'c<N>'.

Returns:

the composed key.

spacr.schema.compose_prc_column(df, columns=None)[source]

Compose escaped plate-row-column identifiers for a frame.

Parameters:
  • df (pandas.DataFrame) – Frame containing the well-key columns.

  • columns (sequence of str, optional) – Plate, row, and column field names. Defaults to WELL_KEY_COLUMNS.

Returns:

pandas.Series – Canonical prc identifiers aligned to df.

Raises:

KeyError – If any required key column is absent.

Notes

Plate values are escaped with the same rules as compose_prc(), so separators and percent characters cannot create ambiguous keys. Legacy unescaped identifiers remain readable through parse_prcf().

spacr.schema.compose_prcf(plate: Any, row: Any, column: Any, field: Any, time: Any = None) → str[source]

Return the prcf field key.

'plate1_r1_c1_f2', or 'plate1_r1_c1_f2_t3' when time is given. The timepoint goes after the field — that is the order _map_wells(timelapse=True) writes and every table on disk carries.

Parameters:
  • plate – plate id.

  • row – row index or 'r<N>'.

  • column – column index or 'c<N>'.

  • field – field token.

  • time – timepoint token, or None outside a timelapse.

Returns:

the composed key.

spacr.schema.compose_prcfo(plate: Any, row: Any, column: Any, field: Any, obj: Any, time: Any = None, object_type: Any = None) → str[source]

Return the prcfo object key: 'plate1_r1_c1_f2_o7'.

With a timepoint the object still goes last: 'plate1_r1_c1_f2_t3_o7'. That matches both utils._map_wells_png(timelapse=True) and the prcf + '_' + 'o' + object_label composition in io._read_and_join_tables.

With an object_type the object component carries it: 'plate1_r1_c1_f2_nucleus7'. The untyped form is unchanged, which is why every prcfo already on disk still composes and parses byte for byte as it did — the type is a refinement of the key, not a new spelling of it.

Parameters:
  • plate – plate id.

  • row – row index or 'r<N>'.

  • column – column index or 'c<N>'.

  • field – field token.

  • obj – object label, bare, 'o<N>' or '<type><N>'.

  • time – timepoint token, or None.

  • object_type – the object table this object came from, or None for “not stated”.

Returns:

the composed key.

spacr.schema.correct_metadata_column_names(df)[source]

Rename legacy metadata columns to the canonical spaCR names.

A thin name over spacr.schema.canonicalise_frame(), which is the one vocabulary: case- and punctuation-insensitive, so Plate, PLATE, plate_id and plateName all become plateID. This function used to carry its own list of six spellings, matched case-sensitively, and that list is why a CSV whose header said Column reached a fit with no columnID in it.

Two things it does that the shared vocabulary deliberately does not:

  • grna_name -> grna. Not in the shared vocabulary because the sequencing CSVs are still read with grna_name by name in spacr.submodules and spacr.plot; renaming it at the database migration would break those reads, and a rename nobody can see is worse than a spelling.

  • plate_row split into plateID and rowID. That is a value split rather than a rename, so it has no place in a name table.

Parameters:

df – DataFrame whose columns may use legacy names.

Returns:

A DataFrame carrying the canonical names. The renames are not applied in place, so the caller must use the returned frame.

spacr.schema.escape_field_stem_plate(name: Any, *, timelapse: bool = False) → str[source]

Escape the plate component of a merged-stack field stem.

Parameters:
  • name (Any) – File name, path, or stem in plate_well_field[_time] form.

  • timelapse (bool, default=False) – Treat the final component as a timepoint. A numeric trailing timepoint is also recognized in non-time-lapse merged-stack names.

Returns:

str – The stem with only its plate component filename-escaped.

Raises:

KeyParseError – If the stem does not contain a plate and the required tail fields.

Notes

Escaping at write time preserves plate names that contain underscores. The tail rules match parse_field_stem().

spacr.schema.escape_filename_component(token: Any) → str[source]

Escape one free-text component for a separator-delimited filename.

This uses the same reversible table as join keys. In particular, 'my_plate' becomes 'my%5Fplate' and a literal percent is escaped first, so parsing cannot merge it with an encoded separator.

spacr.schema.field_id(field: Any, *, strict: bool = False) → str[source]

Return the canonical 'f<N>' field id.

'3', '003', 's3', 'F003' and 3 all give 'f3'. A token holding no integer is preserved ('xy' → 'fxy') rather than becoming 'f0'; see the module docstring for why.

Parameters:
  • field – field token.

  • strict – raise instead of preserving an unparseable token.

Returns:

'f<N>', or 'f<token>' for an unparseable token.

spacr.schema.field_index(value: Any) → int | None[source]

'f2' → 2; 'fxy' → None.

Parameters:

value – prefixed field id or numeric field token.

spacr.schema.fold_column_name(name: Any) → str[source]

The form a column name is looked up by: lower case, no punctuation.

'Plate_ID', 'plate id', 'plate.id' and 'plateID' all fold to 'plateid'. Folding is what keeps LEGACY_COLUMN_NAMES a short list of words rather than a combinatorial table of every separator a plate reader has ever emitted. The supported aliases include Plate, PLATE, plate, plateid, plate_id, plate_name and plateName and this is five entries, not seven.

Parameters:

name – a column name.

Returns:

the folded form.

spacr.schema.is_object_type(object_type: Any) → bool[source]

Whether object_type can be written into an object key.

The question a reader asks about the table it just loaded before stamping OBJECT_TYPE_KEY on the frame. png_list, a summary table or a user’s own table answer False, and the frame stays untyped — which is the key spaCR has always written, so nothing regresses.

Parameters:

object_type – candidate table or object-type name.

spacr.schema.is_positional_pair(row: Any, column: Any) → bool[source]

True when (rowID, columnID) came from the positional passthrough.

parse_well() puts an unrecognisable well into both slots verbatim, so an unprefixed pair of equal values is that passthrough and not a real row and column. Without this check ('12', '12') looks like row 12 / column 12 and well_id() happily renders it 'L12' — a well name for a well that was never identified.

Only strings can be a passthrough. A bare int is unambiguously an index — well_id(1, 1) is a caller asking for well A01, not a well that failed to parse — so an integer pair is never flagged, however equal. (This is not hypothetical: it is the bug the round-trip test in tests/test_schema.py caught in the first version of this function, where well_id(1, 1) raised.)

Parameters:
  • row – the rowID as stored.

  • column – the columnID as stored.

Returns:

whether the pair is a passthrough rather than a position.

spacr.schema.is_positional_well(well: Any) → bool[source]

True when well is a bare number rather than <letters><digits>.

Some acquisitions name wells '12'. There is no way to know whether that means row 1 column 2 or the twelfth well, so parse_well() passes it through into both slots unchanged — which is what all five existing implementations do, and there is data on disk keyed that way. This predicate lets a caller detect the case instead of discovering it from a rowID that does not start with r.

Parameters:

well – well identifier.

Returns:

True when the well holds no row letters.

spacr.schema.is_provenance_column(name: Any) → bool[source]

Return whether name is identity, annotation, or run provenance.

original_* columns, the names a plate had before conversion, count as provenance and never as measured features.

spacr.schema.is_row_column_pair(row: Any, column: Any) → bool[source]

True when (row, column) is recognisably a well’s row and column.

Deliberately narrow. It is the guard that stops a right-to-left key parse from absorbing a deeper key into an underscored plate id, so it must reject a (columnID, fieldID) pair and a (fieldID, objectID) pair: a columnID is never 'f1' and a rowID is never 'c1'.

parse_prcf() and ml._split_prc are both right-to-left parses of a separator-joined key whose leftmost component may itself contain the separator, and both need exactly this test to tell “the plate is called exp1_plate1” from “you handed me a key one level too deep”. It lives here so there is one answer to “is this a row and a column?”.

Parameters:
  • row – candidate rowID token.

  • column – candidate columnID token.

Returns:

whether the pair can be a row and a column.

spacr.schema.is_within_plate_format(row: Any, column: Any, n_wells: int) → bool[source]

True when (row, column) lies inside an n_wells plate.

Parameters:
  • row – row index or 'r<N>'.

  • column – column index or 'c<N>'.

  • n_wells – a key of PLATE_FORMATS.

Returns:

whether the position exists on that plate.

Raises:

KeyParseError – when n_wells is not a standard format.

spacr.schema.legacy_map_wells(file_name: Any, timelapse: bool = False) → Tuple[str, ...][source]

utils._map_wells reproduced bit for bit, 'error' tuple and all.

Used by tests/test_schema.py to assert that the canonical parser agrees with the legacy one on every well that works today, so the migration is provably a strict repair rather than a change of contract.

Parameters:
  • file_name – stack file name.

  • timelapse – parse a trailing timepoint.

Returns:

the same tuple _map_wells returns.

spacr.schema.legacy_safe_int_convert(value: Any, default: Any = 0) → Any[source]

utils._safe_int_convert exactly, including the 0 default.

Kept only for the migration tests. New code calls parse_int_token(), which returns None.

Parameters:
  • value – token to convert.

  • default – what to return on ValueError. Note that TypeError — which is what None raises — is not caught, here or in the original.

Returns:

the int, or default.

spacr.schema.legacy_well_ids(well: Any) → Tuple[str, str][source]

utils._map_wells’ well handling exactly, raising where it raises.

Parameters:

well – well identifier.

Returns:

(rowID, columnID).

Raises:
  • ValueError – on the wells _map_wells swallows into 'error'.

  • IndexError – on an empty well, as _map_wells does.

spacr.schema.letters_from_row_index(index: int) → str[source]

Inverse of row_index_from_letters(). 27 → 'AA'.

Parameters:

index – 1-based row index.

Returns:

the row letters.

Raises:

KeyParseError – when index is not a positive integer.

spacr.schema.model_feature_columns(frame, *, extra_features=(), exclude=(), allow_unknown: bool = False) → list[str][source]

Select numeric model inputs by schema role, never by dtype alone.

A column is eligible when the feature dictionary identifies a measurement, when it belongs to an object-table measurement namespace, or when the caller explicitly names it in extra_features. Identity and provenance are always excluded—even if SQLite/pandas represents them as numbers. allow_unknown is for generic statistics over user-created frames; database-backed model paths should retain the strict default.

Every unusable column is reported in one error, with its dtype and why it is unusable. Refusing the first one and stopping made a user fix them one whole run at a time.

Parameters:
  • frame – the frame to select from; anything else (a Series included) raises ModelFeatureSchemaError rather than TypeError.

  • extra_features – names to declare as features whatever the feature dictionary makes of them. It cannot promote an identity or provenance column — those are dropped before it is consulted — but it does turn a non-numeric column from a silent omission into the error below.

  • exclude – names dropped before any check, so it overrides extra_features and is the escape hatch the error message points at. Iterated, so passing one bare column name excludes its letters only.

  • allow_unknown – also accept unrecognised columns, but only those already of a numeric dtype: an unrecognised non-numeric one is skipped instead of reported. It reaches DERIVED_MODEL_FEATURES as well, so a recruitment column read back as text vanishes from the selection rather than raising.

Raises:

ModelFeatureSchemaError – if a declared feature is non-numeric.

spacr.schema.model_feature_frame(frame, **kwargs)[source]

Return frame restricted to model_feature_columns().

Unlike model_feature_columns() this owns the data it hands back, so it repairs what is losslessly repairable first (coerce_model_feature_types()) instead of refusing a frame whose only fault is that pandas typed an all-NULL measurement object.

Parameters:

frame – pandas frame to coerce and restrict to model features.

spacr.schema.normalise_plate_columns(frame)[source]

Collapse a doubled p prefix in every column that carries a plate.

Applied on READ, so nothing on disk is rewritten and an old database keeps working. Modifies frame in place and returns it, which is what the two callers that predate this function both did.

Parameters:

frame – any frame read from a database or a CSV.

Returns:

the same frame.

spacr.schema.object_id(label: Any, *, object_type: Any = None, strict: bool = False) → str[source]

Return the canonical object id used in prcfo.

'o<N>' when the object’s type is not stated, and '<type><N>' when it is — object_id(7, object_type='nucleus') is 'nucleus7'. The type goes into the key, not beside it, because the key is the only thing that travels: a lasso publishes strings, and a string that cannot say which of a cell’s four children it means is a string that opens the wrong crop.

An already-composed id round-trips, with or without a type (object_id('nucleus7') == 'nucleus7'), so this is idempotent and safe to apply to a value that has already been through it.

Parameters:
  • label – object label — bare, 'o<N>', or '<type><N>'.

  • object_type – the object table this object came from, or None for “not stated”. A type on label that disagrees with this is an error rather than a silent overwrite.

  • strict – raise instead of preserving an unparseable token.

Returns:

the object id.

Raises:

KeyParseError – on an empty token, a conflicting type, any bad token when strict, or a composition that would not read back.

spacr.schema.object_index(value: Any) → int | None[source]

'o41' and 'nucleus41' → 41; an unparseable id → None.

Split through split_object_id() rather than by stripping 'o', so a typed id reads back as the number it is instead of as None.

Parameters:

value – typed, untyped, or bare object-label token.

spacr.schema.object_table_schema(table: str) → ObjectTableSchema[source]

Return the canonical schema for table.

Parameters:

table – canonical object measurement table name.

Raises:

ObjectTableSchemaError – when no canonical contract exists.

spacr.schema.object_type_prefix(object_type: Any) → str[source]

Return the canonical prefix an object of object_type is keyed with.

None — the type is not stated — gives OBJECT_PREFIX. Anything else is folded to lower case and checked, because the prefix has to be separable from the label that follows it with no separator in between:

  • it may not be empty, or every typed key would be an untyped one;

  • it may not contain KEY_SEPARATOR, for the reason _check_plate() gives — the key is separator-joined;

  • it may not contain a digit, or the split is ambiguous. Type 'cell1' with label 7 and type 'cell' with label 17 both write 'cell17', which is the “two identities, one key” failure this whole module exists to end.

The vocabulary is closed: OBJECT_TYPES and nothing else. An open one would mean every unrecognised token in the object slot became a type, and 'plate1_r1_c1_f2_x7' — which is not an object key — would parse as object 7 of type 'x'. Widening what counts as a key is how a malformed key becomes a plausible wrong answer instead of an error, and a closed vocabulary is the same choice KEY_PREFIXES already makes for rows, columns, fields and timepoints.

A caller holding a table name it is not sure about asks is_object_type() first and leaves the frame untyped otherwise — an untyped key is exactly what spaCR wrote before, so that is a no-change, not a failure.

Parameters:

object_type – one of OBJECT_TYPES, or None.

Returns:

the prefix, lower case.

Raises:

KeyParseError – for a type that cannot be written into a key.

spacr.schema.object_type_summary(roles: Sequence[str]) → str[source]

Describe object types once, collapsing organelle slots into a regex.

Parameters:

roles – internal object identifiers, in display order.

Returns:

comma-separated types. Organelle identifiers are represented by one exact pattern, such as organelle(?:[b-z]|[a-z]{2})? for all slots. This changes presentation only, never stored identifiers or parsing.

spacr.schema.parse_field_stem(name: Any, *, timelapse: bool = False, strict: bool = False) → FieldID[source]

Parse a merged-stack file name into a FieldID.

Parameters:
  • name (Any) – File name, path, or stem in plate_well_field[_time] form.

  • timelapse (bool, default=False) – Parse a required trailing timepoint and include it in the identity.

  • strict (bool, default=False) – Reject non-numeric field/time tokens and nonstandard wells.

Returns:

FieldID – Parsed plate, well, field, and optional timepoint identity.

Raises:

Notes

A non-time-lapse call accepts one extra numeric timepoint emitted by the merged-stack writer, but omits it from the returned identity. Other extra components are rejected.

Examples

>>> parse_field_stem('plate1_A01_3').prcf
'plate1_r1_c1_f3'
spacr.schema.parse_int_token(token: Any, *, allow_prefix: bool = True) → int | None[source]

Return the integer token denotes, or None — never 0.

This is the replacement for utils._safe_int_convert, and the whole point of it is the return type. _safe_int_convert answers “what number is this?” with 0 when the honest answer is “there isn’t one”, and 0 is a perfectly good field id, so the lie is unrecoverable downstream. None is not a field id, so every caller is forced to decide what to do — and the callers here do decide, see field_id().

Vendor prefixes are understood, because they are a spelling of a number rather than a different number: an ImageXpress site s3, a CellVoyager field F003 and a bare 3 are the same field, and a pipeline that gave them three different ids would be just as wrong as one that gave them all f0.

Parameters:
  • token – anything — a string, an int, a float, None.

  • allow_prefix – strip one or two leading ASCII letters when what follows is all digits. Default True.

Returns:

the integer, or None when the token holds no integer.

Example

>>> parse_int_token('003'), parse_int_token('s3')
(3, 3)
>>> parse_int_token('T0001'), parse_int_token('x')
(1, None)
>>> parse_int_token('') is None, parse_int_token(None) is None
(True, True)
spacr.schema.parse_object_stem(name: Any, *, timelapse: bool = False, strict: bool = False) → ObjectID[source]

Parse a crop-PNG file name into an ObjectID.

Parameters:
  • name (Any) – File name, path, or stem in plate_well_field[_time]_object form.

  • timelapse (bool, default=False) – Parse a timepoint between the field and object components.

  • strict (bool, default=False) – Reject non-numeric identity tokens and nonstandard wells.

Returns:

ObjectID – Parsed field identity with the final object label attached.

Raises:
spacr.schema.parse_prcf(text: Any) → FieldID[source]

Parse a prcf string back into a FieldID.

Parsed right to left, which is what makes it correct: the components are optional in the middle (timeID may or may not be there), and ml.py splits prcfo left to right into a fixed five columns, so a timelapse key with six parts silently misaligns every column.

Extra components are not automatically an underscored plate. A key with more components than plate_row_column_field[_time] is one of two things, and they mean opposite things:

  • 'exp1_plate1_r2_c12_f1' — a plate id containing the separator. The right-to-left rule handles it, and that is the case the absorption exists for.

  • 'plate1_r1_c1_f1_f2' — a key one level too deep, or a key whose components are not what they claim. Absorbing it would return plateID='plate1_r1', rowID='c1', columnID='f1' — half the well inside the plate and a field id in the column slot — and every per-well figure grouped on that is a plausible wrong number with nothing anywhere saying so.

The two are told apart with is_row_column_pair(), which is the same guard ml._split_prc applies for the same reason. Anything else with extra components is refused rather than guessed at.

Parameters:

text – e.g. 'plate1_r1_c1_f2' or 'plate1_r1_c1_f2_t3'.

Returns:

the FieldID.

Raises:

KeyParseError – when the string is not a prcf.

spacr.schema.parse_prcfo(text: Any) → ObjectID[source]

Parse a prcfo string back into an ObjectID.

The object prefix is stripped before the label is re-canonicalised, which is what makes this the inverse of compose_prcfo() for every label rather than only the numeric ones. object_id() reads an already-prefixed numeric id back out of its prefix ('o7' → 7 → 'o7'), but it cannot do that for a preserved non-numeric token, so handing it 'oxy' used to yield 'ooxy': the key 'p_r1_c1_f1_oxy' parsed to 'p_r1_c1_f1_ooxy', and parsing that yielded 'ooooxy'. A key that grows every time it passes through the parser joins to nothing.

A typed object component is read back with its type: 'plate1_r1_c1_f2_nucleus7' gives objectType == 'nucleus'. An untyped one gives None, which is what every key written before object types existed means and is not the same fact as “cell”.

Parameters:

text – e.g. 'plate1_r1_c1_f2_o7', 'plate1_r1_c1_f2_t3_o7' or 'plate1_r1_c1_f2_nucleus7'.

Returns:

the ObjectID.

Raises:

KeyParseError – when the string is not a prcfo.

spacr.schema.parse_well(well: Any, *, strict: bool = False) → Tuple[str, str][source]

Return (rowID, columnID) for a well identifier.

'A01', 'a1', 'A-01' and ' A01 ' all give ('r1', 'c1'). 'AA01' — a real 1536-plate well — gives ('r27', 'c1'), where utils._map_wells raises into 'error' and utils._map_wells_png returns ('r1', 'c0').

A well with letters but no digits ('A') has no column. Under _map_wells_png it became 'c0', i.e. indistinguishable from a genuine column 0; here it raises, because a well with no column is not a well.

A bare number is passed through into both slots — see is_positional_well().

Parameters:
  • well – well identifier of any of the above shapes.

  • strict – also reject the bare-number passthrough.

Returns:

(rowID, columnID).

Raises:

WellParseError – when the well is empty, has no column, or is a bare number and strict is set.

Example

>>> parse_well('A01'), parse_well('aa1')
(('r1', 'c1'), ('r27', 'c1'))
spacr.schema.plate_format_for(row: Any, column: Any) → int | None[source]

Return the smallest standard plate format containing (row, column).

A column past 24 is not an error — a 1536-well plate has 48 of them — so nothing in this module rejects one. This is how a caller that does care checks.

Parameters:
  • row – row index or 'r<N>'.

  • column – column index or 'c<N>'.

Returns:

the well count of the smallest format that contains the position, or None when it fits no standard plate.

spacr.schema.resolve_metadata_collisions(frame, *, report=None, warn=None)[source]

Collapse every group of columns that mean the same metadata key.

Several columns can normalise to one key – a file carrying well, wellID and well_name has three opinions about which well a row came from, and every join downstream is silently picking one of them. What happens is decided by whether they agree row by row (comparable_key_value()):

  • they agree – keep one, drop the rest, and print. Nothing is wrong, so nothing warns.

  • they disagree – the same action, and a warning naming the columns, the choice, and how many rows differ. That count is the point: “they disagree” is not actionable, “3 of 40 000 rows disagree” is a typo and “40 000 of 40 000” is the wrong file.

Only METADATA_KEYS are collapsed. Two feature columns that normalise alike keep both spellings, because a measurement is data and dropping one to tidy a name is not this function’s call to make.

Duplicate column labels are handled positionally, so a frame that already carries two columns both literally named rowID – which pandas allows and to_sql refuses – is repaired rather than raising.

Parameters:
  • frame – pandas.DataFrame.

  • report – called with each agreeing collision’s message. Pass print to show them; None is silent.

  • warn – called with each disagreeing collision’s message. None routes to warnings.warn(); pass a callable to capture them.

Returns:

(frame, collisions). The frame is new when anything changed and frame itself when nothing did.

spacr.schema.row_id(row: Any, *, strict: bool = False) → str[source]

Return the canonical 'r<N>' row id.

Accepts an index (1, '1'), an already-prefixed id ('r1') — which round-trips rather than becoming 'rr1' — or row letters ('A', 'AA').

Parameters:
  • row – row index, 'r<N>', or row letters.

  • strict – raise instead of preserving an unparseable token.

Returns:

'r<N>'.

Raises:

KeyParseError – on an empty token, or any bad token when strict.

spacr.schema.row_index(value: Any) → int | None[source]

'r3' → 3; 'C' → 3; an unparseable id → None.

Parameters:

value – prefixed row id, row letters, or numeric row token.

spacr.schema.row_index_from_letters(letters: Any) → int | None[source]

'A' → 1, 'Z' → 26, 'AA' → 27, 'AF' → 32.

Bijective base 26. Multi-letter rows are not an edge case: a 1536-well plate has 32 rows and runs A…Z, AA…AF. Both utils._map_wells (which raises, becoming 'error') and utils._map_wells_png (which yields 'c0') get these wrong, in two different ways. This matches plate_qc._alpha_to_index exactly, so the QC module and the database agree.

Parameters:

letters – one or more ASCII letters, any case.

Returns:

the 1-based row index, or None when letters is not purely alphabetic or is empty.

spacr.schema.screen_id(screen: Any = None) → str[source]

Return the canonical screen id, defaulting an absent one.

Free-form text, like the plate id, because it is the name a user gave an experiment. It is not prefixed and not parsed back apart: the whole point of SCREEN_KEY is that it is a dimension you block on, facet by and colour with as it stands.

Absence is the case that matters. None, '' and whitespace all mean “this project has one screen”, and they become DEFAULT_SCREEN rather than raising — every project that exists today has no screen anywhere in its settings, and demanding one would stop all of them from opening. Contrast _check_plate(), which does raise: an empty plate is a broken key, but an empty screen is an ordinary single-screen run.

An empty value is never left empty inside a frame either, because a blank screen groups with every other blank screen — which is exactly the silent pooling spacr.multi_database exists to refuse.

Parameters:

screen – the screen label, or None.

Returns:

the label, stripped, or DEFAULT_SCREEN.

Example

>>> screen_id('tsg101'), screen_id(None)
('tsg101', 'screen1')
spacr.schema.split_object_id(token: Any, *, require_prefix: bool = True) → Tuple[str | None, str][source]

Split an object id into (object type, label).

The inverse of object_id():

split_object_id('nucleus7')  -> ('nucleus', '7')
split_object_id('o7')        -> (None, '7')
split_object_id('omulti')    -> (None, 'multi')
split_object_id('7')         -> (None, '')        # not an object id

A type of None means not stated, which is what 'o' has always meant and is exactly what every key written before object types existed carries. It is not “unknown and therefore probably a cell”.

Parameters:
  • token – the last component of a prcfo.

  • require_prefix – when False a bare label ('7') is accepted and read as an untyped id. That is the shape spacr.selection.object_keys() writes, where the label is joined bare rather than through OBJECT_PREFIX.

Returns:

(type or None, label). An empty label means token is not an object id at all; callers must check it rather than assuming the split succeeded.

spacr.schema.strip_prefix(value: Any, prefix: str) → str[source]

Remove one leading prefix from value if it is there.

Parameters:
  • value – the id, e.g. 'r12'.

  • prefix – the single-letter prefix, e.g. 'r'.

Returns:

the remainder, e.g. '12'.

spacr.schema.table_key_columns(table: str, *, timelapse: bool = False) → Tuple[str, ...][source]

Return the columns that identify a row of table.

Parameters:
  • table – table name.

  • timelapse – include timeID.

Returns:

the key columns, most significant first.

Raises:

KeyParseError – when table is not one spaCR owns.

Example

>>> table_key_columns('cell')
('plateID', 'rowID', 'columnID', 'fieldID', 'object_label')
spacr.schema.time_id(time: Any, *, strict: bool = False) → str[source]

Return the canonical 't<N>' timepoint id.

'T0003' → 't3'. Under the old _safe_int_convert every T#### token became t0, which collapsed a whole timelapse onto one frame.

Parameters:
  • time – timepoint token.

  • strict – raise instead of preserving an unparseable token.

Returns:

't<N>'.

spacr.schema.time_index(value: Any) → int | None[source]

't7' → 7; an unparseable id → None.

Parameters:

value – prefixed timepoint id or numeric time token.

spacr.schema.unescape_filename_component(token: Any) → str[source]

Invert escape_filename_component().

spacr.schema.validate_object_table_frame(frame, table: str, *, timelapse: bool | None = None, metadata_column_map=None, metadata_well_column=None, metadata_pseudo_source=None, allow_pseudo_metadata: bool = False, metadata_prompt=None, metadata_cache_key=None, metadata_mapping_path=None)[source]

Validate an object-table frame against its canonical contract.

Validation is deliberately strict at the writer boundary and compatibility-preserving in shape:

  • required identity/provenance columns must exist and be non-null;

  • labels (and present parent links) must be positive integers;

  • prcf must exactly match the component key columns;

  • one write batch may contain at most one row per object key;

  • measurement stamps are all present or all absent;

  • features from another compartment are rejected, and this table’s own feature namespace must be numeric.

Extra columns are allowed because annotation columns are user-defined and historical databases contain extensions. Legacy metadata spellings are canonicalised on the returned copy before validation.

pandas is imported only when this function is called. Importing spacr.schema itself remains standard-library-only for CLI, multiprocessing, and resume preflight paths.

Parameters:
  • frame – pandas DataFrame to validate.

  • table – one of CANONICAL_OBJECT_TABLES.

  • timelapse – require/forbid timeID; None infers it.

Returns:

canonical-column DataFrame copy.

Raises:

ObjectTableSchemaError – on any contract violation.

spacr.schema.well_id(row: Any, column: Any) → str[source]

Return the canonical well name: ('r3', 'c7') → 'C07'.

The inverse of parse_well() for wells that have one. Matches plate_qc.well_id.

Parameters:
  • row – row index or 'r<N>' or row letters.

  • column – column index or 'c<N>'.

Returns:

the well name, zero padded to two digits.

Raises:

KeyParseError – when either index is unusable, or the pair is a positional passthrough (see is_positional_pair()).

Nested helpers

_non_numeric_feature_error.diagnostic_dtype(dtype) → str

Return stable user-facing text for a pandas feature dtype.

Pandas StringDtype is reported as the established object wording; other extension and NumPy dtypes retain their own names.

spacr/schema.py:2174

validate_object_table_frame._validate_positive_integer(column: str, *, nullable: bool = False)

Require positive integral values in one canonical key column.

When nullable is true, nulls are ignored while every populated value is still checked; failures name the table, column, and examples.

spacr/schema.py:2987