spacr.schema¶
Canonical database keys, filename identities, and table schemas.
spaCR measurement tables share the plateID, rowID, columnID, and
fieldID columns, with timeID added for time-lapse data. This module
defines those names, their composed prc/prcf/prcfo forms, and the
parsers used by database, image, and GUI code.
Numeric tokens may include common vendor prefixes such as s3 or T0003.
Unparseable but non-empty tokens remain distinct in permissive mode and raise
KeyParseError in strict mode; missing identity components always
raise. Legacy helpers are retained for reading and testing older data.
The module has no third-party import-time dependencies. Data-frame helpers import pandas only when called.
Exceptions¶
A key token was absent, or was rejected under |
|
A declared model feature cannot be represented as numeric input. |
|
An object-table frame violates its declared column or row contract. |
|
Base for every failure to build or read a spaCR key. |
|
A well identifier could not be turned into a row and a column. |
Classes¶
What happened when several columns claimed one metadata key. |
|
One imaging field's identity — the key of every measurement row. |
|
One segmented object's identity — the |
|
Declarative contract for one object measurement table. |
Functions¶
|
Parse a name column into the canonical key columns. |
|
Return |
|
Return the canonical spelling of a metadata or feature column name. |
|
The plate id in the form the rest of spaCR keys on. |
|
|
Return |
|
|
The whole vocabulary applied to one frame: the reader's normaliser. |
|
Repair the two representations of a numeric measurement pandas fumbles. |
|
Return the canonical |
|
|
|
One metadata value, reduced to the form two spellings are compared in. |
|
|
|
Return the |
|
Compose escaped plate-row-column identifiers for a frame. |
|
Return the |
|
Return the |
Rename legacy metadata columns to the canonical spaCR names. |
|
|
Escape the plate component of a merged-stack field stem. |
|
Escape one free-text component for a separator-delimited filename. |
|
Return the canonical |
|
|
|
The form a column name is looked up by: lower case, no punctuation. |
|
Whether |
|
True when |
|
True when |
|
Return whether |
|
True when |
|
True when |
|
|
|
|
|
|
|
Inverse of |
|
Select numeric model inputs by schema role, never by dtype alone. |
|
Return |
|
Collapse a doubled |
|
Return the canonical object id used in |
|
|
|
Return the canonical schema for |
|
Return the canonical prefix an object of |
|
Describe object types once, collapsing organelle slots into a regex. |
|
Parse a merged-stack file name into a |
|
Return the integer |
|
Parse a crop-PNG file name into an |
|
Parse a |
|
Parse a |
|
Return |
|
Return the smallest standard plate format containing |
|
Collapse every group of columns that mean the same metadata key. |
|
Return the canonical |
|
|
|
|
|
Return the canonical screen id, defaulting an absent one. |
|
Split an object id into |
|
Remove one leading |
|
Return the columns that identify a row of |
|
Return the canonical |
|
|
|
Invert |
|
Validate an object-table frame against its canonical contract. |
|
Return the canonical well name: |
Module Contents¶
- exception spacr.schema.KeyParseError[source]¶
Bases:
SchemaErrorA key token was absent, or was rejected under
strict=True.Initialize self. See help(type(self)) for accurate signature.
- exception spacr.schema.ModelFeatureSchemaError[source]¶
Bases:
SchemaErrorA declared model feature cannot be represented as numeric input.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.schema.ObjectTableSchemaError[source]¶
Bases:
SchemaErrorAn object-table frame violates its declared column or row contract.
Initialize self. See help(type(self)) for accurate signature.
- exception spacr.schema.SchemaError[source]¶
Bases:
ValueErrorBase for every failure to build or read a spaCR key.
A subclass of
ValueErrorso that call sites which already guard a parse withexcept ValueErrorkeep working.Initialize self. See help(type(self)) for accurate signature.
- exception spacr.schema.WellParseError[source]¶
Bases:
SchemaErrorA well identifier could not be turned into a row and a column.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.schema.ColumnCollision[source]¶
What happened when several columns claimed one metadata key.
- Parameters:
canonical – canonical metadata key claimed by all source columns.
sources – source columns in their original frame order.
chosen – source column retained under the canonical name.
dropped – redundant source columns removed from the frame.
disagreeing_rows – number of rows whose source values disagreed.
rows – total number of rows compared.
- class spacr.schema.FieldID[source]¶
One imaging field’s identity — the key of every measurement row.
- Parameters:
plateID – plate identifier, escaped when composed into a field key.
rowID – stored row component, normally
r<N>but potentially a legacy positional-well passthrough.columnID – stored column component, normally
c<N>but potentially a legacy positional-well passthrough.fieldID – canonical
f<label>imaging-field component.timeID – canonical
t<label>timepoint component inserted after the field, orNoneoutside a timelapse.
- classmethod build(plate: Any, well: Any = None, field: Any = None, time: Any = None, *, row: Any = None, column: Any = None, strict: bool = False) FieldID[source]¶
Construct from a well string, or from a row and column.
- Parameters:
plate – plate id.
well – well identifier, e.g.
'A01'. Mutually exclusive withrow/column.field – field token.
time – timepoint token, or
None.row – row index or
'r<N>', when there is no well string.column – column index or
'c<N>'.strict – reject unparseable field/time tokens and odd wells.
- Returns:
the
FieldID.- Raises:
KeyParseError – when neither a well nor a row/column pair is given.
- to_dict(*, include_prcf: bool = False) Dict[str, str][source]¶
Return the identity as the dict the tables carry.
- Parameters:
include_prcf – also emit
prcandprcf.- Returns:
{plateID, rowID, columnID, fieldID[, timeID][, prc, prcf]}.
- with_object(obj: Any, object_type: Any = None) ObjectID[source]¶
Return the
ObjectIDfor one object in this field.- Parameters:
obj – object label, bare or already prefixed.
object_type – which object table it came from, or
Nonefor “not stated”. A nucleus and a pathogen with the same label in the same field are two objects, and without this they are one key.
- class spacr.schema.ObjectID[source]¶
One segmented object’s identity — the
prcfoa merged row is keyed on.- Parameters:
plateID – plate identifier carried by the containing field and escaped when composed into a field or object key.
rowID – stored row component, normally
r<N>but potentially a legacy positional-well passthrough.columnID – stored column component, normally
c<N>but potentially a legacy positional-well passthrough.fieldID – canonical
f<label>imaging-field component.objectID –
o<label>when the object’s type is not stated, or<type><label>when it is. The type lives inside this field rather than beside it so that twoObjectIDvalues compare equal exactly when they name the same object — a separateobjectTypefield would let('o7', 'nucleus')and('nucleus7', None)describe one object and compare unequal.timeID – canonical
t<label>timepoint component inserted between field and object, orNoneoutside a timelapse.
- to_dict(*, include_prcf: bool = False) Dict[str, str][source]¶
Return the identity as a dict, including
prcfo.object_typeappears only when the object has one. An untyped id emits the same dict it always did, so a reader that never learned about types sees no new column on data that has no type to report.
- class spacr.schema.ObjectTableSchema[source]¶
Declarative contract for one object measurement table.
Feature columns are open-ended because channel counts and enabled measurements vary per run, but they are not untyped: a feature written by a table starts with
<object_type>_and is numeric. Unknown annotation or provenance columns remain permitted so older databases and user-added labels are not destroyed by validation.- Parameters:
table – SQLite table name.
object_type – required feature-column prefix.
parent_column – optional link to a parent cell.
- feature_column(name: Any) bool[source]¶
Return whether
namebelongs to this table’s feature namespace.- Parameters:
name – candidate column name.
- row_key_columns(*, timelapse: bool = False) Tuple[str, ...][source]¶
Return the columns that must be unique within one write batch.
- validate(frame, *, timelapse: bool | None = None)[source]¶
Validate and return a canonical-column copy of
frame.- Parameters:
frame – pandas frame to validate against this table contract.
- property identifier_columns: Tuple[str, ...][source]¶
Prefixed link/label columns emitted by morphology measurement.
- spacr.schema.add_identity_columns(df, source: str = 'file_name', *, timelapse: bool = False, objects: bool = False, strict: bool = False, include_prcf: bool = True)[source]¶
Parse a name column into the canonical key columns.
The vectorised form of
parse_field_stem()/parse_object_stem(), for the writers that today dodf[[...]] = df[col].apply(lambda x: pd.Series(_map_wells(x)))— a line that positionally unpacks a tuple whose length changes withtimelapse, so a mismatched flag misaligns every column.- Parameters:
df –
pandas.DataFramewith a column of file names.source – name of that column. Default
'file_name'.timelapse – names carry a timepoint.
objects – names are crop PNGs, so also emit
prcfo.strict – reject unparseable tokens.
include_prcf – also emit
prcandprcf.
- Returns:
a new frame with the key columns added.
- Raises:
KeyParseError – when
sourceis not a column ofdf.
- spacr.schema.add_screen_column(df, screen: Any = None, *, overwrite: bool = False)[source]¶
Return
dfcarrying a filled-inSCREEN_KEYcolumn.The three cases, and why each behaves as it does:
The frame has no screen column. It gains one, holding
screenorDEFAULT_SCREEN. That is a single-screen project, which is every project written before and it must keep working.The frame has one, and
screenisNone. Its labels are kept. Relabelling a frame that already knows which experiment it came from would move rows between screens with nothing on screen to say so.The frame has one and
screenwas given. The caller is looking at the files and has said which screen this is, so it wins — but only because they said so.overwrite=False(the default) still fills blank values only; passoverwrite=Trueto restamp every row.
Blanks are never left blank:
None,NaNand''becomeDEFAULT_SCREEN, because an empty screen id is not an identity and every row carrying one would group with every other.- Parameters:
df –
pandas.DataFrame.screen – the screen label for rows that do not have one.
overwrite – replace existing labels instead of filling blanks.
- Returns:
a new frame;
dfis not modified.
- spacr.schema.canonical_column_name(name: Any) str[source]¶
Return the canonical spelling of a metadata or feature column name.
Metadata lookup folds case and punctuation (
fold_column_name()), so'RowID','rowid','Row Name'and'row_name'all resolve to'rowID'. Feature rewrites (LEGACY_COLUMN_PATTERNS) are case-sensitive. A name matching neither is returned unchanged.Note
What is deliberately not in the vocabulary
Aliases are whole words.
colis here because spaCR itself wrote it; a one-lettercis not, and never will be, because it would capture a measurement column calledcand rename a real variable into a plate key. A name that is not normalised is visible the moment a user looks at the picker; a measurement silently renamed tocolumnIDis not, and it corrupts the join it lands in.This is the only implementation.
spacr.database_schemaandspacr.utilsre-export this function rather than defining their own.Note
Migration note (one function, was two)
database_schema.canonical_column_nameused to be a second, narrower implementation: 11 aliases, matched case-sensitively. Which one a caller got depended on whether it had importedspacr.schemaorspacr.utils, and the two disagreed. Adopting this one widens what a database migration renames:plate_id,row_id,column_id,col_name,field_id,rowid,time,timepoint,channel_name,chan_idandslice_idare now renamed to their canonical spellings;any case variant is now renamed too, so a database column spelled
RoworRowIDis repaired instead of being left for a reader to trip over. SQLite compares identifiers case-insensitively, soRowID->rowIDis a pure respelling of one column, but pandas reports whatever spelling is stored — which is how a frame ends up with norowIDcolumn on a database that has one.
The non-destructive rule is unchanged: a table already carrying the canonical name keeps both columns.
- Parameters:
name – column name as it appears in a table or CSV.
- Returns:
the canonical name, or
nameunchanged.
Example
>>> canonical_column_name('column_name') 'columnID' >>> canonical_column_name('cell_periphery_25_percentile') 'cell_periphery_percentile_25' >>> canonical_column_name('cell_area') 'cell_area'
- spacr.schema.canonical_plate_id(plate: Any) str[source]¶
The plate id in the form the rest of spaCR keys on.
A legacy score CSV stamps its plate
pplate1while the sequencing counts stamp itplate1. The two then do not join, the merge returns zero rows, and the run dies several steps later inside a plot withKeyError: 0– nowhere near the mismatch and with nothing on screen naming a plate.This lives here, and not in
spacr.multi_databasewhere it was written, because there must be one normaliser:utils.correct_metadatahad the rule for frames,multi_databasehad it for scalars and for database reads, and the pair had to be pinned against each other by test precisely because they were two. Both now call this.spacr.multi_databasere-exports the name, so no caller moved.- Parameters:
plate – a plate id, from anywhere.
- Returns:
the id with a doubled
pprefix collapsed; everything else unchanged.
- spacr.schema.canonical_rename_plan(columns, requested=None)[source]¶
{old: canonical}for the columns that can safely be renamed.The one definition of the “target already exists, keep both” rule for frames, shared by
canonicalise_columns()andutils.canonicalize_measurement_columnsso the two cannot drift apart again — they have already disagreed once about which spellings they fix.The test folds case, because SQLite compares identifiers case-insensitively and these frames are written with
to_sql. A frame holdingrowandrowidalready has the canonical column: renamingrowtorowIDproduced['rowID', 'rowid'], which looks fine in pandas and makesto_sqlraiseduplicate column name: rowid. This is the same rule, and the same reasoning, as the database-level rename inspacr.database_schema— see the comment there about a plate whose database could not be opened at all.A column is excluded from its own comparison, so a pure respelling (
rowid->rowID, one column, one identifier as far as SQLite is concerned) is still made rather than being read as a collision with itself.- Parameters:
columns – the frame’s column names, in order.
requested – optional explicit
{source: canonical_target}choices supplied by the metadata resolver. They pass through the same case-folded collision guard as built-in aliases.
- Returns:
dictmapping each renameable name to its canonical form; empty when there is nothing to do.
- spacr.schema.canonicalise_columns(df)[source]¶
Return
dfwith every legacy column name renamed canonically.Metadata aliases (
LEGACY_COLUMN_NAMES) and the legacy feature spellings (LEGACY_COLUMN_PATTERNS) — this used to fix only the former whileutils.canonicalize_measurement_columns, the other frame-level canonicaliser, fixed both. Both now call onecanonical_column_name(), so a frame gets the same columns whichever of the two a caller reached for.A rename is skipped when the canonical name is already present, which is the same rule
utils.rename_columns_in_dbfollows: a frame carrying both spellings keeps both untouched rather than having one silently overwrite the other. Dropping data to tidy a name is never the right trade — a human can decide which column is authoritative, and until then both stay reachable.“Already present” is decided case-insensitively by
canonical_rename_plan(), because these frames get written withto_sqland SQLite compares identifiers case-insensitively.- Parameters:
df –
pandas.DataFramewhose columns may use legacy names.- Returns:
a new frame with canonical column names.
- spacr.schema.canonicalise_frame(frame, *, report=None, warn=None, repair_plate_ids: bool = True)[source]¶
The whole vocabulary applied to one frame: the reader’s normaliser.
Three steps, in this order and for this reason:
resolve_metadata_collisions()– collapse duplicate opinions before renaming, because renaming first is what produces two columns calledrowIDand ato_sqlthat refuses the table.canonical_rename_plan()– the surviving legacy spellings, plus the legacy feature spellings, renamed.normalise_plate_columns()– thepplate1value repair, on every column that carries a plate.
Every collision is recorded on
frame.attrs['column_collisions']as well as reported, so a GUI can show what a read decided without having intercepted the callbacks.- Parameters:
frame –
pandas.DataFrame.report – see
resolve_metadata_collisions().warn – see
resolve_metadata_collisions().repair_plate_ids – apply step 3. Off for a caller that must see the stored plate id exactly as written.
- Returns:
a new frame.
- spacr.schema.coerce_model_feature_types(frame, *, extra_features=(), exclude=(), allow_unknown: bool = False)[source]¶
Repair the two representations of a numeric measurement pandas fumbles.
A column with no values at all comes back as
object. This is the common one and it is not a data problem:pandas.read_sqlbuilds the frame from the rows it gets, so a column that isNULLin every row arrives as a column ofNoneand pandas types thatobject– even though SQLite declares itREAL. spaCR writesNULLfor an honest NaN, and whole measurements are legitimately NaN for a whole database:skew_intensity/kurtosis_intensityare NaN for every uniform object, andmode_intensitywas NaN for every object in every database written before the SciPy shim inmeasure._extended_regionprops_table. Such a column is converted tofloat64NaN, which is what it always meant. Nothing is lost, so nothing is warned about – and the caller’s own all-NaN filter then drops it and says so.Numeric text is coerced, loudly. Values inserted as text make pandas use
objectfor the whole column even when every one of them is a valid number;'12.0'is recoverable and is recovered, with aUserWarningnaming the columns, because a measurement stored as text is a database that wants looking at.'n/a'is not recoverable and is never quietly turned into NaN and fitted on: it raisesModelFeatureSchemaErrornaming every offending column at once.The input frame is returned unchanged when no conversion is needed. A shallow copy is made on the first conversion, so callers do not have their source data mutated and wide database joins do not get copied needlessly.
- Parameters:
frame – the measurements frame. Anything else — a Series included — raises
ModelFeatureSchemaError, notTypeError.extra_features – names to treat as declared features whatever the feature dictionary makes of them. The only way to get an unrecognised text column repaired, and it opts that column into the error above too.
exclude – names to leave alone entirely — never repaired, never reported. Tested first, so it overrides
extra_features. It is iterated, so a bare string excludes its letters and hence nothing.allow_unknown – widen what counts as a feature to unrecognised columns — but an unrecognised column is then skipped rather than repaired, so this only ever converts fewer columns. It reaches
DERIVED_MODEL_FEATURESas well: with it set, a text or all-NULLrecruitmentstaysobjectand is then silently dropped bymodel_feature_columns()instead of being read as numbers.
- spacr.schema.column_id(column: Any, *, strict: bool = False) str[source]¶
Return the canonical
'c<N>'column id.- Parameters:
column – column index, or an already-prefixed
'c<N>'.strict – raise instead of preserving an unparseable token.
- Returns:
'c<N>'.
- spacr.schema.column_index(value: Any) int | None[source]¶
'c12'→12; an unparseable id →None.- Parameters:
value – prefixed column id or numeric column token.
- spacr.schema.comparable_key_value(value: Any) str[source]¶
One metadata value, reduced to the form two spellings are compared in.
1,1.0,'01'and' 1 'are the same well. A dtype difference is not a disagreement, and this is the whole reason the comparison is notSeries.equals: a naive equality warns on every file that stored one copy of the well as text and the other as a number, and a warning that fires every time teaches the user to ignore the one that matters.Missing is its own value:
None,NaNand''all reduce to'', so two columns that are both blank on a row agree there.- Parameters:
value – a single cell.
- Returns:
the comparison string.
- spacr.schema.comparable_key_values(values) Tuple[str, ...][source]¶
comparable_key_value()over a column.- Parameters:
values – iterable of metadata cells to reduce for comparison.
- spacr.schema.compose_prc(plate: Any, row: Any, column: Any) str[source]¶
Return the
prcwell key:'plate1_r1_c1'.- Parameters:
plate – plate id.
row – row index or
'r<N>'.column – column index or
'c<N>'.
- Returns:
the composed key.
- spacr.schema.compose_prc_column(df, columns=None)[source]¶
Compose escaped plate-row-column identifiers for a frame.
- Parameters:
df (pandas.DataFrame) – Frame containing the well-key columns.
columns (sequence of str, optional) – Plate, row, and column field names. Defaults to
WELL_KEY_COLUMNS.
- Returns:
pandas.Series – Canonical
prcidentifiers aligned todf.- Raises:
KeyError – If any required key column is absent.
Notes
Plate values are escaped with the same rules as
compose_prc(), so separators and percent characters cannot create ambiguous keys. Legacy unescaped identifiers remain readable throughparse_prcf().
- spacr.schema.compose_prcf(plate: Any, row: Any, column: Any, field: Any, time: Any = None) str[source]¶
Return the
prcffield key.'plate1_r1_c1_f2', or'plate1_r1_c1_f2_t3'whentimeis given. The timepoint goes after the field — that is the order_map_wells(timelapse=True)writes and every table on disk carries.- Parameters:
plate – plate id.
row – row index or
'r<N>'.column – column index or
'c<N>'.field – field token.
time – timepoint token, or
Noneoutside a timelapse.
- Returns:
the composed key.
- spacr.schema.compose_prcfo(plate: Any, row: Any, column: Any, field: Any, obj: Any, time: Any = None, object_type: Any = None) str[source]¶
Return the
prcfoobject key:'plate1_r1_c1_f2_o7'.With a timepoint the object still goes last:
'plate1_r1_c1_f2_t3_o7'. That matches bothutils._map_wells_png(timelapse=True)and theprcf + '_' + 'o' + object_labelcomposition inio._read_and_join_tables.With an
object_typethe object component carries it:'plate1_r1_c1_f2_nucleus7'. The untyped form is unchanged, which is why everyprcfoalready on disk still composes and parses byte for byte as it did — the type is a refinement of the key, not a new spelling of it.- Parameters:
plate – plate id.
row – row index or
'r<N>'.column – column index or
'c<N>'.field – field token.
obj – object label, bare,
'o<N>'or'<type><N>'.time – timepoint token, or
None.object_type – the object table this object came from, or
Nonefor “not stated”.
- Returns:
the composed key.
- spacr.schema.correct_metadata_column_names(df)[source]¶
Rename legacy metadata columns to the canonical spaCR names.
A thin name over
spacr.schema.canonicalise_frame(), which is the one vocabulary: case- and punctuation-insensitive, soPlate,PLATE,plate_idandplateNameall becomeplateID. This function used to carry its own list of six spellings, matched case-sensitively, and that list is why a CSV whose header saidColumnreached a fit with nocolumnIDin it.Two things it does that the shared vocabulary deliberately does not:
grna_name->grna. Not in the shared vocabulary because the sequencing CSVs are still read withgrna_nameby name inspacr.submodulesandspacr.plot; renaming it at the database migration would break those reads, and a rename nobody can see is worse than a spelling.plate_rowsplit intoplateIDandrowID. That is a value split rather than a rename, so it has no place in a name table.
- Parameters:
df – DataFrame whose columns may use legacy names.
- Returns:
A DataFrame carrying the canonical names. The renames are not applied in place, so the caller must use the returned frame.
- spacr.schema.escape_field_stem_plate(name: Any, *, timelapse: bool = False) str[source]¶
Escape the plate component of a merged-stack field stem.
- Parameters:
name (Any) – File name, path, or stem in
plate_well_field[_time]form.timelapse (bool, default=False) – Treat the final component as a timepoint. A numeric trailing timepoint is also recognized in non-time-lapse merged-stack names.
- Returns:
str – The stem with only its plate component filename-escaped.
- Raises:
KeyParseError – If the stem does not contain a plate and the required tail fields.
Notes
Escaping at write time preserves plate names that contain underscores. The tail rules match
parse_field_stem().
- spacr.schema.escape_filename_component(token: Any) str[source]¶
Escape one free-text component for a separator-delimited filename.
This uses the same reversible table as join keys. In particular,
'my_plate'becomes'my%5Fplate'and a literal percent is escaped first, so parsing cannot merge it with an encoded separator.
- spacr.schema.field_id(field: Any, *, strict: bool = False) str[source]¶
Return the canonical
'f<N>'field id.'3','003','s3','F003'and3all give'f3'. A token holding no integer is preserved ('xy'→'fxy') rather than becoming'f0'; see the module docstring for why.- Parameters:
field – field token.
strict – raise instead of preserving an unparseable token.
- Returns:
'f<N>', or'f<token>'for an unparseable token.
- spacr.schema.field_index(value: Any) int | None[source]¶
'f2'→2;'fxy'→None.- Parameters:
value – prefixed field id or numeric field token.
- spacr.schema.fold_column_name(name: Any) str[source]¶
The form a column name is looked up by: lower case, no punctuation.
'Plate_ID','plate id','plate.id'and'plateID'all fold to'plateid'. Folding is what keepsLEGACY_COLUMN_NAMESa short list of words rather than a combinatorial table of every separator a plate reader has ever emitted. The supported aliases includePlate,PLATE,plate,plateid,plate_id,plate_nameandplateNameand this is five entries, not seven.- Parameters:
name – a column name.
- Returns:
the folded form.
- spacr.schema.is_object_type(object_type: Any) bool[source]¶
Whether
object_typecan be written into an object key.The question a reader asks about the table it just loaded before stamping
OBJECT_TYPE_KEYon the frame.png_list, a summary table or a user’s own table answer False, and the frame stays untyped — which is the key spaCR has always written, so nothing regresses.- Parameters:
object_type – candidate table or object-type name.
- spacr.schema.is_positional_pair(row: Any, column: Any) bool[source]¶
True when
(rowID, columnID)came from the positional passthrough.parse_well()puts an unrecognisable well into both slots verbatim, so an unprefixed pair of equal values is that passthrough and not a real row and column. Without this check('12', '12')looks like row 12 / column 12 andwell_id()happily renders it'L12'— a well name for a well that was never identified.Only strings can be a passthrough. A bare
intis unambiguously an index —well_id(1, 1)is a caller asking for well A01, not a well that failed to parse — so an integer pair is never flagged, however equal. (This is not hypothetical: it is the bug the round-trip test intests/test_schema.pycaught in the first version of this function, wherewell_id(1, 1)raised.)- Parameters:
row – the
rowIDas stored.column – the
columnIDas stored.
- Returns:
whether the pair is a passthrough rather than a position.
- spacr.schema.is_positional_well(well: Any) bool[source]¶
True when
wellis a bare number rather than<letters><digits>.Some acquisitions name wells
'12'. There is no way to know whether that means row 1 column 2 or the twelfth well, soparse_well()passes it through into both slots unchanged — which is what all five existing implementations do, and there is data on disk keyed that way. This predicate lets a caller detect the case instead of discovering it from arowIDthat does not start withr.- Parameters:
well – well identifier.
- Returns:
True when the well holds no row letters.
- spacr.schema.is_provenance_column(name: Any) bool[source]¶
Return whether
nameis identity, annotation, or run provenance.original_*columns, the names a plate had before conversion, count as provenance and never as measured features.
- spacr.schema.is_row_column_pair(row: Any, column: Any) bool[source]¶
True when
(row, column)is recognisably a well’s row and column.Deliberately narrow. It is the guard that stops a right-to-left key parse from absorbing a deeper key into an underscored plate id, so it must reject a
(columnID, fieldID)pair and a(fieldID, objectID)pair: acolumnIDis never'f1'and arowIDis never'c1'.parse_prcf()andml._split_prcare both right-to-left parses of a separator-joined key whose leftmost component may itself contain the separator, and both need exactly this test to tell “the plate is calledexp1_plate1” from “you handed me a key one level too deep”. It lives here so there is one answer to “is this a row and a column?”.- Parameters:
row – candidate
rowIDtoken.column – candidate
columnIDtoken.
- Returns:
whether the pair can be a row and a column.
- spacr.schema.is_within_plate_format(row: Any, column: Any, n_wells: int) bool[source]¶
True when
(row, column)lies inside ann_wellsplate.- Parameters:
row – row index or
'r<N>'.column – column index or
'c<N>'.n_wells – a key of
PLATE_FORMATS.
- Returns:
whether the position exists on that plate.
- Raises:
KeyParseError – when
n_wellsis not a standard format.
- spacr.schema.legacy_map_wells(file_name: Any, timelapse: bool = False) Tuple[str, ...][source]¶
utils._map_wellsreproduced bit for bit,'error'tuple and all.Used by
tests/test_schema.pyto assert that the canonical parser agrees with the legacy one on every well that works today, so the migration is provably a strict repair rather than a change of contract.- Parameters:
file_name – stack file name.
timelapse – parse a trailing timepoint.
- Returns:
the same tuple
_map_wellsreturns.
- spacr.schema.legacy_safe_int_convert(value: Any, default: Any = 0) Any[source]¶
utils._safe_int_convertexactly, including the0default.Kept only for the migration tests. New code calls
parse_int_token(), which returnsNone.- Parameters:
value – token to convert.
default – what to return on
ValueError. Note thatTypeError— which is whatNoneraises — is not caught, here or in the original.
- Returns:
the int, or
default.
- spacr.schema.legacy_well_ids(well: Any) Tuple[str, str][source]¶
utils._map_wells’ well handling exactly, raising where it raises.- Parameters:
well – well identifier.
- Returns:
(rowID, columnID).- Raises:
ValueError – on the wells
_map_wellsswallows into'error'.IndexError – on an empty well, as
_map_wellsdoes.
- spacr.schema.letters_from_row_index(index: int) str[source]¶
Inverse of
row_index_from_letters().27→'AA'.- Parameters:
index – 1-based row index.
- Returns:
the row letters.
- Raises:
KeyParseError – when
indexis not a positive integer.
- spacr.schema.model_feature_columns(frame, *, extra_features=(), exclude=(), allow_unknown: bool = False) list[str][source]¶
Select numeric model inputs by schema role, never by dtype alone.
A column is eligible when the feature dictionary identifies a measurement, when it belongs to an object-table measurement namespace, or when the caller explicitly names it in
extra_features. Identity and provenance are always excluded—even if SQLite/pandas represents them as numbers.allow_unknownis for generic statistics over user-created frames; database-backed model paths should retain the strict default.Every unusable column is reported in one error, with its dtype and why it is unusable. Refusing the first one and stopping made a user fix them one whole run at a time.
- Parameters:
frame – the frame to select from; anything else (a Series included) raises
ModelFeatureSchemaErrorrather thanTypeError.extra_features – names to declare as features whatever the feature dictionary makes of them. It cannot promote an identity or provenance column — those are dropped before it is consulted — but it does turn a non-numeric column from a silent omission into the error below.
exclude – names dropped before any check, so it overrides
extra_featuresand is the escape hatch the error message points at. Iterated, so passing one bare column name excludes its letters only.allow_unknown – also accept unrecognised columns, but only those already of a numeric dtype: an unrecognised non-numeric one is skipped instead of reported. It reaches
DERIVED_MODEL_FEATURESas well, so arecruitmentcolumn read back as text vanishes from the selection rather than raising.
- Raises:
ModelFeatureSchemaError – if a declared feature is non-numeric.
- spacr.schema.model_feature_frame(frame, **kwargs)[source]¶
Return
framerestricted tomodel_feature_columns().Unlike
model_feature_columns()this owns the data it hands back, so it repairs what is losslessly repairable first (coerce_model_feature_types()) instead of refusing a frame whose only fault is that pandas typed an all-NULL measurementobject.- Parameters:
frame – pandas frame to coerce and restrict to model features.
- spacr.schema.normalise_plate_columns(frame)[source]¶
Collapse a doubled
pprefix in every column that carries a plate.Applied on READ, so nothing on disk is rewritten and an old database keeps working. Modifies
framein place and returns it, which is what the two callers that predate this function both did.- Parameters:
frame – any frame read from a database or a CSV.
- Returns:
the same frame.
- spacr.schema.object_id(label: Any, *, object_type: Any = None, strict: bool = False) str[source]¶
Return the canonical object id used in
prcfo.'o<N>'when the object’s type is not stated, and'<type><N>'when it is —object_id(7, object_type='nucleus')is'nucleus7'. The type goes into the key, not beside it, because the key is the only thing that travels: a lasso publishes strings, and a string that cannot say which of a cell’s four children it means is a string that opens the wrong crop.An already-composed id round-trips, with or without a type (
object_id('nucleus7') == 'nucleus7'), so this is idempotent and safe to apply to a value that has already been through it.- Parameters:
label – object label — bare,
'o<N>', or'<type><N>'.object_type – the object table this object came from, or
Nonefor “not stated”. A type onlabelthat disagrees with this is an error rather than a silent overwrite.strict – raise instead of preserving an unparseable token.
- Returns:
the object id.
- Raises:
KeyParseError – on an empty token, a conflicting type, any bad token when
strict, or a composition that would not read back.
- spacr.schema.object_index(value: Any) int | None[source]¶
'o41'and'nucleus41'→41; an unparseable id →None.Split through
split_object_id()rather than by stripping'o', so a typed id reads back as the number it is instead of asNone.- Parameters:
value – typed, untyped, or bare object-label token.
- spacr.schema.object_table_schema(table: str) ObjectTableSchema[source]¶
Return the canonical schema for
table.- Parameters:
table – canonical object measurement table name.
- Raises:
ObjectTableSchemaError – when no canonical contract exists.
- spacr.schema.object_type_prefix(object_type: Any) str[source]¶
Return the canonical prefix an object of
object_typeis keyed with.None— the type is not stated — givesOBJECT_PREFIX. Anything else is folded to lower case and checked, because the prefix has to be separable from the label that follows it with no separator in between:it may not be empty, or every typed key would be an untyped one;
it may not contain
KEY_SEPARATOR, for the reason_check_plate()gives — the key is separator-joined;it may not contain a digit, or the split is ambiguous. Type
'cell1'with label7and type'cell'with label17both write'cell17', which is the “two identities, one key” failure this whole module exists to end.
The vocabulary is closed:
OBJECT_TYPESand nothing else. An open one would mean every unrecognised token in the object slot became a type, and'plate1_r1_c1_f2_x7'— which is not an object key — would parse as object 7 of type'x'. Widening what counts as a key is how a malformed key becomes a plausible wrong answer instead of an error, and a closed vocabulary is the same choiceKEY_PREFIXESalready makes for rows, columns, fields and timepoints.A caller holding a table name it is not sure about asks
is_object_type()first and leaves the frame untyped otherwise — an untyped key is exactly what spaCR wrote before, so that is a no-change, not a failure.- Parameters:
object_type – one of
OBJECT_TYPES, orNone.- Returns:
the prefix, lower case.
- Raises:
KeyParseError – for a type that cannot be written into a key.
- spacr.schema.object_type_summary(roles: Sequence[str]) str[source]¶
Describe object types once, collapsing organelle slots into a regex.
- Parameters:
roles – internal object identifiers, in display order.
- Returns:
comma-separated types. Organelle identifiers are represented by one exact pattern, such as
organelle(?:[b-z]|[a-z]{2})?for all slots. This changes presentation only, never stored identifiers or parsing.
- spacr.schema.parse_field_stem(name: Any, *, timelapse: bool = False, strict: bool = False) FieldID[source]¶
Parse a merged-stack file name into a
FieldID.- Parameters:
- Returns:
FieldID – Parsed plate, well, field, and optional timepoint identity.
- Raises:
KeyParseError – If the stem has the wrong number of components.
WellParseError – If the well component cannot be parsed.
Notes
A non-time-lapse call accepts one extra numeric timepoint emitted by the merged-stack writer, but omits it from the returned identity. Other extra components are rejected.
Examples
>>> parse_field_stem('plate1_A01_3').prcf 'plate1_r1_c1_f3'
- spacr.schema.parse_int_token(token: Any, *, allow_prefix: bool = True) int | None[source]¶
Return the integer
tokendenotes, orNone— never0.This is the replacement for
utils._safe_int_convert, and the whole point of it is the return type._safe_int_convertanswers “what number is this?” with0when the honest answer is “there isn’t one”, and0is a perfectly good field id, so the lie is unrecoverable downstream.Noneis not a field id, so every caller is forced to decide what to do — and the callers here do decide, seefield_id().Vendor prefixes are understood, because they are a spelling of a number rather than a different number: an ImageXpress site
s3, a CellVoyager fieldF003and a bare3are the same field, and a pipeline that gave them three different ids would be just as wrong as one that gave them allf0.- Parameters:
token – anything — a string, an int, a float,
None.allow_prefix – strip one or two leading ASCII letters when what follows is all digits. Default
True.
- Returns:
the integer, or
Nonewhen the token holds no integer.
Example
>>> parse_int_token('003'), parse_int_token('s3') (3, 3) >>> parse_int_token('T0001'), parse_int_token('x') (1, None) >>> parse_int_token('') is None, parse_int_token(None) is None (True, True)
- spacr.schema.parse_object_stem(name: Any, *, timelapse: bool = False, strict: bool = False) ObjectID[source]¶
Parse a crop-PNG file name into an
ObjectID.- Parameters:
- Returns:
ObjectID – Parsed field identity with the final object label attached.
- Raises:
KeyParseError – If the stem has too few components.
WellParseError – If the well component cannot be parsed.
- spacr.schema.parse_prcf(text: Any) FieldID[source]¶
Parse a
prcfstring back into aFieldID.Parsed right to left, which is what makes it correct: the components are optional in the middle (
timeIDmay or may not be there), andml.pysplitsprcfoleft to right into a fixed five columns, so a timelapse key with six parts silently misaligns every column.Extra components are not automatically an underscored plate. A key with more components than
plate_row_column_field[_time]is one of two things, and they mean opposite things:'exp1_plate1_r2_c12_f1'— a plate id containing the separator. The right-to-left rule handles it, and that is the case the absorption exists for.'plate1_r1_c1_f1_f2'— a key one level too deep, or a key whose components are not what they claim. Absorbing it would returnplateID='plate1_r1',rowID='c1',columnID='f1'— half the well inside the plate and a field id in the column slot — and every per-well figure grouped on that is a plausible wrong number with nothing anywhere saying so.
The two are told apart with
is_row_column_pair(), which is the same guardml._split_prcapplies for the same reason. Anything else with extra components is refused rather than guessed at.- Parameters:
text – e.g.
'plate1_r1_c1_f2'or'plate1_r1_c1_f2_t3'.- Returns:
the
FieldID.- Raises:
KeyParseError – when the string is not a
prcf.
- spacr.schema.parse_prcfo(text: Any) ObjectID[source]¶
Parse a
prcfostring back into anObjectID.The object prefix is stripped before the label is re-canonicalised, which is what makes this the inverse of
compose_prcfo()for every label rather than only the numeric ones.object_id()reads an already-prefixed numeric id back out of its prefix ('o7'→7→'o7'), but it cannot do that for a preserved non-numeric token, so handing it'oxy'used to yield'ooxy': the key'p_r1_c1_f1_oxy'parsed to'p_r1_c1_f1_ooxy', and parsing that yielded'ooooxy'. A key that grows every time it passes through the parser joins to nothing.A typed object component is read back with its type:
'plate1_r1_c1_f2_nucleus7'givesobjectType == 'nucleus'. An untyped one givesNone, which is what every key written before object types existed means and is not the same fact as “cell”.- Parameters:
text – e.g.
'plate1_r1_c1_f2_o7','plate1_r1_c1_f2_t3_o7'or'plate1_r1_c1_f2_nucleus7'.- Returns:
the
ObjectID.- Raises:
KeyParseError – when the string is not a
prcfo.
- spacr.schema.parse_well(well: Any, *, strict: bool = False) Tuple[str, str][source]¶
Return
(rowID, columnID)for a well identifier.'A01','a1','A-01'and' A01 'all give('r1', 'c1').'AA01'— a real 1536-plate well — gives('r27', 'c1'), whereutils._map_wellsraises into'error'andutils._map_wells_pngreturns('r1', 'c0').A well with letters but no digits (
'A') has no column. Under_map_wells_pngit became'c0', i.e. indistinguishable from a genuine column 0; here it raises, because a well with no column is not a well.A bare number is passed through into both slots — see
is_positional_well().- Parameters:
well – well identifier of any of the above shapes.
strict – also reject the bare-number passthrough.
- Returns:
(rowID, columnID).- Raises:
WellParseError – when the well is empty, has no column, or is a bare number and
strictis set.
Example
>>> parse_well('A01'), parse_well('aa1') (('r1', 'c1'), ('r27', 'c1'))
- spacr.schema.plate_format_for(row: Any, column: Any) int | None[source]¶
Return the smallest standard plate format containing
(row, column).A column past 24 is not an error — a 1536-well plate has 48 of them — so nothing in this module rejects one. This is how a caller that does care checks.
- Parameters:
row – row index or
'r<N>'.column – column index or
'c<N>'.
- Returns:
the well count of the smallest format that contains the position, or
Nonewhen it fits no standard plate.
- spacr.schema.resolve_metadata_collisions(frame, *, report=None, warn=None)[source]¶
Collapse every group of columns that mean the same metadata key.
Several columns can normalise to one key – a file carrying
well,wellIDandwell_namehas three opinions about which well a row came from, and every join downstream is silently picking one of them. What happens is decided by whether they agree row by row (comparable_key_value()):they agree – keep one, drop the rest, and print. Nothing is wrong, so nothing warns.
they disagree – the same action, and a warning naming the columns, the choice, and how many rows differ. That count is the point: “they disagree” is not actionable, “3 of 40 000 rows disagree” is a typo and “40 000 of 40 000” is the wrong file.
Only
METADATA_KEYSare collapsed. Two feature columns that normalise alike keep both spellings, because a measurement is data and dropping one to tidy a name is not this function’s call to make.Duplicate column labels are handled positionally, so a frame that already carries two columns both literally named
rowID– whichpandasallows andto_sqlrefuses – is repaired rather than raising.- Parameters:
frame –
pandas.DataFrame.report – called with each agreeing collision’s message. Pass
printto show them;Noneis silent.warn – called with each disagreeing collision’s message.
Noneroutes towarnings.warn(); pass a callable to capture them.
- Returns:
(frame, collisions). The frame is new when anything changed andframeitself when nothing did.
- spacr.schema.row_id(row: Any, *, strict: bool = False) str[source]¶
Return the canonical
'r<N>'row id.Accepts an index (
1,'1'), an already-prefixed id ('r1') — which round-trips rather than becoming'rr1'— or row letters ('A','AA').- Parameters:
row – row index,
'r<N>', or row letters.strict – raise instead of preserving an unparseable token.
- Returns:
'r<N>'.- Raises:
KeyParseError – on an empty token, or any bad token when
strict.
- spacr.schema.row_index(value: Any) int | None[source]¶
'r3'→3;'C'→3; an unparseable id →None.- Parameters:
value – prefixed row id, row letters, or numeric row token.
- spacr.schema.row_index_from_letters(letters: Any) int | None[source]¶
'A'→ 1,'Z'→ 26,'AA'→ 27,'AF'→ 32.Bijective base 26. Multi-letter rows are not an edge case: a 1536-well plate has 32 rows and runs
A…Z,AA…AF. Bothutils._map_wells(which raises, becoming'error') andutils._map_wells_png(which yields'c0') get these wrong, in two different ways. This matchesplate_qc._alpha_to_indexexactly, so the QC module and the database agree.- Parameters:
letters – one or more ASCII letters, any case.
- Returns:
the 1-based row index, or
Nonewhenlettersis not purely alphabetic or is empty.
- spacr.schema.screen_id(screen: Any = None) str[source]¶
Return the canonical screen id, defaulting an absent one.
Free-form text, like the plate id, because it is the name a user gave an experiment. It is not prefixed and not parsed back apart: the whole point of
SCREEN_KEYis that it is a dimension you block on, facet by and colour with as it stands.Absence is the case that matters.
None,''and whitespace all mean “this project has one screen”, and they becomeDEFAULT_SCREENrather than raising — every project that exists today has no screen anywhere in its settings, and demanding one would stop all of them from opening. Contrast_check_plate(), which does raise: an empty plate is a broken key, but an empty screen is an ordinary single-screen run.An empty value is never left empty inside a frame either, because a blank screen groups with every other blank screen — which is exactly the silent pooling
spacr.multi_databaseexists to refuse.- Parameters:
screen – the screen label, or
None.- Returns:
the label, stripped, or
DEFAULT_SCREEN.
Example
>>> screen_id('tsg101'), screen_id(None) ('tsg101', 'screen1')
- spacr.schema.split_object_id(token: Any, *, require_prefix: bool = True) Tuple[str | None, str][source]¶
Split an object id into
(object type, label).The inverse of
object_id():split_object_id('nucleus7') -> ('nucleus', '7') split_object_id('o7') -> (None, '7') split_object_id('omulti') -> (None, 'multi') split_object_id('7') -> (None, '') # not an object id
A type of
Nonemeans not stated, which is what'o'has always meant and is exactly what every key written before object types existed carries. It is not “unknown and therefore probably a cell”.- Parameters:
token – the last component of a
prcfo.require_prefix – when False a bare label (
'7') is accepted and read as an untyped id. That is the shapespacr.selection.object_keys()writes, where the label is joined bare rather than throughOBJECT_PREFIX.
- Returns:
(type or None, label). An empty label meanstokenis not an object id at all; callers must check it rather than assuming the split succeeded.
- spacr.schema.strip_prefix(value: Any, prefix: str) str[source]¶
Remove one leading
prefixfromvalueif it is there.- Parameters:
value – the id, e.g.
'r12'.prefix – the single-letter prefix, e.g.
'r'.
- Returns:
the remainder, e.g.
'12'.
- spacr.schema.table_key_columns(table: str, *, timelapse: bool = False) Tuple[str, ...][source]¶
Return the columns that identify a row of
table.- Parameters:
table – table name.
timelapse – include
timeID.
- Returns:
the key columns, most significant first.
- Raises:
KeyParseError – when
tableis not one spaCR owns.
Example
>>> table_key_columns('cell') ('plateID', 'rowID', 'columnID', 'fieldID', 'object_label')
- spacr.schema.time_id(time: Any, *, strict: bool = False) str[source]¶
Return the canonical
't<N>'timepoint id.'T0003'→'t3'. Under the old_safe_int_converteveryT####token becamet0, which collapsed a whole timelapse onto one frame.- Parameters:
time – timepoint token.
strict – raise instead of preserving an unparseable token.
- Returns:
't<N>'.
- spacr.schema.time_index(value: Any) int | None[source]¶
't7'→7; an unparseable id →None.- Parameters:
value – prefixed timepoint id or numeric time token.
- spacr.schema.unescape_filename_component(token: Any) str[source]¶
Invert
escape_filename_component().
- spacr.schema.validate_object_table_frame(frame, table: str, *, timelapse: bool | None = None, metadata_column_map=None, metadata_well_column=None, metadata_pseudo_source=None, allow_pseudo_metadata: bool = False, metadata_prompt=None, metadata_cache_key=None, metadata_mapping_path=None)[source]¶
Validate an object-table frame against its canonical contract.
Validation is deliberately strict at the writer boundary and compatibility-preserving in shape:
required identity/provenance columns must exist and be non-null;
labels (and present parent links) must be positive integers;
prcfmust exactly match the component key columns;one write batch may contain at most one row per object key;
measurement stamps are all present or all absent;
features from another compartment are rejected, and this table’s own feature namespace must be numeric.
Extra columns are allowed because annotation columns are user-defined and historical databases contain extensions. Legacy metadata spellings are canonicalised on the returned copy before validation.
pandas is imported only when this function is called. Importing
spacr.schemaitself remains standard-library-only for CLI, multiprocessing, and resume preflight paths.- Parameters:
frame – pandas DataFrame to validate.
table – one of
CANONICAL_OBJECT_TABLES.timelapse – require/forbid
timeID;Noneinfers it.
- Returns:
canonical-column DataFrame copy.
- Raises:
ObjectTableSchemaError – on any contract violation.
- spacr.schema.well_id(row: Any, column: Any) str[source]¶
Return the canonical well name:
('r3', 'c7')→'C07'.The inverse of
parse_well()for wells that have one. Matchesplate_qc.well_id.- Parameters:
row – row index or
'r<N>'or row letters.column – column index or
'c<N>'.
- Returns:
the well name, zero padded to two digits.
- Raises:
KeyParseError – when either index is unusable, or the pair is a positional passthrough (see
is_positional_pair()).
Nested helpers¶
- _non_numeric_feature_error.diagnostic_dtype(dtype) str¶
Return stable user-facing text for a pandas feature dtype.
Pandas
StringDtypeis reported as the establishedobjectwording; other extension and NumPy dtypes retain their own names.spacr/schema.py:2174
- validate_object_table_frame._validate_positive_integer(column: str, *, nullable: bool = False)¶
Require positive integral values in one canonical key column.
When
nullableis true, nulls are ignored while every populated value is still checked; failures name the table, column, and examples.spacr/schema.py:2987