spacr.foreign

Workflow inputs and outputs

Import

Import external measurements with explicit object/column mappings, or choose Import Images, Format Converter or External Masks. Imported measurements are not a fresh Measure run. With matching source images and external integer masks, Import builds the merged project arrays used by Measure. Choose the route matching your files and inspect object identities before measuring; imported feature tables can be used directly when they already contain the required measurements.

Open: Home → Import.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.

  • Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.

  • External measurements — Third-party measurement CSV/database tables and their image, mask and object-identity mappings.

Outputs

  • Measured objects — measurements/measurements.db; object tables depend on the enabled cell, nucleus, pathogen and organelle masks. Relevant tables, depending on the route: cell, nucleus, pathogen, cytoplasm. Relevant columns, depending on the route: plateID, rowID, columnID, fieldID.

  • Object crops — data/**/*_png when save_png is enabled; png_list indexes saved crops. Supported workflows can instead stream crops from merged arrays and masks. Relevant tables, depending on the route: png_list. Relevant columns, depending on the route: png_path, prcfo.

  • Images and label masks — merged/*.npy in the project; channels and integer label planes share each field array.

  • Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.

After this module

  • Annotate: Imported tables require explicit column and object-identity mappings.

  • Measure: Import matching images and external integer masks to build merged project arrays, then open Measure on that project. Skip this step when compatible measurements have already been imported or computed. Do not append duplicate measurements to an existing imported table.

  • Mask: For image-only imports, use Import Images or Format Converter and point Mask at the formatted image project. External measurements alone are not segmentation input.

API reference.

Module tutorial.

Foreign-data importer — somebody else’s images, masks and measurements turned into a working spaCR project.

The problem

A collaborator sends a folder of TIFFs, a folder of label images, and a results.csv out of CellProfiler / Fiji / QuPath / their own script. None of it is shaped like a spaCR experiment, and none of their column names mean what spaCR’s mean. What this module produces is a real project root:

dst/
  images/                       Yokogawa-named TIFFs + conversion_map.csv
  stack/<plate>_<well>_<field>.npy
  masks/cell_mask_stack/<plate>_<well>_<field>.npy
  masks/nucleus_mask_stack/…
  merged/<plate>_<well>_<field>.npy
  measurements/measurements.db
  crops/<object>/…
  column_map.csv                the mapping that was actually applied

…which Mask, Measure, Annotate, Classify, the Plate Viewer and the Database Browser all read without knowing it was imported.

What is actually hard

Moving files is the easy half, and spacr.convert already does it: spacr.convert.scan() / plan() / convert() handle the image formats, the Yokogawa naming, the well assignment and the conversion_map.csv that maps every converted name back to the file it came from. This module reuses all of it rather than growing a second copy.

The hard half is mapping an arbitrary external schema onto spaCR’s, and the four ways that goes silently wrong:

  1. A guessed mapping applied without review. Their Area in µm² written into spaCR’s cell_area in px² is a number that is wrong by a factor of a few hundred and looks completely plausible. So infer_column_map() only ever proposes; the proposal is a file the user edits (save_column_map() / load_column_map()); and run_import() applies what was agreed, nothing else.

  2. Units. spaCR’s intensities are raw uncalibrated counts, and its geometry is px²/px for a 2-D run — but a 3-D run measures volumes, in µm³ when it was given voxel_size_z_um/voxel_size_xy_um, under the same column names; the row’s measurement_units says which, and spacr.feature_dict resolves it per table. A foreign table in µm needs a scale factor unless the target rows are µm too, and every mapping therefore carries unit_in / unit_out. When a conversion is declared but the pixel size is unknown, the value is not multiplied by 1.0 and pretended to be pixels: the column is redirected to the foreign_ prefix, recorded with its own unit and calibrated = 0, and named in the plan and the summary.

  3. Columns that could not be mapped. They are listed by name in the plan, in ImportResult.summary(), and in the foreign_columns table — and they are still imported, under the foreign_ prefix. Dropping a column the user cared about, quietly, is the failure mode this exists to prevent.

  4. Name collisions. A foreign column called cell_area that means something else would corrupt every downstream analysis with no error at all. Targets are checked against spacr.feature_dict.parse_column() and against spaCR’s reserved key columns; a collision is a Conflict that either refuses the import (on_conflict='refuse', the default) or renames the column (on_conflict='rename') — never an overwrite.

  5. A destination that is already somebody’s project. The same argument one level up: a table collision. See below.

Their table never replaces spaCR’s

Their measurements are always written to foreign_<object>. The canonical cell / nucleus / pathogen tables are spaCR’s, and run_import() writes one only when it is empty of anyone else’s work — either because the destination is a fresh project (an import-only project needs a cell table for the rest of spaCR to read, and gets a copy of the foreign rows) or because a previous run of this importer wrote it and it has not been added to since.

When the destination already holds measurements:

  • the same fields — their columns arrive beside spaCR’s, not on top of them. <object> is left byte-for-byte as it was, foreign_<object> holds their rows, and the <object>_with_foreign view joins the two on (prcf, object_label) — the same key spacr.utils._merge_and_save_to_database() writes.

  • different fields — there is nothing to reconcile and no shape in which two experiments belong in one database, so the import refuses, naming the tables, their row counts and the fields on each side, and writes nothing at all.

This is a merge onto the canonical keys rather than a substitution for them because the canonical tables carry things a foreign table cannot reproduce: every spaCR feature column, and cell_id, the link from a nucleus or pathogen back to the cell that contains it. Replacing the table dropped both, and a parent-child link is not recoverable from what is left. conversion_map gets the same treatment — merged on the output filename, never replaced — so a project’s provenance back to its own original files survives an import into it.

The copy in the canonical table is a convenience, and it is handed back

The copy an import-only project gets exists so the rest of spaCR has a cell table to read. It stops being useful the moment spaCR measures the same objects itself: measure_crop appends, so its rows would land in that table beside theirs and every per-well count would be the sum of two populations, with nothing marking the seam.

release_canonical_copy() hands the table back. It removes only rows it has matched, in SQL, against foreign_<object> — same value in every shared column, NULL in every unshared one — so what it deletes provably still exists somewhere; it un-claims the table in the same transaction, so no record goes on saying the importer owns a table whose rows it no longer has; and it builds the <object>_with_foreign view, so their numbers stay one query away. It runs from two places: before run_import(measure=True) measures, and from spacr.resume.supersede_imported_copies() when a measure resume is about to fill the table. Neither ever runs it half way — a table that is half released is worse than one that is not.

The rows are identified by that predicate and by nothing else — no rowid, no key lookup. Both alternatives were measured against a real mixed table and both destroyed spaCR’s measurements: an object table declares a column called rowID, which shadows SQLite’s own row identity, and the import’s row for an object and measure_crop’s row for the same object carry identical plateID/rowID/columnID/fieldID/object_label. What tells the two apart is the only thing that ever did — which columns each holds a value in. See _RELEASE_ALIAS.

Object identity is the join

Their measurement rows have to line up with the objects in their masks. The key is (field, object label): an image_key column that says which image a row belongs to, and a label_key column holding the integer label of the object in that image’s mask. Both are stated in the plan, and both are verified against the label images before anything is written — JoinReport counts the rows that resolve to no field, the rows whose label is in no mask, and the mask objects that no row measures. An import where 40% of the rows match nothing is broken, and it says so with the number rather than quietly inner-joining it away.

Typical use:

from spacr import foreign as fg

plan = fg.plan_import('/data/theirs/images',
                      {'cell': '/data/theirs/cell_masks'},
                      '/data/theirs/results.csv',
                      um_per_px=0.65)
print(fg.format_plan(plan))          # nothing has been written
fg.save_column_map(plan, '/data/column_map.csv')
# …the user edits that file…
plan = fg.plan_import(..., column_maps=fg.load_column_map('/data/column_map.csv'))
result = fg.run_import(plan, '/data/imported')
print(result.summary())

Classes

ColumnMap

One foreign column and what it becomes. The unit of review.

Conflict

A foreign column whose target would collide with something of spaCR's.

ImportPlan

Everything the import would do, before any of it is done.

ImportResult

What run_import() actually did.

JoinReport

How their measurement rows line up with the objects in their masks.

MaskMapping

One foreign label image and the field it belongs to.

PairingReport

Which image fields have which masks, and everything that did not pair.

ResolvedColumn

A ColumnMap after conflicts and units have been settled.

Functions

default_settings(→ Dict[str, Any])

Return the settings import_project() understands, with defaults.

format_plan(→ str)

Render an ImportPlan as the block a user reads before agreeing.

import_project(→ ImportResult)

Plan and run one foreign import from a settings dict.

infer_column_map(→ List[ColumnMap])

Propose a mapping for every column of a foreign measurement table.

is_spacr_name(→ bool)

True when name is a column spaCR itself writes.

load_column_map(→ List[ColumnMap])

Read a column-map CSV back.

plan_import(→ ImportPlan)

Work out the whole import and write nothing.

read_measurements(→ pandas.DataFrame)

Read a foreign measurement table into a DataFrame.

release_canonical_copy(→ int)

Hand <object> back to spaCR: remove the copy this importer put in it.

run_import(, crop_limit, progress, int, str], ...)

Execute a reviewed plan: build the project at dst.

save_column_map(→ pathlib.Path)

Write the reviewable column-map CSV.

Module Contents

class spacr.foreign.ColumnMap[source]

One foreign column and what it becomes. The unit of review.

A ColumnMap is a proposal until a human has looked at it. infer_column_map() writes them, save_column_map() puts them in a CSV the user edits, and run_import() applies exactly what comes back from load_column_map() — there is no path by which an inferred mapping reaches the database unreviewed.

Parameters:
  • source – source-table column name, preserved verbatim.

  • target – destination column in measurements.db; an empty value is reported as unmapped and routed under the foreign prefix rather than dropped.

  • transform – "identity", a pixel-size "length"/"area"/ "volume" conversion, or a literal factor such as "*0.65" or "/1000".

  • unit_in – declared source unit; required for non-identity, non-literal conversions.

  • unit_out – stored-value unit; scaling transforms default to spaCR’s corresponding pixel unit when omitted.

  • note – reviewer prose retained in the map file, import plan, and foreign_columns provenance table.

classmethod from_row(row: Mapping[str, Any]) → ColumnMap[source]

Build a mapping from one row of the column-map file.

Parameters:

row – column-map record keyed by the serialised field names.

resolve(um_per_px: float | None = None) → Tuple[float | None, str][source]

Return (factor, reason) for this mapping.

factor is None when the conversion cannot be performed — and that is the whole point of returning it separately from a number. reason is empty on success and names the missing piece otherwise, in words that go straight into the plan.

Parameters:

um_per_px – micrometres per pixel. None means unknown, and an unknown pixel size never silently becomes 1.0.

to_row() → Dict[str, str][source]

This mapping as one row of the column-map file.

property declares_conversion: bool[source]

True when the two units are different scales of one quantity.

The check that catches a half-filled row: units saying um^2 -> px^2 with transform='identity' is not a copy, it is a conversion somebody forgot to declare.

property is_literal: bool[source]

True when the transform carries its own number.

property literal_factor: float | None[source]

The explicit numeric factor of a '*k' / '/k' transform.

property normalised_unit_in: str[source]

unit_in, normalised.

property normalised_unit_out: str[source]

unit_out, normalised, defaulted to spaCR’s own unit.

property power: int[source]

How many powers of the pixel size this transform applies.

class spacr.foreign.Conflict[source]

A foreign column whose target would collide with something of spaCR’s.

Parameters:
  • kind – collision category: "reserved", "spacr_name", "duplicate_target", or the non-blocking "shadows_spacr" notice.

  • source – foreign source-column name involved in the collision.

  • target – requested destination-column name that conflicts.

  • detail – human-readable explanation of the collision.

  • blocking – whether the import must refuse the plan until the collision is resolved.

__str__() → str[source]

Return a compact description of the conflicting column.

Returns:

Conflict kind, source, target, and explanatory detail.

class spacr.foreign.ImportPlan[source]

Everything the import would do, before any of it is done.

Nothing here has touched the destination. format_plan() renders it, ok says whether run_import() will accept it, and the three lists that matter — unmapped, conflicts and warnings — are the ones a user has to read before agreeing.

Parameters:
  • images – the spacr.convert.ConversionPlan for their image files. Built by spacr.convert.plan(); this module adds no second naming scheme.

  • masks – PairingReport — which mask belongs to which field, and every file on either side that did not pair.

  • measurements – their table, as read.

  • column_maps – the reviewed mapping that will be applied.

  • unmapped – source columns with no mapping, by name. They are still imported, under FOREIGN_PREFIX.

  • conflicts – Conflict entries; a blocking one makes ok False.

  • warnings – non-blocking things the user must see — an uncalibrated column, a low join match rate, a lossy z handling.

  • resolved – the derived, executable form of column_maps.

  • join – JoinReport.

  • errors – blocking planning problems that make ok false.

  • notes – non-problem planning facts shown before the user confirms the import.

  • object_types – mask/object classes to import, in spaCR mask-plane order.

  • n_channels – common number of intensity channels in each imported image field.

  • mask_dims – zero-based merged-array mask-plane index keyed by object type.

  • um_per_px – image calibration in micrometres per pixel, or None when physical length and area conversions must remain uncalibrated.

  • prefix – namespace prepended to foreign target columns that do not use a reviewed spaCR name.

  • on_conflict – "refuse" to block colliding targets or "rename" to assign an unused prefixed name.

  • allow_spacr_targets – explicit opt-in allowing reviewed foreign columns to use names owned by spaCR.

  • sources – absolute source locations keyed by "images", "measurements", and "mask:<object_type>".

  • base_warnings – the warnings that do not come from the column mapping (unpaired masks, the join, z handling). Kept apart so with_column_maps() can rebuild the mapping’s own warnings without losing them or duplicating them.

  • base_errors – likewise for blocking problems.

  • proposed – true while the column mapping is inferred and has not yet been returned through with_column_maps() for review.

target_for(source: str) → str[source]

The column name source will actually be written under.

Parameters:

source – foreign source column to look up.

targets() → List[str][source]

Every column name the import will write, in order.

with_column_maps(column_maps: Sequence[ColumnMap], *, um_per_px: Any = '<keep>', on_conflict: str | None = None, allow_spacr_targets: bool | None = None) → ImportPlan[source]

Return this plan with a different column mapping applied.

Pure CPU — no folder is rescanned and no file reopened — which is what lets a GUI re-run the conflict and unit checks on every keystroke in the mapping table. The join, the pairing and the image plan are carried over untouched, because none of them depends on how the columns are named.

Parameters:
  • column_maps – the mapping to apply instead.

  • um_per_px – a new pixel size; omitted keeps the plan’s.

  • on_conflict – 'refuse' / 'rename'; omitted keeps.

  • allow_spacr_targets – omitted keeps.

Returns:

a new ImportPlan.

property blocking_conflicts: List[Conflict][source]

Conflicts that stop the import.

property ok: bool[source]

True when run_import() will accept this plan.

property stems: List[str][source]

Field stems that will be imported, sorted.

property uncalibrated: List[ResolvedColumn][source]

Columns whose values are not in the unit they were meant to be.

class spacr.foreign.ImportResult[source]

What run_import() actually did.

Parameters:
  • plan – import plan represented by this result.

  • dst – destination project directory.

  • conversion – the spacr.convert.ConversionResult for their images — the provenance back to the original filenames.

  • db_path – the measurements.db that was written.

  • column_map_path – path of the applied column mapping saved beside the imported project.

  • stacks – per-field intensity-stack .npy files written for the project.

  • mask_files – per-field label-mask .npy files written for the project.

  • merged – merged .npy paths, one per imported field.

  • rows – rows written into each foreign object table.

  • crops – PNG paths cut from the merged arrays, if any.

  • measured – True when spaCR’s own measurements were re-extracted.

  • ledger – RunLedger carrying per-item outcomes and overall completeness, or None when no ledger was produced.

  • warnings – non-fatal execution problems, including fields skipped after planning.

  • notes – things that happened and are not problems — chiefly a canonical object table that was already populated and was therefore left exactly as it was found.

summary() → str[source]

The end-of-run block: what landed, and everything that did not.

property is_complete: bool[source]

True when nothing was skipped and nothing failed.

property n_fields: int[source]

Fields that produced a merged array.

property n_rows: int[source]

Measurement rows written, across every object table.

class spacr.foreign.JoinReport[source]

How their measurement rows line up with the objects in their masks.

The join key is (field, object label). Both halves are verified against the label images at plan time, and every failure is counted: an import where 40% of the rows match no mask object is broken, and the number is the only thing that makes that visible before the database exists.

Variables:
  • image_key – their column naming the image a row belongs to, or '' when the whole table is one field.

  • label_key – their column holding the object’s integer label.

  • object_type – mask object class the measurement rows describe.

  • rows_total – rows in their table.

  • rows_matched – rows whose (field, label) exists in a mask.

  • unresolved_fields – (value, count) for image-key values that matched no converted image.

  • rows_no_object – (stem, count) for rows whose label is absent from that field’s mask.

  • objects_unmeasured – (stem, count) for mask objects that no row measures.

  • ambiguous_keys – image-key spellings that matched more than one field and were therefore not used.

  • examples – representative row-level failures shown after the counts.

summary() → str[source]

A multi-line rendering of every count above.

property key_description: str[source]

The join, in one sentence, for the plan and the summary.

property match_rate: float[source]

Fraction of their rows that found an object. 1.0 when empty.

property n_no_object: int[source]

Rows whose label is in no mask.

property n_objects_unmeasured: int[source]

Mask objects with no measurement row.

property n_unresolved: int[source]

Rows whose image key resolved to no field at all.

property rows_unmatched: int[source]

Rows that will carry no object — the honest headline number.

class spacr.foreign.MaskMapping[source]

One foreign label image and the field it belongs to.

Parameters:
  • source – Filesystem path of the foreign label-mask image.

  • object_type – Segmented object role from spacr.crops.MASK_PLANE_ORDER (cell, nucleus, pathogen, or one of the organelle slots).

  • stem – Canonical matched spaCR field stem, such as plate1_A01_3, used for imported per-field artifacts.

  • plate – Canonical plate identifier of the matched image field.

  • well – Canonical well identifier of the matched image field.

  • field – Integer field number of the matched image field.

  • source_field – Field token parsed from the original mask filename before matching.

  • match – "exact" when the source field token matched unchanged, or "normalised" when mask suffix stripping was required. This records filename matching independently of directory-based fallback.

  • labels – Sorted positive object labels read from the mask during verified planning; empty when the mask has no positive labels or label verification was disabled.

class spacr.foreign.PairingReport[source]

Which image fields have which masks, and everything that did not pair.

The rule is field-for-field: a mask with no image and an image with no mask are both reported per file. A converter that quietly kept the intersection would produce a smaller, perfectly consistent, wrong experiment.

Variables:
  • fields – {stem: {object_type: MaskMapping}} for the fields that have a full set of masks.

  • images_without_masks – (image path, object_type) per source image whose field has no mask of that type.

  • masks_without_images – (mask path, object_type) per mask that matched no image field.

  • unreadable_masks – (mask path, reason).

  • excluded – stems dropped because they lacked a required mask.

property ok: bool[source]

True when every image has every mask and every mask an image.

class spacr.foreign.ResolvedColumn[source]

A ColumnMap after conflicts and units have been settled.

This — not ColumnMap — is what run_import() executes, and what the foreign_columns table records. Keeping the two apart is what makes “the mapping you saved is the mapping that ran” a property that can be checked: the mapping is the input, the resolution is the derivation, and the derivation is deterministic.

Parameters:
  • mapping – reviewed source-to-target column mapping from which this resolution was derived.

  • target – column name actually written after conflict handling and any rename.

  • factor – multiplier applied to numeric values, or None when a requested conversion could not be performed.

  • calibrated – whether stored values are valid in the unit reported by unit.

  • unit – unit of the stored values, or an empty string when unknown.

  • status – resolution state: "mapped", "renamed", "uncalibrated", or "unmapped".

  • reason – explanation for a non-"mapped" resolution, otherwise an empty string.

apply(values: pandas.Series) → pandas.Series[source]

Return values with this resolution’s factor applied.

Parameters:

values – foreign-column values to copy or scale.

A non-numeric column is passed through untouched however the factor reads: multiplying a string column of treatment names by 0.65 is not a unit conversion, it is a crash.

to_record(table: str) → Dict[str, Any][source]

One row of the foreign_columns provenance table.

Parameters:

table – destination table that owns the resolved column.

property source: str[source]

The foreign column name.

spacr.foreign.default_settings(settings: Mapping[str, Any] | None = None) → Dict[str, Any][source]

Return the settings import_project() understands, with defaults.

Shaped like every other spacr settings factory — pass a partial dict, get it back filled in — so a CLI or a GUI can build a panel from it without special-casing this module.

spacr.foreign.format_plan(plan: ImportPlan) → str[source]

Render an ImportPlan as the block a user reads before agreeing.

Parameters:

plan – proposed import plan to render.

Ordered by what can hurt them: blocking problems, then conflicts, then the columns that could not be mapped, then the join, then the plain counts.

spacr.foreign.import_project(settings: Mapping[str, Any] | None = None, **overrides: Any) → ImportResult[source]

Plan and run one foreign import from a settings dict.

Always prints the plan before writing anything, so even a headless run leaves the mapping, the unmapped columns and the join counts in the log where a surprised user can find them.

column_map is the path to a reviewed save_column_map() file. Leaving it None runs with the inferred proposal, which the printed plan says in as many words — use preview_only first, save the map, read it, then run.

Returns:

the ImportResult; for preview_only an empty one carrying the plan.

Raises:

ConfigurationError – a missing input, or a plan with blocking problems.

spacr.foreign.infer_column_map(df: pandas.DataFrame, image_key: str | None = None, label_key: str | None = None, prefix: str = FOREIGN_PREFIX, skip: Iterable[str] | None = None) → List[ColumnMap][source]

Propose a mapping for every column of a foreign measurement table.

A proposal, never an application. Nothing in this module writes a database from the output of this function without it having passed through save_column_map() / load_column_map() or having been handed back explicitly — because the one mistake this module exists to prevent is an inferred Area -> cell_area that nobody read.

The proposal is deliberately conservative:

  • Nothing is proposed onto a spaCR feature name. Their Area is not spaCR’s cell_area: different segmentation, different definition, different unit. Every feature column is proposed as foreign_<name>. A user who genuinely wants their column in spaCR’s slot edits the target and passes allow_spacr_targets=True to plan_import(), which is a decision with a name on it.

  • A unit read out of the header becomes a declared conversion. Area (µm²) is proposed as transform='area', unit_in='um^2', unit_out='px^2' — which then needs a pixel size, and says so if there is none.

  • The join keys are left out: they are consumed by the join, not imported as measurements, and JoinReport states them.

Parameters:
  • df – their measurement table.

  • image_key – the column identifying the image; inferred from the headers when None.

  • label_key – the column holding the object label; likewise.

  • prefix – prefix for the proposed targets.

  • skip – extra columns to leave out of the proposal.

Returns:

one ColumnMap per remaining column, in table order.

spacr.foreign.is_spacr_name(name: str) → bool[source]

True when name is a column spaCR itself writes.

Parameters:

name – candidate column name to classify.

Delegated to spacr.feature_dict.parse_column() rather than a second parser: that module already implements the whole grammar (<object>_channel_<i>_<stat>, the radial-distribution and organelle-summary forms, the merge suffixes) and knows every metadata column. family == 'unknown' is exactly “spaCR would not write this”.

spacr.foreign.load_column_map(path: str) → List[ColumnMap][source]

Read a column-map CSV back.

Parameters:

path – a CSV written by save_column_map(), possibly edited.

Returns:

the mappings, in file order.

Raises:

ConfigurationError – the file is missing, is not a column map, or names the same source column twice — which would make “what was applied” ambiguous.

spacr.foreign.plan_import(images: str, masks: str | Mapping[str, str] | Sequence[Any] | None, measurements: str | pandas.DataFrame, *, layout: str = 'auto', z_handling: str = cv.Z_MAX, plate_naming: str = 'index', mask_layout: str | None = None, mask_suffixes: Sequence[str] | None = None, measurement_table: str | None = None, measurement_object: str | None = None, image_key: str | None = None, label_key: str | None = None, column_maps: Sequence[ColumnMap] | None = None, um_per_px: float | None = None, on_conflict: str = 'refuse', allow_spacr_targets: bool = False, prefix: str = FOREIGN_PREFIX, verify_labels: bool = True, metadata_type: str | None = None, custom_regex: str | None = None) → ImportPlan[source]

Work out the whole import and write nothing.

Five things happen here, in order, and every one of them can only produce a report:

  1. Images are scanned and planned by spacr.convert — the Yokogawa naming, the well assignment and the conversion_map.csv come from there, unchanged.

  2. Masks are scanned the same way and paired to image fields. A mask with no image and an image with no mask are both listed by path in ImportPlan.masks.

  3. Their table is read and, when column_maps is not given, infer_column_map() proposes one. A proposal is not an application: it is what you save, edit and hand back.

  4. Columns are resolved — collisions with spaCR names settled, unit conversions computed or refused.

  5. The join is verified against the label images: every count in JoinReport is real, not assumed.

Parameters:
  • images – folder of their image files.

  • masks – folder (taken as cell), or {object_type: folder}.

  • measurements – their table, or a path to it.

  • layout – source layout for spacr.convert.scan().

  • z_handling – spacr.convert.Z_MAX by default, because a merged array holds one plane per channel; keeping every z would produce fields spaCR cannot merge, and the plan says so.

  • mask_layout – layout for the mask folders; defaults to layout.

  • mask_suffixes – tokens stripped from mask filenames before matching (MASK_SUFFIXES by default).

  • measurement_object – which object type their table measures; defaults to the first mask type given.

  • image_key – their column naming the image. Inferred when None.

  • label_key – their column holding the object label. Inferred when None.

  • column_maps – the reviewed mapping. When None, an inferred proposal is used and the plan says loudly that it is a proposal.

  • um_per_px – micrometres per pixel. None means unknown, and any column needing it is reported uncalibrated rather than scaled by 1.

  • on_conflict – 'refuse' (default) or 'rename'.

  • allow_spacr_targets – opt in to writing foreign values into spaCR’s own column names.

  • verify_labels – read the masks to check the join. On by default.

  • metadata_type – the filename convention their images AND masks are named by – any of Mask’s metadata_type values. None or 'auto' reads plate / well / field from the folders, as before. See spacr.convert.scan().

  • custom_regex – the pattern for metadata_type='custom'.

Returns:

an ImportPlan.

Raises:

ConfigurationError – for an unreadable input or an unknown option — a setup mistake, not a per-item failure.

spacr.foreign.read_measurements(source: str | pandas.DataFrame, table: str | None = None) → pandas.DataFrame[source]

Read a foreign measurement table into a DataFrame.

CSV / TSV / Excel / Parquet / SQLite, or a DataFrame straight through.

Parameters:
  • source – path, or an already-loaded DataFrame.

  • table – table name, for a SQLite source. The only table is used when there is exactly one; otherwise the name is required, because picking one at random is how you import the wrong 40 000 rows.

Returns:

the table.

Raises:

ConfigurationError – unreadable, unknown extension, or an ambiguous SQLite source.

spacr.foreign.release_canonical_copy(db_path: str, object_type: str, dry_run: bool = False) → int[source]

Hand <object> back to spaCR: remove the copy this importer put in it.

run_import copies the imported frame into the canonical cell / nucleus / pathogen table when nothing of anyone else’s is there, so that a project built purely by import is readable by every spaCR tool. That copy is a convenience, and it stops being one the moment spaCR measures the same project itself: measure appends, so its rows would land in the same table beside theirs, and every per-well count downstream would be the sum of two populations with nothing marking the seam.

This removes them — losslessly, and it proves that rather than assuming it. Every row it deletes is checked, in SQL, against foreign_<object>: same value in every column the two tables share, NULL in every column they do not. A row that cannot be matched stops the whole call, because the copy would then be the only copy.

It also un-claims the table, which is the half a previous attempt left out: the foreign_columns rows for <object> are deleted and foreign_import.canonical_table_written is set to 0, so nothing in the database goes on saying the importer owns a table whose rows it no longer has. Without that the claim outlives the rows and no later run can tell the difference.

Their numbers stay exactly where they were, in foreign_<object>, and the <object>_with_foreign view is created so they are still one query away — joined to whatever spaCR measures next on (prcf, object_label).

Running it twice is a no-op: the second call finds nothing to release and returns 0.

Parameters:
  • db_path – path to measurements.db.

  • object_type – canonical table to release, e.g. 'cell'.

  • dry_run – run every check and report the count, writing nothing. This is how a caller asks “could this be released?” without a second implementation of the question that could answer differently from the one that does the work.

Returns:

number of imported rows removed from <object> (or, with dry_run, that would be).

Raises:

ConfigurationError – when the rows cannot be shown to survive the removal — foreign_<object> missing, or holding no twin for some row — or when the DELETE turns out to have removed a different number of rows than the checks that cleared it counted, which means the statement did not act on the rows that were verified. Nothing is written in any of those cases; the last one rolls the delete and the un-claim back together.

Example

from spacr.foreign import release_canonical_copy
release_canonical_copy('exp/measurements/measurements.db', 'cell')
# 4  -> `cell` is spaCR's again; theirs are in foreign_cell,
#       joined by the view cell_with_foreign
spacr.foreign.run_import(plan: ImportPlan, dst: str, *, overwrite: bool = False, measure: bool = False, measure_settings: Mapping[str, Any] | None = None, crops: bool = False, crop_channels: Sequence[int] | None = None, crop_size: Tuple[int, int] = (224, 224), crop_limit: int | None = None, progress: Callable[[int, int, str], None] | None = None, ledger: spacr.errors.RunLedger | None = None) → ImportResult[source]

Execute a reviewed plan: build the project at dst.

In order: convert their images (spacr.convert.convert()), write the intensity stacks, write the mask stacks, build the merged arrays with spaCR’s own spacr.io._load_and_concatenate_arrays(), load the conversion map into the database, write their measurements next to spaCR’s metadata, record the provenance, and — optionally, and separately — re-extract spaCR’s own measurements and cut crops.

Re-running is a no-op rather than a duplication: existing TIFFs are left alone by the converter, and every table this owns is replaced, never appended to.

Owning is the whole of it. Their rows go to foreign_<object>, and the canonical <object> table is written only when it does not exist or when a previous run of this importer wrote it and nothing has been added to it since. A <object> table holding somebody’s measurements is left exactly as it is and joined to theirs by the <object>_with_foreign view; a destination whose measurements are of a different experiment is refused before a single file is written. See the module docstring.

Parameters:
  • plan – a plan from plan_import() whose ok is True.

  • dst – destination project root. Created if missing. It may be an existing spaCR project of the same experiment, and nothing in it will be overwritten.

  • overwrite – rewrite converted images that already exist.

  • measure – also run spacr.measure.measure_crop() over the imported project. Off by default, and separate: when it is on, the standard object tables are spaCR’s and theirs stay in the foreign_* tables, joined by a <object>_with_foreign view. A convenience copy an earlier import left in <object> is released first (release_canonical_copy()) rather than appended to — that path used to record “the importer did not write this table” while leaving every one of its rows in it.

  • measure_settings – extra settings for measure_crop.

  • crops – cut one PNG per imported object.

  • progress – progress(done, total, message).

  • ledger – reuse an existing spacr.errors.RunLedger.

Returns:

an ImportResult.

Raises:

ConfigurationError – the plan is not ok, or dst already holds measurements of a different experiment.

spacr.foreign.save_column_map(plan_or_maps: ImportPlan | Sequence[ColumnMap], path: str) → pathlib.Path[source]

Write the reviewable column-map CSV.

Six columns — COLUMN_MAP_COLUMNS — and a few # comment lines above them explaining what the file is for. It opens in a spreadsheet, and the round trip through load_column_map() is exact.

Parameters:
Returns:

the written path.

Nested helpers

ColumnMap.from_row._get(key: str) → str

Return one serialised mapping field as stripped text.

Parameters:

key – Column-map field name to read from row.

Returns:

Stripped text, or an empty string for a missing, null, or NaN value.

spacr/foreign.py:631

_resolve_columns._foreign(source: str) → str

Build an unused prefixed target for one foreign column.

Parameters:

source – Foreign source-column name to sanitise.

Returns:

SQL-safe prefixed target not present in taken.

spacr/foreign.py:1548

_resolve_columns._note_shadow(source: str, target: str) → None

Record that their column is named like one of spaCR’s.

Not blocking — the value lands under the foreign prefix, so nothing is overwritten — but a user reading cell_area in the source table and cell_area in a spaCR database has to be told they are two different numbers.

spacr/foreign.py:1556

run_import._step(index: int, message: str) → None

Report one import step when a progress callback is available.

Parameters:
  • index – One-based completed-step position.

  • message – Human-readable description of the current step.

Returns:

None.

spacr/foreign.py:3237