spacr.convert¶
Workflow inputs and outputs¶
Format Converter¶
Convert supported microscope formats to a configured TIFF layout while preserving source mappings. Conversion does not generate object measurements.
Open: Import → Format Converter.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.
Outputs
Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.
After this module
Mask: Use the converted layout and preserve source identity mappings.
Format converter / importer — vendor microscopy files into Yokogawa TIFFs.
This is the standalone version of “get my images into a shape spaCR can
read”. It is deliberately decoupled from the mask pipeline: nothing
here imports torch, cellpose or spacr.core, so it can be run,
tested and previewed on a laptop with no GPU and no segmentation
settings in sight.
What it does¶
A folder tree like:
run1/
wt/
fov01_C1.tif fov01_C2.tif
fov02_C1.tif fov02_C2.tif
…
becomes one flat folder of Yokogawa-named TIFFs:
plate1_A01_T0001F001L01A01Z01C01.tif
plate1_A01_T0001F001L01A01Z01C02.tif
plate1_A01_T0001F002L01A01Z01C01.tif
…
run1 is the plate, wt is the well, each field-set gets its own
field id, each channel its own channel id. That name is exactly what
spacr.utils._get_regex('cellvoyager', 'tif') parses, so the output
folder drops straight into Mask/Measure with
metadata_type='cellvoyager'.
Why a second converter¶
spacr.io.convert_to_yokogawa() already converts ND2/CZI/LIF/TIFF,
and it is still the right tool when you want an in-place rename of a
flat folder. It cannot do the three things this module exists for:
Preview. It writes as it walks. There is no point at which you can look at “
run1/wt/fov01_C1.tif→plate1_A01_…C01.tif” and say no.plan()produces that table and writes nothing.A map you can read back. It writes a
rename_log.csvwhose columns differ per input format (the LIF and TIFF branches record two columns; the ND2 branch records seven; the CZI branch records nine), and nothing in spaCR ever reads it.write_map()emits one fixed schema andread_map()/populate_db_from_map()read it back, so after a run the original filename for any measurement is one SQL join away.Z that is not silently flattened.
convert_to_yokogawamax-projects unconditionally in three of its four branches (ND2, LIF, and both the 3-D and 4-D TIFF cases), and says so nowhere in its output. Here the default isZ_KEEP— every plane is written with its ownZ##— and choosing to project is an explicitz_handling='max', announced in the plan and recorded per row in the map file.
Design rules¶
No new required dependencies.
nd2reader/czifile/readlifare probed withimportlib.util.find_spec(); when one is missing the affected files are reported as skipped with the package name and thepip installline, never as an ImportError traceback.Never overwrite. Two sources landing on one target name is a plan-time error naming both, not a last-writer-wins at convert time. An output that already exists on disk is skipped and counted, so a re-run is a no-op rather than a silent rewrite.
Atomic writes. Every TIFF is written to a temp file in the destination and
os.replace()d into place, so an interrupted run leaves no half-written file that the next run mistakes for done.Ledgered. A conversion that skipped 12 unreadable files says so in one grouped block at the end and stamps
conversion_map.run_status.jsonnext to its output — seespacr.errors.
Typical use:
from spacr import convert as cv
sources = cv.scan('/data/run1')
plan = cv.plan(sources, z_handling=cv.Z_KEEP)
print(plan.to_frame()) # the preview — nothing written yet
if plan.ok:
result = cv.convert(plan, '/data/run1_yokogawa')
print(result.summary())
cv.populate_db_from_map('/data/run1_yokogawa/measurements.db',
result.map_path)
Classes¶
The preview: what would be written, and what would go wrong. |
|
What |
|
One source plane and the one output TIFF it becomes. |
|
One readable unit of input: a file, or one series inside a file. |
Functions¶
|
Map arbitrary well-folder names onto plate well ids. |
|
Execute a plan: write the TIFFs, the map file and the run stamp. |
|
Scan, preview, convert and map one folder in a single call. |
|
Return the settings |
|
Return the user-facing sentence for an unreadable format. |
|
Return |
|
Say why |
|
Turn scanned sources into the preview table. Writes nothing. |
|
Return the plate format that has to be used for one plate's wells. |
|
Load a map file into |
|
Read a map file back. |
|
Read a Bio-Rad Image Lab |
|
True when |
|
Return |
|
Walk |
|
Render a |
|
Build one Yokogawa filename. |
|
Return every well id of an |
|
Write the map file for |
Module Contents¶
- class spacr.convert.ConversionPlan[source]¶
The preview: what would be written, and what would go wrong.
Nothing in here has touched the disk.
plan.to_frame()is the table a user reads before deciding;plan.okis what a caller checks before handing it toconvert().- Variables:
mappings – one entry per output TIFF.
errors – blocking problems. A non-empty list means
convert()refuses — the only member today is a target-name collision, which is exactly the case where writing anyway would destroy data.warnings – non-blocking but load-bearing: z projection, files that cannot be read, an already-converted-looking source.
notes – neutral facts about the plan (counts, assumptions).
unreadable – sources
scan()could not open, carried through so the preview shows them and the ledger counts them.well_map –
{(plate key, well key): well id}— the record of how folder names became wells.plate_map –
{plate key: plate token}.sources – scanned source images used to build and later validate the plan before any conversion starts.
channel_map – source plate/channel identities mapped to their stable one-based output channel numbers.
z_handling – plan-wide z policy applied to every source stack.
- to_frame() pandas.DataFrame[source]¶
Return the preview table.
One row per output, plus one row per unreadable source with an empty
targetand the reason instatus— a preview that quietly omitted the files it could not read would be exactly the wrong shape of honest.
- class spacr.convert.ConversionResult[source]¶
What
convert()actually did.- Parameters:
plan – conversion plan supplied to the run or returned with a preview-only result.
dst – destination directory associated with the conversion.
written – mappings whose TIFF targets were created by this run.
existing – mappings whose valid targets were already present, including fields accepted from a resume checkpoint, and were left untouched.
failed – planned mappings left unwritten after their source or source series raised during conversion.
skipped –
(item, reason)pairs for unreadable sources and for attempted source or source-series groups that failed.ledger – run ledger recording conversion successes and failures, or
Nonefor a preview-only or manually constructed result.map_path – path of the written conversion-map CSV, or
""when no map was written.checkpoint_path – path of the atomic field-level checkpoint, or
""when no checkpoint is associated with the result.resumed_fields – field identifiers accepted from a compatible checkpoint after all recorded TIFF targets were revalidated.
- class spacr.convert.Mapping[source]¶
One source plane and the one output TIFF it becomes.
Every field needed to walk the arrow backwards is here, which is why the map file can be a straight dump of these.
- Parameters:
source – path of the source image as retained by the scan.
target – basename of the output TIFF.
plate – assigned and sanitized output plate token.
well – assigned output well identifier.
field – assigned one-based output field number.
channel – assigned one-based output channel number.
z – assigned one-based output z number.
t – assigned one-based output timepoint number.
source_plate – plate key inferred from the original input layout.
source_well – well key inferred from the original input layout.
source_field – field key parsed from the original source.
source_channel – channel key parsed or inferred from the original source.
source_z – provenance text for the selected source z plane, including the original filename token, a one-based plane ordinal, or
"max(1..N)"for a projection.source_t – provenance text for the selected source timepoint: the original filename token or a one-based plane ordinal.
z_handling – z-plane policy applied to this output:
"keep","max", or"first".n_z_planes – number of z planes in the source before any projection or first-plane reduction.
n_timepoints – number of timepoints recorded for the source.
plane – zero-based
(t, z, c)source-array index; a z index of-1requests projection over every z plane.meta – selected source provenance serialized into the conversion map, including series, extension, reader, relative path, axes, and axis assumptions when available.
- class spacr.convert.SourceImage[source]¶
One readable unit of input: a file, or one series inside a file.
A vendor file holding six scenes produces six
SourceImages, so “each field a unique field id” holds whether the fields arrived as separate files or as series inside one.- Parameters:
path – source-image path discovered beneath
scan()’s input directory. It remains relative when the input directory was relative.plate – source plate key inferred from the input layout: the plate directory for
plate_well, or the source-directory basename forwellandflatlayouts.well – source well key inferred from the input layout, or
DEFAULT_WELLfor a flat layout.field – source field key formed from the filename stem after removing channel, z, and time tokens; nested directory components and a
#s<n>series suffix are included when applicable.channel – channel token parsed from the filename, such as
"C2", orNonewhen channels are stored inside the source file.z – number of z planes contributed by the source, or zero for a source that could not be read during scanning.
t – number of timepoints contributed by the source, or zero for a source that could not be read during scanning.
meta – remaining scan metadata, including extension, dimensions, axes, series and source-relative path, assumptions, and any read error.
n_channels – number of channels reported or inferred inside the file; defaults to one when no internal channel dimension is available.
- spacr.convert.assign_wells(names: Sequence[str], *, n_wells: int | None = None) Dict[str, str][source]¶
Map arbitrary well-folder names onto plate well ids.
The rule, in full:
A name that already is a well address keeps it —
normalise_well()handlesa1/A-1/A01/aa1.The plate format is chosen by
plate_format_for_names(): 384 unless a claimed address or the sheer number of names needs the 1536, in which case rows run toAFand columns to 48.Every remaining name is sorted with
_natural_key()and handed the next free id of that plate, skipping any id claimed in step 1.
All three are deterministic — the same folder names always produce the same wells — and step 3 is only reversible because the map file records
source_wellnext towell. Which is the point: after conversionplate1_A01means nothing without the map, soplan()also reports every synthetic assignment by name.- Parameters:
names – the well-folder names found for one plate.
n_wells – force a plate format instead of choosing one.
- Returns:
{original name: well id}.- Raises:
ConfigurationError – when the names fit no standard plate — the real limit is 1536, not 384.
- spacr.convert.convert(conversion_plan: ConversionPlan, dst: str, overwrite: bool = False, map_name: str = MAP_FILENAME, progress: Callable[[int, int, str], None] | None = None, ledger: spacr.errors.RunLedger | None = None, resume: bool = False, checkpoint_path: str | None = None) ConversionResult[source]¶
Execute a plan: write the TIFFs, the map file and the run stamp.
Each source is opened once and all of its planes written from that one read. A source that raises is recorded on the ledger and the batch carries on, so one corrupt file out of 400 costs one file.
- Parameters:
conversion_plan – the plan from
plan(), already reviewed.dst – destination folder, created if missing. Must not be the source folder: converting in place is what makes
spacr.io.convert_to_yokogawa()’s output impossible to re-scan, since its own outputs land next to its inputs.overwrite – when False (default) a target that already exists is left alone and counted in
ConversionResult.existing.map_name – filename for the map, written inside
dst.progress – optional
progress(done, total, message), called once per source.ledger – reuse an existing ledger instead of making one.
resume – reuse fields recorded by a compatible checkpoint. Every target in a recorded field is revalidated as a readable TIFF before that field is skipped; missing or corrupt targets are repaired.
checkpoint_path – checkpoint JSON path. Defaults to
dst/.spacr_conversion.checkpoint.json. A checkpoint is written after every complete field even whenresumeis False, so a later invocation can opt in after a crash.
- Returns:
- Raises:
ConfigurationError – when the plan has blocking errors (a target-name collision), or
dstis the source folder.
- spacr.convert.convert_folder(settings: Mapping[str, Any] | None = None, **overrides: Any) ConversionResult[source]¶
Scan, preview, convert and map one folder in a single call.
The entry point the CLI and the Qt bridge want: one function, one settings dict, the same three phases underneath. It always prints the plan before writing anything, so even a headless
spacr-run format_convertleaves the source → target table in the log where a surprised user can find it.Settings (see
default_settings()):srcfolder to convert. Required.
dstdestination; defaults to
<src>_yokogawa. Neversrc.layout/z_handling/plate_namingoverwriteFalse by default — existing targets are left alone.
db_pathwhen set,
populate_db_from_map()loads the map into it.preview_onlyprint the plan and stop. The returned result has written nothing and has no
map_path.resumeaccept complete fields from a compatible checkpoint after validating every output TIFF.
checkpoint_pathoptional checkpoint JSON path; defaults inside
dst.
- Returns:
the
ConversionResult; forpreview_onlyan empty one carrying the plan.- Raises:
ConfigurationError – no
src, or a plan with blocking errors — a name collision must stop the run, not be printed and walked past.
- spacr.convert.default_settings(settings: Mapping[str, Any] | None = None) Dict[str, Any][source]¶
Return the settings
convert_folder()understands, with defaults.Shaped like every other
spacr.settingsfactory — pass a partial dict, get it back filled in — so the CLI and the GUIs can build a panel from it without special-casing this module.- Parameters:
settings – partial settings; keys given here win.
- spacr.convert.missing_reader_message(ext: str) str[source]¶
Return the user-facing sentence for an unreadable format.
Names the package and the exact install command. This is what the plan and the ledger carry instead of an ImportError traceback.
- Parameters:
ext – file extension including the dot.
- spacr.convert.normalise_well(name: str, *, n_wells: int | None = None) str | None[source]¶
Return
nameas a canonical well id (A01,AA48), or None.a1,A-1andA01all normalise toA01;aa1givesAA01, which is a real well of a 1536-plate. Anything that is not a well address at all (wt,KO_clone3,fov01) returns None and gets a synthetic id fromassign_wells()instead.This used to be a private
^([A-Pa-p])[ _\-]?(\d{1,2})$with an extra1 <= column <= 24check, soQ01(row 17) andA25(column 25) — both perfectly good wells of the 1536-plate the heatmap already draws — came back None and were renamed to the next free synthetic address.spacr.schema.parse_well()is now the parser, andspacr.schema.plate_format_for()decides whether the position it produced is a well that exists.- Parameters:
name – the source well-folder name.
n_wells – restrict to one plate format, e.g.
384. The default accepts any address that exists on some standard plate, which is to say up to 1536.
- Returns:
the canonical well id, or None when
nameis not one.
- spacr.convert.off_plate_reason(name: str) str | None[source]¶
Say why
namelooks like a well but is not one, or return None.The dangerous middle case.
wtis obviously not a well and nobody is surprised when it gets a synthetic address;ZZ99andA0parse into a row and a column and then turn out to sit on no plate that exists, so handing them a synthetic address silently is how a typo becomes a well name nobody can trace.- Parameters:
name – the source well-folder name.
- Returns:
a sentence naming the name and the position it read as, or None when the name either is a real well or is not well-shaped.
- spacr.convert.plan(sources: Sequence[SourceImage], z_handling: str = Z_KEEP, plate_naming: str = 'index', well_map: Mapping[Any, str] | None = None, plate_map: Mapping[str, str] | None = None) ConversionPlan[source]¶
Turn scanned sources into the preview table. Writes nothing.
Ids are handed out deterministically, always from a
_natural_key()sort so that the same tree yields the same numbers on every machine:plate —
plate1, plate2, …in sorted order of the source plate folder (plate_naming='index', the default, and what produces theplate1_A01_…the converter is specified against), or the sanitised folder name withplate_naming='name'.well — see
assign_wells().field — 1..N per well, over the distinct field keys; or, for sources
scan()read by ametadata_type, the field numbers the filenames state, when every field has one and none repeats.channel — 1..N per plate, over the distinct channel keys, so
C01means the same stain in every well of a plate.
z_handlingis explicit on purpose.Z_KEEP(default) writes every plane;Z_MAXmax-projects andZ_FIRSTkeeps only plane 1 — both are announced inConversionPlan.warningsand recorded per row in the map file, because a converter that silently flattens a z-stack loses data nobody notices for months.- Parameters:
sources – output of
scan().z_handling – one of
Z_HANDLING.plate_naming –
'index'or'name'.well_map – explicit
{well key: well id}or{(plate key, well key): well id}overrides.plate_map – explicit
{plate key: plate token}overrides.
- Returns:
- Raises:
ConfigurationError – for an unknown
z_handlingorplate_naming.
- spacr.convert.plate_format_for_names(n_names: int, wells: Sequence[str], minimum: int = DEFAULT_PLATE_FORMAT) int | None[source]¶
Return the plate format that has to be used for one plate’s wells.
The smallest standard format that is at least
minimum, holds every address inwells, and has room forn_namesdistinct sources. That second clause is the 1536 fix: one folder namedAA01means the plate is a 1536, whether or not there are 1536 of them.- Parameters:
n_names – how many distinct source names must be given a well.
wells – the canonical addresses already claimed by name.
minimum – never return a format smaller than this.
- Returns:
the well count of the format to use, or None when the names fit no standard plate.
- spacr.convert.populate_db_from_map(db_path: str, map_path: str, table: str = CONVERSION_TABLE) int[source]¶
Load a map file into
measurements.dbso the run can be joined back.This is the read-back that closes the loop. After spaCR has measured the converted images, every row of every measurement table carries
plateID/rowID/columnID/fieldID— but nothing that says the field came fromrun1/wt/fov07_C2.tif. This writes the map into the same database asconversion_map, keyed exactly the way the measurement tables are, so:SELECT c.*, m.source, m.source_well, m.source_field FROM cell AS c JOIN conversion_map AS m ON c.plateID = m.plateID AND c.rowID = m.rowID AND c.columnID = m.columnID AND c.fieldID = m.fieldID
joins the original metadata onto the measurements.
prc(well level,plate_r1_c1) andprcf(field level,plate_r1_c1_f1) are there for the same join in one column.The table is replaced, not appended: re-running a conversion and re-populating must not leave two generations of rows behind.
- Parameters:
db_path – SQLite database to write into; created if missing.
map_path – the CSV from
write_map().table – table name,
conversion_mapby default.
- Returns:
number of rows written.
- Raises:
ConfigurationError – when the map is missing or malformed.
- spacr.convert.read_map(path: str) pandas.DataFrame[source]¶
Read a map file back.
- Parameters:
path – a CSV written by
write_map().- Returns:
the map as a DataFrame.
- Raises:
ConfigurationError – when the file is missing or is not a spaCR conversion map — a wrong path here would otherwise populate a database with somebody else’s columns.
- spacr.convert.read_scn(path: Any, index: int | None = 0) Tuple[numpy.ndarray, Dict[str, Any]][source]¶
Read a Bio-Rad Image Lab
.scn(Gel Doc, ChemiDoc) image.Pure numpy and the standard library: the file is a MIME multipart document whose
ScanImageTagNparts each hold rawImageData(width x height uint16 in the stated endianness) and an XMLImageHeader. Image Lab storeszero_is="white"data, so such an image is inverted asdata_ceiling - raw; the result then looks like Image Lab’s own PDF and TIFF exports (dark objects on a light background).zero_is="black"data is returned as stored.- Parameters:
path – the
.scnfile.index – which image of a multi-image file;
Nonestacks every image into(N, H, W)(they must share one size).
- Returns:
(image, meta).imageisuint16.metaholdswidth,height,endian,data_ceiling,zero_is,inverted,size_mmandpixel_size_mm((x, y)mm or None when the file does not know its physical size),pixels_per_um(or None),n_images, and the scan attributes Image Lab recorded (imager,image_date,exposure_timein seconds,application,excitation_source, …). Withindex=Noneit is the first image’s metadata plusimages, a list of every image’s.- Raises:
ValueError – for a file that is not an Image Lab document, holds no image, or is truncated.
IndexError – for an
indexpast the last image.
- spacr.convert.reader_available(ext: str) bool[source]¶
True when
extcan actually be read on this machine.- Parameters:
ext – file extension including the dot.
- spacr.convert.reader_requirement(ext: str) Tuple[str, str] | None[source]¶
Return
(module, pip command)needed to readext, or None.- Parameters:
ext – file extension including the dot, case-insensitive.
- Returns:
None for formats served by always-present dependencies (TIFF, PNG, JPEG, BMP).
- spacr.convert.scan(src: str, layout: str = 'auto', extensions: Sequence[str] | None = None, metadata_type: str | None = None, custom_regex: str | None = None) List[SourceImage][source]¶
Walk
srcand describe every image it holds. Writes nothing.This is the read-only half of the converter: it opens headers, not pixels, and produces the
SourceImagelist thatplan()turns into a preview.Layouts:
'plate_well'src/<plate>/<well>/…— the user’srun1/wt/case.'well'src/<well>/…, withsrc’s own name as the plate.'flat'images directly in
src; one plate, one well (A01), one field per file.'auto'(default)chooses from how deep the images actually sit.
A file whose reader is not installed, or that will not open, comes back as a
SourceImagewithmeta['error']set rather than raising — the plan shows it, the ledger counts it and the summary names it.- Parameters:
src – folder to scan.
layout – one of
LAYOUTS.extensions – override the scanned extensions.
metadata_type – a filename convention from Mask’s
metadata_typelist ('opera_phenix','cq1','zeiss_zen_split_tiles', …). When given, plate, well, field, channel, z and t are READ FROM EACH NAME by that convention’s pattern, and the folders only supply what the name does not carry – the plate for a convention with no plate group, the well for one whose “well” is a scene index or a literal word. A file whose name does not follow the convention is reported unreadable, with the reason, rather than guessed at.None,''and'auto'keep the folder-and-token inference.custom_regex – the pattern for
metadata_type='custom', with named groupswellID,fieldIDandchanID(and optionallyplateID,timeID,sliceID).
- Returns:
one
SourceImageper file (or per series in a file).- Raises:
ConfigurationError – when
srcis not a directory,layoutis not recognised, ormetadata_typenames no convention.
- spacr.convert.scn_to_rgb8(image: numpy.ndarray, meta: Mapping[str, Any]) numpy.ndarray[source]¶
Render a
read_scn()image asH x W x 3uint8 RGB.The scale is linear from 0 to the scanner’s
data_ceiling, never a per-image stretch, so every plate photographed on one imager keeps the same grey levels.- Parameters:
image – the uint16 array
read_scn()returned.meta – its metadata.
- Returns:
grey RGB pixels.
- spacr.convert.target_name(plate: str, well: str, field: int, channel: int, z: int = 1, t: int = 1, action: int = 1) str[source]¶
Build one Yokogawa filename.
plate1_A01_T0001F001L01A01Z01C01.tif— the exact shapespacr.utils._get_regex('cellvoyager', 'tif')parses, so a folder of these can be handed to Mask/Measure withmetadata_type='cellvoyager'and nothing else.- Parameters:
plate – plate token (no underscores — see
_sanitise()).well – canonical well id, e.g.
A01.field – 1-based field id.
channel – 1-based channel id.
z – 1-based z-slice id.
t – 1-based timepoint id.
action – the
A##action id; spaCR ignores it, so it is 1.
- spacr.convert.well_sequence(n_wells: int = DEFAULT_PLATE_FORMAT) Tuple[str, ...][source]¶
Return every well id of an
n_wellsplate, row-major.Built from
spacr.schema.PLATE_FORMATSand rendered byspacr.schema.well_id(), so a 1536-well plate’s rows pastZcome outAA…AFand its columns run to 48. This module used to carry its own'ABCDEFGHIJKLMNOP'andrange(1, 25), which is exactly why it could not name a well on a plate bigger than 384.- Parameters:
n_wells – a key of
spacr.schema.PLATE_FORMATS.- Returns:
the well ids,
A01first.- Raises:
ConfigurationError – for a non-standard plate format.
- spacr.convert.write_map(result: ConversionResult, path: str) pathlib.Path[source]¶
Write the map file for
result.The map is the whole point of the exercise: once the images are called
plate1_A01_T0001F001L01A01Z01C01.tif, the only thing that can say which microscope file that came from is this table.One row per output TIFF, columns
MAP_COLUMNS:target/target_path— the converted file;source/source_relpath— the original file;plate/well/field/channel/z/t— the assigned ids;source_plate/source_well/source_field/source_channel/source_z/source_t— what those ids were before, which is what makes the renaming reversible;plateID/rowID/columnID/fieldID/prc/prcf— the spaCR join keys, in the exact formspacr.utils._map_wells()produces;z_handling/n_z_planes/n_timepoints— how the third and fourth dimensions were treated;status—converted,existingorfailed;meta_json— reader, extension, series and axes.
- Parameters:
result – a finished
ConversionResult.path – destination CSV.
- Returns:
the written path.