spacr.image_import

Workflow inputs and outputs

Import Images

Review field identities, image channels and optional external masks before writing a separate project. Image-only imports still need segmentation before measurement.

Open: Import → Import Images.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.

Outputs

  • Microscope images — Source image folder; original files, supported vendor files or imported TIFFs.

  • Images and label masks — merged/*.npy in the project; channels and integer label planes share each field array.

  • Label masks — masks/ when retained, or explicitly saved image/mask pairs. Intermediate masks may be removed by cleanup.

After this module

  • Mask: Use imported image planes and identities; image-only imports still need segmentation.

API reference.

Module tutorial.

Work out how a folder of images is named, instead of being told.

spaCR’s import path used to ask the user to pick a filename convention from a closed list – cellvoyager, cq1, custom (write your own regular expression), auto – and the corpus in tests/import_corpus.py measures what that costs: of ten real acquisition layouts, two parse and eight recover nothing. Opera Phenix and ImageXpress are among the eight.

THE FIXED LIST IS THE PROBLEM, NOT THE REGULAR EXPRESSIONS IN IT. Every one of the ten encodes the SAME six facts – plate, well, field, channel, z, t – and differs only in how. So rather than matching a whole filename against one of N templates, this module reads the parts of the name that VARY across the folder and works out what each varying part means.

WHY VARIANCE IS THE RIGHT SIGNAL. A token that is the same in every file carries no information: L01 in plate1_A01_T0001F001L01A01Z01C01.tif is a constant for the whole plate and identifies nothing. A token that takes two values across eight files is an axis with two positions. Reading the folder tells you which tokens are axes; nothing about the filename alone can.

WHY MARKERS ARE STILL NEEDED. Variance says a token IS an axis; it cannot say WHICH axis. F001 and C01 both vary and are not interchangeable. The letter in front is what the convention uses to say so, and every convention in the corpus uses one – so the marker table below is the vocabulary, and it is short because the conventions agree more than they disagree.

WHAT THIS DELIBERATELY DOES NOT DO. It never guesses a role it has no evidence for. A token that varies and carries no recognisable marker is reported as UNPLACED, with its values, so the caller can ask rather than assume. The criterion is that a partly-unparseable tree imports what it can and REPORTS the rest: a wrong answer that looks plausible is the failure mode this whole module exists to avoid, and it is the one the consolidate bugs demonstrated – images disappearing with no error rather than an import refusing.

Classes

ImportPlan

What WOULD be imported, for a human to check before anything is written.

ImportResult

What apply_import() did.

InferredLayout

What was worked out about a folder, and what was not.

InsideFile

What one file's own metadata says about the axes it holds.

TokenSlot

One position in a filename, and what was seen there across the folder.

Functions

apply_import(→ ImportResult)

Write the spaCR project plan describes.

canonical_name(→ str)

The spaCR filename for one resolved image.

infer_layout() → InferredLayout)

Work out what the filenames under root encode.

load_plan(→ ImportPlan)

Reload a saved plan, so the second import of the week is one press.

plan_import(→ ImportPlan)

Propose an import of root. Writes nothing.

read_axes_inside(→ InsideFile)

What path's own metadata says its pages are.

save_plan(→ pathlib.Path)

Write plan where load_plan() can read it back.

tokenise(→ List[Tuple[str, str]])

Split a name into (kind, text) runs, kind "alpha" or "digit".

Module Contents

class spacr.image_import.ImportPlan[source]

What WOULD be imported, for a human to check before anything is written.

THE PLAN IS THE FEATURE. metadata_type is unusable today not because its regular expressions are bad but because the user cannot see what they did until masks come out wrong, so the whole point of this module is that the proposal is visible and correctable BEFORE it is acted on. This mirrors spacr.foreign, whose column mapping is the middle of its screen and is editable in place – the same split, applied to images.

Nothing here touches the destination. problems() says what would stop an import; with_mapping() returns a NEW plan, resolved in memory, so editing is instant and a rejected edit costs nothing.

Parameters:
  • layout – what the names gave.

  • inside – per relative path, what the file’s own metadata gave.

  • mapping – user answers for axes inference could not name, as {token position: {value: index}} – e.g. {0: {"DAPI": 1, "GFP": 2}} for a tree with dye folders.

columns() → List[str][source]

file then every axis anything resolved, counts last.

ONLY THE AXES THIS FOLDER HAS. A column of empty cells for an axis no file carries reads as “spaCR looked for a Z and found none”, which is a different claim from “this acquisition has no Z” and is the one a user acts on. The tiled tree gains a tile column and the flat OME one gains c_count; neither shows the other’s.

counts() → Dict[str, int][source]

Distinct values per axis – the numbers a user checks first.

A plate with the wrong number of wells or channels is visible here and nowhere else until the run has finished.

problems() → List[str][source]

Everything that would make this import wrong, in plain sentences.

REPORTED, NOT RAISED, and all of them at once. The first version of the translation audit raised on its first finding and reported one problem where there were three; a plan that stops at the first complaint makes the user fix things one round-trip at a time.

rows(limit: int | None = None) → List[List[str]][source]

One row of strings per file, cells aligned to columns().

SEPARATE FROM table() SO THE GUI AND THE TEXT CANNOT DISAGREE. Both are the same proposal seen twice – the Qt model draws these rows into a widget and table() pads them into a monospace block – and a screen that decided its own columns would be free to show a different answer from the one the CLI and the saved plan show.

table(limit: int = 8) → str[source]

The proposal, as the table a user reads before pressing anything.

Example filenames BESIDE the fields they were parsed into, because a wrong guess is only visible next to the name it came from – reading a column of numbers cannot tell you the field and the channel were swapped.

with_mapping(mapping: Dict[int, Dict[str, int]]) → ImportPlan[source]

A new plan with mapping merged in. Nothing on disk is touched.

Parameters:

mapping – further answers as {token position: {value: index}}; for each position they are added to, and override, the answers this plan already holds.

property files: Dict[str, Dict[str, object]][source]

Relative path -> every axis known, from all three sources.

Names first, then the file’s own metadata, then the user’s mapping – in that order because each is more specific than the last, and the user is the only one who can be right about the last of them.

property root: pathlib.Path[source]

The folder every path in this plan is relative to.

Returns:

the root path.

property unmapped: Dict[int, List[str]][source]

Axes still waiting on an answer, after the mapping is applied.

class spacr.image_import.ImportResult[source]

What apply_import() did.

Parameters:
  • destination – the plate folder written.

  • written – how many images the project now names.

  • linked – how many were symlinked rather than copied.

  • bytes_saved – source bytes not duplicated by linking.

  • skipped – source paths that were not written, with the reason.

  • stitched – fields assembled from several tiles.

  • unverified – written name -> why its seams are not believed. A stitch whose tiles correlate on nothing is still the best answer available, and saying so is the difference between a field a user can check and one they cannot.

summary() → str[source]

What was written, what was not, and why.

EVERY SKIPPED FILE IS NAMED WITH ITS REASON, not counted. A count says an import was incomplete; only the name and the reason say what to do about it, and the reasons differ – a tiled tree needs tiles_as_fields, an unwritable destination needs a different folder. The consolidate bugs this module replaces lost images with no message at all, so silence here would be the same failure with a nicer table in front of it.

class spacr.image_import.InferredLayout[source]

What was worked out about a folder, and what was not.

Parameters:
  • root – the folder examined.

  • per_file – relative path -> {axis: value} for what was resolved.

  • slots – the token slots, resolved and unresolved alike.

  • unplaced – axes-shaped tokens that vary but carry no known marker, as {position: [values]}. THE IMPORTANT FIELD: it is what the caller shows the user instead of guessing.

  • skipped – files whose token shape did not match the majority, so nothing was claimed about them.

  • sampled – how many files were read to decide.

summary() → str[source]

One line per axis, plus what could not be placed.

property axes: set[source]

Which axes were resolved anywhere.

class spacr.image_import.InsideFile[source]

What one file’s own metadata says about the axes it holds.

Parameters:
  • pages – how many pages the file has. 1 means a plain 2-D image and nothing below matters.

  • axes – the axis letters the file declares, e.g. "CYX", "ZYX", "TYX". Empty when the file declares none.

  • sizes – {axis: length} for the non-spatial axes only, so {"c": 2} or {"z": 5}.

  • declared – whether the axes came from the FILE or were guessed. False with pages > 1 is the honest unknown – a multi-page TIFF with no metadata could be Z, T or C and nothing can say which.

property is_ambiguous: bool[source]

Several pages and nothing saying what they are.

class spacr.image_import.TokenSlot[source]

One position in a filename, and what was seen there across the folder.

Parameters:
  • index – position among the tokens of a name.

  • marker – the alphabetic run immediately before it, lower-cased.

  • values – every value seen at this position, in first-seen order.

  • axis – the axis it was resolved to, or "" when unplaced.

property varies: bool[source]

Whether this slot takes more than one value across the folder.

spacr.image_import.apply_import(plan: ImportPlan, destination, *, link: bool = True, plate: str = 'plate1', tiles_as_fields: bool = False, stitch_tiles: bool = True) → ImportResult[source]

Write the spaCR project plan describes.

LINKS BY DEFAULT, AND THAT IS THE POINT. consolidate – the closest thing spaCR had to this – COPIES every image to rearrange its name, so a 300 GB plate costs 600 GB to import. Nothing about renaming requires duplicating bytes. Where symlinks are unavailable the copy still happens, and the result says which was used and what linking saved.

REFUSES ON A PLAN WITH PROBLEMS. An import is the irreversible half, and every problem the plan states is a way for the result to be quietly wrong – an unnamed axis means images that cannot be told apart. Fix the plan or answer its questions; do not write past it.

Parameters:
  • plan – a plan whose ImportPlan.problems() is empty.

  • destination – the plate folder to create.

  • link – symlink rather than copy. Falls back to copying per file.

  • plate – the plate name to write into the filenames.

  • stitch_tiles –

    put each field’s tiles back together into the one image the field is. ON BY DEFAULT: a field arrives as one image, and turning this off is the exception rather than the rule.

    THIS IS WHY THE FILENAME NEEDS NO TILE SLOT. The convention the core modules read is plate/well/T/field/L/A/Z/channel, so four tiles of one field produce four images with one name and three would be overwritten. Stitching removes the question rather than answering it: a stitched field IS one image with one name, and nothing downstream has to learn what a tile is. See spacr.image_stitch for how the placement is worked out, and for what it does when the tiles give it nothing to work with.

  • tiles_as_fields –

    give each tile its own field number instead.

    THE OPT-OUT, and it takes precedence over stitch_tiles because it is the more specific request: a caller who says “each tile is a field” has said what they want the tiles to be. It discards the fact that they are ONE field, so anything measuring per field is then counting quarters of one – which is why it is not the default.

    With both off, tiled images are SKIPPED with a reason rather than overwritten. That was the only honest answer before stitching existed, and it is kept because refusing to write is still better than writing three of four images over each other.

Raises:

ValueError – when the plan still has problems.

spacr.image_import.canonical_name(entry: Dict[str, object], *, plate: str = 'plate1') → str[source]

The spaCR filename for one resolved image.

Missing axes take 1 rather than 0: spaCR’s convention is one-based, and a plate whose only timepoint is T0000 reads as a bug in the acquisition rather than as an absence.

Parameters:
  • entry – one resolved file’s axes, read by the keys plate, well, t, field, z and channel; a missing well becomes A01 and a missing or zero index becomes 1.

  • plate – the plate name used when entry has no plate key.

spacr.image_import.infer_layout(root, *, sample: int = 400, extensions: Sequence[str] = ('.tif', '.tiff', '.png', '.jpg', '.jpeg', '.bmp')) → InferredLayout[source]

Work out what the filenames under root encode.

Parameters:
  • root – the folder to inspect.

  • sample – how many files to read before deciding. A SAMPLE, not the whole tree: the answer is a naming convention, and a convention is visible in a few hundred names. This is what keeps inspecting a 400-plate archive as fast as inspecting one plate.

  • extensions – which files count as images.

Returns:

an InferredLayout. Never raises for an unrecognised tree – an empty per_file with populated unplaced is the honest answer and the caller is expected to show it.

spacr.image_import.load_plan(path) → ImportPlan[source]

Reload a saved plan, so the second import of the week is one press.

A lab images the same way every week, and re-answering the same questions every time is how a tool stops being used. Reloading also makes an import SCRIPTABLE – the saved file is the whole answer, so a cluster job needs no GUI – and testable, which is why this exists rather than a cache.

The plan is re-derived from the folder and the saved MAPPING rather than trusting the saved per-file table: the folder may have gained images since, and a stale table would silently import last week’s files.

Parameters:

path – the JSON file written by save_plan(); its root and mapping keys are read and the rest is ignored.

spacr.image_import.plan_import(root, *, sample: int = 400, mapping: Dict[int, Dict[str, int]] | None = None, read_files: bool = True) → ImportPlan[source]

Propose an import of root. Writes nothing.

Parameters:
  • root – the folder of images.

  • sample – how many files to inspect; see infer_layout().

  • mapping – answers for axes the names cannot name.

  • read_files – also open each file for its axis metadata. On by default because a channel that lives inside a file is invisible without it; turn it off for a fast first look at a large archive.

spacr.image_import.read_axes_inside(path) → InsideFile[source]

What path’s own metadata says its pages are.

THE AXIS THAT IS NOT IN THE NAME. infer_layout reads names, and a name cannot carry what an acquisition put inside the file: an OME-TIFF holding two channels is one filename, and a Z-stack and a timelapse of the same field have IDENTICAL names. Only the file says which.

NEVER GUESSES. A multi-page TIFF with no axis metadata could be Z, T or C, and this returns declared=False with the page count rather than picking one. The caller shows that to the user: a page index treated as meaningful on its own is exactly the guess this avoids.

Parameters:

path – an image file.

Returns:

an InsideFile. A file that cannot be opened at all comes back as InsideFile(pages=0) rather than raising – a folder of thousands should not fail wholesale because one file is truncated.

spacr.image_import.save_plan(plan: ImportPlan, path) → pathlib.Path[source]

Write plan where load_plan() can read it back.

THE MAPPING IS THE ANSWER WORTH KEEPING. The per-file table is written too, because a saved plan a person cannot read is not reviewable, but load_plan() re-derives the files from the folder and trusts only the mapping – see its own docstring for why replaying a stale table would import last week’s images.

Parameters:
  • plan – the plan to save.

  • path – the JSON file to write; missing parent folders are created and an existing file is overwritten.

Returns:

the path written, so a caller can report it.

spacr.image_import.tokenise(name: str) → List[Tuple[str, str]][source]

Split a name into (kind, text) runs, kind "alpha" or "digit".

"plate1_A01_F001" becomes alpha/digit pairs. Separators are dropped: conventions disagree about _ versus - versus nothing, and the disagreement carries no information.

Parameters:

name – the file or folder name to split; every character that is not an ASCII letter or a digit is dropped.