spacr.channel_sorting

Sort a folder of images (and their masks) into channels, rename, and merge.

The case this module exists for: masks were drawn in Make Masks on a folder holding two kinds of image at once – nucleus-stain images with nucleus masks and cell-stain images with cell masks, say – and those have to become one multi-channel field per position, in the layout spaCR’s Mask/Measure pipeline reads.

It happens in three separate steps, and nothing is moved until the last:

  1. Channels and sets. Each image gets a CHANNEL, and images that show the same field in different channels share a SET. Channels come from the user’s selection (parse_names() channels=) or from a regex’s chanID group; sets come from the regex’s other named groups (parse_names(), checked by check_sets()), from a regex spaCR proposes (infer_regex()), or, ignoring names, from pairing each image with the most similar image of the same size in every other channel (detect_sets()).

  2. The plan. build_plan() names every set in the Yokogawa / CellVoyager form plate1_A01_T0001F001L01A01Z01C01.tif – exactly what spacr.utils._get_regex('cellvoyager', ...) parses – and lists every move. It writes nothing.

  3. Apply. apply_plan() moves each image into C01/, C02/… under a new folder, its mask into that channel folder’s masks/ under the new name (and its .curation.json ledger with it, so the pairing survives), records every move in channel_sorting_manifest.csv, and then merge_sorted() builds stack/, masks/<role>_mask_stack/ and merged/ with spaCR’s own functions (spacr.utils._extract_filename_metadata(), spacr.io._escaped_field_stem() and spacr.io._load_and_concatenate_arrays()), so each merged array has the pipeline’s own layout: intensity channels first, then one label plane per mask role in the order cell, nucleus, pathogen.

Regex conventions: chanID is the channel; plateID, wellID, fieldID and timeID say where a set goes; any other named group only helps tell sets apart. A group that did not take part in a match counts as the empty string, so an optional (?:_(?P<fieldID>\d+))? reads the numbered copies spacr.folder_consolidation makes.

Classes

ApplyResult

What apply_plan() did.

DetectedSets

What detect_sets() paired.

ParsedName

One image name read through a regex.

PlanRow

One image to move, with its mask and the mask's ledger.

SetName

Where one set goes in the Yokogawa naming.

SetReport

Whether a set of parsed names forms complete, unique sets.

SortPlan

Everything apply_plan() will do, decided before it does any.

Functions

apply_plan(→ ApplyResult)

MOVE the images and masks as planned, then merge.

build_plan(→ SortPlan)

Decide every move, name and check before anything is touched.

canonical_well(→ Optional[str])

A1, a01 or r01c01 as A01; None for anything else.

check_sets(→ SetReport)

Group parsed names into sets and report every way they fail to fit.

compile_regex(pattern)

Compile pattern, or return the error message.

default_mask_roles(→ Dict[int, str])

Guess which role each masked channel's masks have.

detect_sets(→ DetectedSets)

Pair images across channels by name similarity, order and size.

image_shape(→ Optional[Tuple[int, ...]])

Read an image's pixel shape from its header, without its pixels.

infer_regex(→ Optional[str])

Propose a regex that maps every image to a unique, complete set.

list_folder_images(→ List[str])

Return the image file names directly in folder, naturally sorted.

mask_for(→ Optional[str])

Return the path of the saved mask of image name, or None.

mask_thumbnail(→ Optional[numpy.ndarray])

A small preview of a label mask: 255 where any object is, else 0.

masks_folder(→ str)

Where Make Masks keeps the masks of folder's images.

merge_sorted(→ Tuple[List[str], List[str]])

Build stack/, the mask stacks and merged/ from sorted channels.

name_distance(→ float)

How different two names are, 0 (same tokens) to 1 (nothing shared).

name_sets(→ Dict[SetKey, SetName])

Give every set a unique plate, well, field and time.

natural_key(→ tuple)

Sort key that orders img2 before img10.

parse_names(→ List[ParsedName])

Read each name's channel and set through pattern.

set_label(→ str)

A set key as short text, wellID=A01 fieldID=2.

split_extension(→ Tuple[str, str])

Split name into stem and extension, keeping .ome.tif whole.

thumbnail(→ Optional[numpy.ndarray])

A small 8-bit grey preview of an image, contrast-stretched.

unused_folder(→ str)

parent/name, or name_2, name_3... when that exists.

well_name(→ str)

The index-th well of a 384-well plate, row by row: 0 is A01.

yokogawa_name(→ str)

The Yokogawa file name of one channel of one set.

Module Contents

class spacr.channel_sorting.ApplyResult[source]

What apply_plan() did.

Variables:
  • dest – the folder written.

  • manifest – the move/rename record.

  • moved – files moved (images, masks and ledgers).

  • stacks – stack/*.npy written.

  • merged – merged/*.npy written.

class spacr.channel_sorting.DetectedSets[source]

What detect_sets() paired.

Variables:
  • sets – {set key: {channel: name}}, keyed (("set", "000001"),).

  • unpaired – images left without a partner in some channel.

  • problems – reasons the pairing cannot be used as it is.

class spacr.channel_sorting.ParsedName[source]

One image name read through a regex.

Variables:
  • name – the file name.

  • matched – whether the regex matched it.

  • groups – every named group’s value; one that did not take part is the empty string.

  • channel – the 1-based channel, from the selection or chanID; None when neither gives one.

  • set_key – the set the image belongs to; None when unmatched.

class spacr.channel_sorting.PlanRow[source]

One image to move, with its mask and the mask’s ledger.

Variables:
  • channel – 1-based channel.

  • name – the set’s place.

  • set_key – the set the image is in.

  • source_image – current image path.

  • target_image – where it goes.

  • source_mask – current mask path, or None.

  • target_mask – where the mask goes, or None.

  • source_ledger – the mask’s .curation.json, or None.

  • target_ledger – where the ledger goes, or None.

  • convert – "scn" for a Bio-Rad .scn rewritten as a TIFF, or "rgb" or "zstack" when the image is written as one converted 2-D plane (the original is kept under originals/).

class spacr.channel_sorting.SetName[source]

Where one set goes in the Yokogawa naming.

Variables:
  • plate – the plate token, with no underscore.

  • well – e.g. A01.

  • field – 1-based field.

  • time – 1-based timepoint.

class spacr.channel_sorting.SetReport[source]

Whether a set of parsed names forms complete, unique sets.

Variables:
  • channels – the channels present, sorted.

  • complete – {set key: {channel: name}} for sets with exactly one image in every channel.

  • incomplete – {set key: [missing channels]}.

  • duplicated – {(set key, channel): [names]} – two images claiming the same place.

  • unmatched – names the regex did not match.

  • no_channel – matched names with no channel.

summary() → str[source]

A few plain lines saying what is and is not right.

property ok: bool[source]

True when every image is in exactly one complete set.

class spacr.channel_sorting.SortPlan[source]

Everything apply_plan() will do, decided before it does any.

Variables:
  • folder – the image folder.

  • dest – the new folder the channels, stacks and merged/ go in.

  • rows – one per image.

  • channels – channels, sorted.

  • mask_roles – {channel: role} for channels whose masks are merged.

  • problems – reasons the plan cannot be applied.

  • warnings – things the user should know before it is.

  • convertible – RGB images and z-stacks that build_plan() would convert when asked to (convert=True).

summary() → str[source]

Plain lines saying what applying the plan will do.

property n_sets: int[source]

How many sets the plan holds.

property ok: bool[source]

True when there is something to do and nothing stops it.

spacr.channel_sorting.apply_plan(plan: SortPlan, *, merge: bool = True, log: Callable[[str], None] | None = None) → ApplyResult[source]

MOVE the images and masks as planned, then merge.

Every move is written to <dest>/channel_sorting_manifest.csv as it happens, so a run that stops half-way still says which file went where. Nothing is overwritten: every target is checked to be free before the first move.

Parameters:
  • plan – an ok plan from build_plan().

  • merge – run merge_sorted() afterwards.

  • log – called with progress lines; default prints.

Returns:

an ApplyResult.

Raises:
  • ValueError – when the plan has problems, a source is gone or a target is taken.

  • OSError – when a move fails; the manifest records it and every move before it.

spacr.channel_sorting.build_plan(folder: str, sets: Dict[SetKey, Dict[int, str]], *, masks_dir: str | None = None, mask_roles: Dict[int, str] | None = None, dest: str | None = None, check_shapes: bool = True, convert: bool = False, masks: Dict[str, str | None] | None = None) → SortPlan[source]

Decide every move, name and check before anything is touched.

Parameters:
  • folder – the image folder.

  • sets – {set key: {channel: name}}, complete sets only.

  • masks_dir – an explicit masks folder.

  • mask_roles – {channel: role}; None guesses (default_mask_roles()); a channel missing from it, or given "none", has its masks moved but not merged.

  • dest – the new folder; default <folder>/sorted_channels.

  • check_shapes – read each image’s and mask’s size and refuse a set whose images differ in size, a mask that is not its image’s size or an image that is not 2-D.

  • masks – {image: mask path} naming each image’s mask outright (the organizer’s mask columns); an image missing from it has none. None looks each mask up with mask_for().

  • convert – write RGB images as grey and z-stacks as their maximum projection instead of refusing them; without it they are listed in SortPlan.convertible so the caller can ask.

Returns:

a SortPlan.

spacr.channel_sorting.canonical_well(text: str) → str | None[source]

A1, a01 or r01c01 as A01; None for anything else.

Parameters:

text – a well token.

spacr.channel_sorting.check_sets(parsed: Sequence[ParsedName]) → SetReport[source]

Group parsed names into sets and report every way they fail to fit.

Parameters:

parsed – the output of parse_names().

Returns:

a SetReport.

spacr.channel_sorting.compile_regex(pattern: str)[source]

Compile pattern, or return the error message.

Parameters:

pattern – the regex text.

Returns:

(compiled, None) or (None, message).

spacr.channel_sorting.default_mask_roles(members: Dict[int, List[str]], with_masks: Iterable[int]) → Dict[int, str][source]

Guess which role each masked channel’s masks have.

A channel whose names say nuc/dapi/hoechst is nucleus, cell/cyto is cell, parasite/pathogen is pathogen; the rest take the first unused of cell, nucleus, pathogen.

Parameters:
  • members – {channel: [names]}.

  • with_masks – the channels that have masks.

Returns:

{channel: role} for those channels, roles distinct.

spacr.channel_sorting.detect_sets(folder: str, channels: Dict[str, int], masks_dir: str | None = None, shapes: Dict[str, tuple | None] | None = None) → DetectedSets[source]

Pair images across channels by name similarity, order and size.

The channel with the most images is the reference. Each other channel’s images are matched one-to-one to the reference images by the smallest total cost, where the cost is name_distance() plus a small term for how far apart the two sit in their channel’s sorted order, and a pair whose images differ in size is not allowed. Names play no other part: each resulting set is keyed by its number alone. A mask whose size is not its image’s is reported as a problem.

Parameters:
  • folder – the image folder.

  • channels – {name: channel} for every image to pair.

  • masks_dir – an explicit masks folder.

  • shapes – {name: shape} already read; read here otherwise.

Returns:

a DetectedSets.

spacr.channel_sorting.image_shape(path: str) → Tuple[int, ...] | None[source]

Read an image’s pixel shape from its header, without its pixels.

Singleton axes are dropped, so a (1, H, W) TIFF and an (H, W) PNG compare equal.

Parameters:

path – an image file.

Returns:

the shape, or None when the file cannot be read.

spacr.channel_sorting.infer_regex(names: Sequence[str], channels: Dict[str, int] | None = None) → str | None[source]

Propose a regex that maps every image to a unique, complete set.

Names are split into tokens (at _ - . space, and failing that into digit and letter runs); the tokens that vary are found; one is tried as the channel – or, when the user assigned channels, the tokens that only follow the channel are set aside – and the rest must identify a set that holds exactly one image per channel. Each candidate is checked with check_sets() before it is returned. spaCR’s own Yokogawa pattern is tried first, and spacr.regex_infer.propose() last. Of the candidates that pass, the one with the FEWEST channels wins: any name can be read as a channel of its own in a single set, which passes and means nothing.

Parameters:
  • names – image file names.

  • channels – {name: channel} from the selection strategy; when every name has one, the regex needs no chanID group.

Returns:

the first regex that passes, or None.

spacr.channel_sorting.list_folder_images(folder: str) → List[str][source]

Return the image file names directly in folder, naturally sorted.

Parameters:

folder – the folder; a missing one holds none.

spacr.channel_sorting.mask_for(folder: str, name: str, masks_dir: str | None = None) → str | None[source]

Return the path of the saved mask of image name, or None.

Make Masks saves every mask as <masks folder>/<stem>.tif (spacr.qt.mask_engine.mask_save_path()). An image named by an absolute path outside folder – one dropped into the sort dialog from elsewhere, or a channel subfolder’s – has its mask in its own folder’s masks/.

Parameters:
  • folder – the image folder.

  • name – the image file name, or an absolute path.

  • masks_dir – an explicit masks folder, for images in folder.

spacr.channel_sorting.mask_thumbnail(path: str, size: int = 96) → numpy.ndarray | None[source]

A small preview of a label mask: 255 where any object is, else 0.

Parameters:
  • path – the mask file.

  • size – the longest side, in pixels.

Returns:

a 2-D uint8 array, or None when it cannot be read.

spacr.channel_sorting.masks_folder(folder: str, masks_dir: str | None = None) → str[source]

Where Make Masks keeps the masks of folder’s images.

Parameters:
  • folder – the image folder.

  • masks_dir – an explicit masks folder, when not <folder>/masks.

spacr.channel_sorting.merge_sorted(dest: str, mask_roles: Dict[int, str], *, log: Callable[[str], None] | None = None) → Tuple[List[str], List[str]][source]

Build stack/, the mask stacks and merged/ from sorted channels.

The Yokogawa names in C01/, C02/… are read with spaCR’s own cellvoyager pattern and spacr.utils._extract_filename_metadata(), each field’s stem is spacr.io._escaped_field_stem() – the stem the pipeline gives it – and its channels are stacked in channel order into stack/<stem>.npy. Each channel’s masks are copied to masks/<role>_mask_stack/<stem>.tif, and spacr.io._load_and_concatenate_arrays() writes merged/.

Parameters:
  • dest – the folder apply_plan() filled.

  • mask_roles – {channel: role}.

  • log – called with progress lines.

Returns:

(stack files, merged files).

spacr.channel_sorting.name_distance(a: str, b: str) → float[source]

How different two names are, 0 (same tokens) to 1 (nothing shared).

Parameters:
  • a – a file name.

  • b – another.

spacr.channel_sorting.name_sets(keys: Iterable[SetKey]) → Dict[SetKey, SetName][source]

Give every set a unique plate, well, field and time.

Sets keyed by detect_sets() get one well each, A01 onward, field 1, 384 wells to a plate. Otherwise the plate is plateID (or plate1); the well is wellID when every value reads as a well, else wells are numbered in natural order; the field is fieldID when it is the only other group and every value is a whole number from 1 to 999, else fields are numbered within their well; without a wellID group the fields fill A01 to 999 and then move on to A02. The time is timeID when every value is a whole number from 1, else numbered; 1 without one.

Parameters:

keys – set keys, as check_sets() or detect_sets() return them.

Returns:

{key: SetName}.

Raises:

ValueError – when two sets would get one name.

spacr.channel_sorting.natural_key(text) → tuple[source]

Sort key that orders img2 before img10.

Parameters:

text – a name.

Returns:

a tuple comparing numbers numerically and text case-blind.

spacr.channel_sorting.parse_names(names: Sequence[str], pattern: str | None, channels: Dict[str, int] | None = None) → List[ParsedName][source]

Read each name’s channel and set through pattern.

The channel of a name the user assigned (channels) is that assignment; otherwise it is its chanID value’s position among all chanID values, naturally sorted, from 1. The set key is every other named group, in the pattern’s order. A name the pattern does not match has no set.

Parameters:
  • names – image file names, or absolute paths for images outside the folder; the regex reads the file name alone.

  • pattern – the regex, matched from the start of each name; empty or None matches nothing.

  • channels – {name: channel} from the selection strategy.

Returns:

one ParsedName per name, in order.

spacr.channel_sorting.set_label(key: SetKey | None) → str[source]

A set key as short text, wellID=A01 fieldID=2.

Parameters:

key – the set key.

spacr.channel_sorting.split_extension(name: str) → Tuple[str, str][source]

Split name into stem and extension, keeping .ome.tif whole.

Parameters:

name – a file name.

Returns:

(stem, extension); the extension keeps its dot.

spacr.channel_sorting.thumbnail(path: str, size: int = 96) → numpy.ndarray | None[source]

A small 8-bit grey preview of an image, contrast-stretched.

Parameters:
  • path – the image file.

  • size – the longest side, in pixels.

Returns:

a 2-D uint8 array, or None when it cannot be read.

spacr.channel_sorting.unused_folder(parent: str, name: str) → str[source]

parent/name, or name_2, name_3… when that exists.

Parameters:
  • parent – the folder to create in.

  • name – the name wanted.

spacr.channel_sorting.well_name(index: int) → str[source]

The index-th well of a 384-well plate, row by row: 0 is A01.

Parameters:

index – 0-based, below WELLS_PER_PLATE.

spacr.channel_sorting.yokogawa_name(name: SetName, channel: int, extension: str = '.tif') → str[source]

The Yokogawa file name of one channel of one set.

Built by spacr.convert.target_name(), so it is the exact form spacr.utils._get_regex('cellvoyager', ...) parses.

Parameters:
  • name – the set’s place.

  • channel – 1-based channel.

  • extension – the image’s own extension, kept.

Nested helpers

_match_columns.plane(path: str | None)

A file’s 2-D size, for comparing an image with a mask.

Parameters:

path – a file, or None.

spacr/channel_sorting.py:1314

_regexes_for_family.complete(identity: Sequence[int], labels: Sequence) → bool

Whether identity columns give each label exactly once per set.

Parameters:
  • identity – the columns that identify a set.

  • labels – the channel label of each name.

spacr/channel_sorting.py:647

_regexes_for_family.function_of(position: int, labels: Sequence) → bool

Whether column position is fixed by labels.

Parameters:
  • position – a token column.

  • labels – one label per name.

spacr/channel_sorting.py:635

infer_regex.channel_count(pattern: str) → int | None

How many channels pattern gives, when it puts every name in a set.

Parameters:

pattern – a candidate regex.

Returns:

the channel count, or None when the pattern fails.

spacr/channel_sorting.py:762

name_sets.identity(d: Dict[str, str]) → tuple

The groups that tell fields apart within a well.

Groups with one value across every set (the L01/A01 of a Yokogawa name) tell nothing apart and are left out.

Parameters:

d – one set’s groups.

spacr/channel_sorting.py:1038