spacr.diameter

Propose Cellpose diameter values from blob statistics, without Cellpose.

Why this exists

diameter is the most consequential Cellpose knob spaCR exposes and it is the one users guess at. Under Cellpose 4 (cpsam) the constructor argument diam_mean= is ignored, but CellposeModel.eval(diameter=...) is very much alive: it rescales the input by 30.0 / diameter so objects land near the network’s ~30 px working size. Feed it a value two-fold off and every downstream mask, count and measurement inherits the error. So a defensible starting number is worth a few seconds of arithmetic.

Design constraint: this module must be usable before committing to a segmentation run, which means it may not cost what a segmentation run costs. It therefore imports no torch and no cellpose — that is a tested property, not an aspiration (see tests/test_diameter_estimator.py). It reuses spacr.validate for filename-metadata parsing, which is the other deliberately dependency-light module in the package (spacr.utils, where _get_regex and _extract_filename_metadata live, imports torch and cellpose at module scope and so cannot be touched from here). The regexes in spacr.validate.METADATA_REGEXES mirror spacr.utils._get_regex, and the channel ordering used below mirrors spacr.io._rename_and_organize_image_files, which concatenates one plane per distinct chanID in sorted order — so channel index i is the i-th sorted chanID, exactly as cell_channel and friends mean it.

Method

For each requested object channel, on each sampled field:

  1. Flatten. Denoise with a 1 px Gaussian, then subtract a heavily smoothed copy (sigma = max(32, min(H, W) / 4)) to remove illumination gradients. The sigma is deliberately far larger than any plausible object so that flattening removes vignetting without eating the objects — the opposite trade-off (a tight rolling ball) shrinks what it is trying to measure.

  2. Reject noise. Compare the structural amplitude of the flattened plane (p99 - p30) against the pixel-level noise scale, measured as 1.4826 * MAD(img - gaussian(img, 1)) on the raw plane. A pure-noise plane scores below 1; a plane with real objects scores in the tens. Below min_snr the field is discarded rather than thresholded, because Otsu will happily bisect pure noise and hand back a confident-looking number.

  3. Threshold and label. Otsu, fill holes, label, drop components that touch the image border (truncated, so their size is a lie) and components that are absurd (equivalent diameter below min_object_diameter, or area above max_object_fraction of the field). Sizes are equivalent diameters, 2 * sqrt(area / pi).

  4. Drop specks, then take the median. A textured or punctate stain (a mitochondrial cell stain, a parasite marker) thresholds into a few whole objects plus hundreds of 4-10 px specks, and a plain median counts every speck as an object: on spaCR’s own test plate that put the cell at 8 px and the pathogen at 11 px against 175 px and 32 px in the curated masks, while the nucleus, whose stain is solid, came out right at 91 px. So the typical object is first located by the stained area it carries (the area-weighted median, which specks cannot move because they hold almost no area), everything narrower than SPECK_FRACTION of it is dropped, and the characteristic size is the median of what remains. Small components that together hold SPECK_MAX_AREA_SHARE of the stained area or more are a second population, not specks, and stay.

  5. Cross-check by distance transform. Step 3 has one dominant failure mode: a confluent monolayer fuses into a single component, that component touches the border and is dropped, and the estimate is then computed from whatever debris survived — biased low, and silently. So the Euclidean distance transform of the (unfilled) foreground is computed as well, its local maxima are taken as one seed per object (two passes: a coarse pass sets the suppression radius for the refined pass), and a watershed on -EDT splits the fused foreground back into objects whose equivalent diameters are measured the same way.

Both estimates are always computed, and choosing between them is where the care goes. The watershed estimate is taken when the transform resolved something and either the threshold path kept nothing at all, or the field is dense enough for fusion to be the explanation (foreground at or above fused_fraction) and the transform resolves at least five times as many objects. A count disagreement alone is therefore never fusion — a hollow, membrane-only object is one correct component by area but shatters into dozens of arcs under the distance transform, so a bare ratio test would discard the right answer in favour of the wall thickness, which is the same silent collapse entered from the other side. A high foreground fraction alone is not fusion either — a dense but well-separated field can reach 30% foreground and still be measured correctly by thresholding — so on its own it only brings the confidence down while the threshold estimate stands. An empty threshold result is taken as fusion on its own, with no density check, because no threshold estimate is left to stand; if the transform resolved nothing either, no estimate is returned for that channel. When fusion is not declared and the two estimates disagree by more than 1.5-fold, the confidence is downgraded and the note says so, because that disagreement is the honest signal that neither number should be trusted blind.

Public API

DiameterEstimate

One proposal, with its plausible range, provenance and confidence.

estimate_diameters(src, channels, n_fields=5, ...)

{object_type: DiameterEstimate}.

format_estimates(estimates)

Human-readable block, ready to print.

channels_from_settings(settings)

Pull cell_channel / nucleus_channel / … out of a settings dict.

Where a user meets this

spacr.qt.prerun.DiameterPanel sits above the Run row on the Mask screen: it reads the channels off the settings form with channels_from_settings(), calls estimate_diameters() on a worker thread, and shows one row per object type carrying the proposal and its evidence — the 10th-90th percentile range, how many objects it was pooled from, how many fields contributed, the method and the confidence. A proposal without those is just a different guess. Nothing is written into <object>_diameter until the user presses Use, and DiameterEstimate.usable is checked first, so the NaN an unmeasurable channel returns cannot reach a settings field.

Classes

DiameterEstimate

One proposed diameter, with everything needed to disbelieve it.

Functions

channels_from_settings(→ Dict[str, int])

Extract {object_type: channel_index} from a spaCR settings dict.

estimate_diameters(→ Dict[str, DiameterEstimate])

Propose a Cellpose diameter per object type from blob statistics.

format_estimates(→ str)

Render estimate_diameters() output as a printable block.

Module Contents

class spacr.diameter.DiameterEstimate[source]

One proposed diameter, with everything needed to disbelieve it.

Parameters:
  • object_type – 'cell', 'nucleus', 'pathogen' or 'organelle'.

  • diameter – proposed value in pixels — the median object diameter across the sampled fields. float('nan') when nothing usable was found; see usable. It is never a fabricated fallback.

  • low – 10th percentile of the measured object diameters, in pixels.

  • high – 90th percentile of the measured object diameters, in pixels.

  • n_objects – how many objects the estimate was pooled from.

  • n_fields – how many fields actually contributed (which can be fewer than requested).

  • method – which measurement produced diameter — 'threshold_otsu', 'watershed_edt', or 'none'.

  • confidence – 'high', 'medium' or 'low'.

  • note – why the number is what it is, which confidence downgrades applied, and what to check if it looks wrong.

__str__() → str[source]

Return the estimate and interval, or the reason none is usable.

Returns:

Compact human-readable diameter-estimate summary.

property usable: bool[source]

False when no defensible number could be produced.

Callers must check this before writing diameter into a settings dict: an unusable estimate carries NaN precisely so that a fabricated number cannot leak into a run.

spacr.diameter.channels_from_settings(settings: Dict[str, Any]) → Dict[str, int][source]

Extract {object_type: channel_index} from a spaCR settings dict.

Reads cell_channel, nucleus_channel, pathogen_channel and organelle_channel; object types whose channel is None or unparseable are absent from the result.

Parameters:

settings – a spaCR settings dict (or anything dict-like).

Returns:

mapping suitable for estimate_diameters().

spacr.diameter.estimate_diameters(src: Any, channels: Dict[str, Any], n_fields: int = 5, *, metadata_type: Any = 'cellvoyager', custom_regex: Any = None, random_state: int | None = None, max_pixels: int = 2250000, min_snr: float = 4.0, fused_fraction: float = 0.35, min_object_diameter: float = 4.0, max_object_fraction: float = 0.25, verbose: bool = False) → Dict[str, DiameterEstimate][source]

Propose a Cellpose diameter per object type from blob statistics.

Samples n_fields fields spread across src (never the first N — see _sample_indices()), measures characteristic object size in each requested channel by thresholding and by a distance-transform watershed, and pools the two into one proposal per object type. Cellpose is never loaded; neither is torch.

Parameters:
  • src – a source folder, or a list of them. Merged stack/ / merged/ .npy arrays are used when present, otherwise raw acquisition files are grouped by the metadata regex.

  • channels – {object_type: channel_index}, e.g. {'cell': 2, 'nucleus': 0}. Values that are not integers are skipped. channels_from_settings() builds this from a settings dict.

  • n_fields – how many fields to sample. More is slower and steadier.

  • metadata_type – 'cellvoyager' / 'cq1' / 'auto' / 'custom'; only consulted when raw files must be parsed.

  • custom_regex – named-group regex used when metadata_type is 'custom'.

  • random_state – None (default) samples on an even stride, which is deterministic; an int draws a reproducible random sample instead.

  • max_pixels – fields larger than this are centre-cropped, which loses peripheral objects but does not bias their measured size.

  • min_snr – minimum ratio of structural amplitude to pixel noise for a field to be measured at all. Below it the field is discarded rather than thresholded, so a background-only channel yields no number.

  • fused_fraction – foreground fraction at or above which the field is called confluent, which downgrades the confidence and is recorded in the note. It does not by itself switch the measurement method; the distance-transform estimate takes over only when it resolves substantially more objects than plain labelling did.

  • min_object_diameter – components thinner than this many pixels are discarded as debris.

  • max_object_fraction – components covering more than this fraction of a field are discarded as fused blobs rather than measured.

  • verbose – print each estimate as it is produced.

Returns:

{object_type: DiameterEstimate} for every object type given an integer channel. Entries whose DiameterEstimate.usable is False carry NaN — check it before writing the value into settings.

spacr.diameter.format_estimates(estimates: Dict[str, DiameterEstimate]) → str[source]

Render estimate_diameters() output as a printable block.

Parameters:

estimates – the mapping returned by estimate_diameters().

Returns:

a multi-line string: one table row per object type carrying the proposed diameter, its plausible range, object and field counts, confidence and method, followed by the note behind each row.

Nested helpers

format_estimates._line(cells: Sequence[str]) → str

Pad one report row to the captured column widths.

Parameters:

cells – cell text in the same order as header.

Returns:

two-space-indented, column-aligned text with trailing padding removed.

spacr/diameter.py:1067