spacr.qt.synthetic

Synthetic datasets + saved settings for exercising every pipeline app.

The goal: give a developer (or a bug reporter) a one-line way to generate a demo folder that flows cleanly through every spacr pipeline — mask, measure, crop, classify, timelapse, map_barcodes — plus a matching settings CSV that plugs into the “Import settings…” button on each app screen.

Everything is reverse-engineered from what the pipelines actually consume:

  • Filenames match the cellvoyager regex in spacr.utils._get_regex('.tif', 'cellvoyager'):

    <plateID>_<wellID>_T<timeID>F<fieldID>L<laserID>A<AID>Z<sliceID>C<chanID>.tif
    
  • Channels are laid out in the order every mask default expects:

    C0 = nucleus, C1 = cell, C2 = pathogen, C3 = organelle
    
  • Images are 16-bit uint16. Every channel of a field is drawn from one shared cell layout, so the nucleus really is inside the cell and the pathogen really is inside the same cell — the relationships measure_crop goes looking for when it links objects.

  • Measure/crop demos ship the merged/*.npy stacks measure_crop actually reads (image planes first, then the label-mask planes), not a stand-in.

  • Settings CSVs are written in the two-column Key,Value format, the first of the header pairs spacr.qt.screens.app_screen.AppScreen ._load_settings_csv tries, so the “Import settings…” button restores every value into the form. Reading one straight from Python means naming the columns — load_settings(path, setting_key="Key", setting_value="Value") — because spacr.utils.load_settings defaults to the other spelling (setting_key/setting_value, what spacr.io.save_settings_to_db writes) and raises rather than guessing.

Every generator is reproducible: identical inputs give byte-identical output on any machine. That is not decoration. Two people comparing “the demo fails here” have to be looking at the same pixels, and the previous seeding (hash((well, field, time, chan))) was salted by PYTHONHASHSEED, so it changed on every interpreter start.

Public API:

generate_mask_demo(dst, ...) -> DemoLayout
generate_measure_demo(dst, ...) -> DemoLayout
generate_crop_demo(dst, ...) -> DemoLayout
generate_classify_demo(dst, ...) -> DemoLayout
generate_timelapse_demo(dst, ...) -> DemoLayout
generate_map_barcodes_demo(dst, ...) -> DemoLayout
save_settings_csv(dst, settings) -> Path
demo_settings(app_key, src, channels=None) -> Dict[str, Any]

CLI:

python -m spacr.qt.synthetic mask /tmp/demo
python -m spacr.qt.synthetic all  /tmp/demo

Classes

DemoLayout

What a demo generator produced. Absolute paths only.

Functions

barcode_pool(→ List[str])

Return n distinct synthetic barcodes of length bases.

cellvoyager_filename(→ str)

Return a filename matching:

demo_settings(→ Dict[str, Any])

Return a spacr settings dict tailored for the demo dataset

generate_barcode_csv(→ pathlib.Path)

Write a name,sequence barcode CSV.

generate_classify_demo() → DemoLayout)

Classify wants PNG single-object crops + a measurements.db

generate_crop_demo(, fields, channels)

Same dataset as measure — Crop is measure with save_png on, and

generate_map_barcodes_demo(→ DemoLayout)

Populate dst with a self-contained map_barcodes demo:

generate_mask_demo(, fields, channels)

Populate dst with a folder that runs cleanly through the Mask

generate_measure_demo(, fields, channels)

Measure consumes what Mask produces: a merged/ folder of

generate_synthetic_fastq(→ List[pathlib.Path])

Write a gzip-compressed synthetic FASTQ pair carrying known barcodes.

generate_timelapse_demo(, fields, times, channels)

Timelapse needs multi-T frames per (well, field) so tracking

main(→ int)

Generate one (or every) demo dataset via the python -m CLI.

save_settings_csv(→ pathlib.Path)

Write settings in the two-column Key,Value format that

synthetic_read(→ str)

Build one 150-base read carrying a (column, gRNA, row) triplet.

Module Contents

class spacr.qt.synthetic.DemoLayout[source]

What a demo generator produced. Absolute paths only.

merged_files replaces the old mask_files: nothing in spaCR reads a folder of standalone label tiffs, and the measure/crop demos now ship the merged/*.npy stacks measure_crop actually opens — the label planes are the trailing planes of those arrays.

Parameters:
  • src – absolute path of the demo folder, the value the demo’s settings use as src.

  • image_dir – absolute path of the folder holding the generated raw images; the same as src except for demos that put them in a data subfolder.

spacr.qt.synthetic.barcode_pool(n: int, length: int, seed: int = 0) → List[str][source]

Return n distinct synthetic barcodes of length bases.

Parameters:
  • n – how many to draw.

  • length – barcode length in bases.

  • seed – RNG seed — the same seed always returns the same pool.

Returns:

list of n unique uppercase DNA strings.

spacr.qt.synthetic.cellvoyager_filename(plate: str = 'plate1', well: str = 'A01', time: int = 1, field: int = 1, laser: int = 1, a: int = 1, slice_: int = 1, chan: int = 1, ext: str = 'tif') → str[source]

Return a filename matching: <plateID>_<wellID>_T<timeID>F<fieldID>L<laserID>A<AID>Z<sliceID>C<chanID>.<ext>

spacr.qt.synthetic.demo_settings(app_key: str, src: str, channels: Sequence[int] | None = None) → Dict[str, Any][source]

Return a spacr settings dict tailored for the demo dataset generated by generate_<app>_demo.

Parameters:
  • app_key – which app the settings are for.

  • src – the demo folder (for map_barcodes, the demo root — its barcode CSVs are resolved relative to it).

  • channels – the acquisition channels the dataset actually holds. Defaults to all four. Every *_channel and *_mask_dim key is derived from it, so a two-channel demo cannot advertise a pathogen channel it never acquired.

Values are the minimum needed to make the pipeline flow — real users will tweak thresholds + channel numbers to fit their data.

spacr.qt.synthetic.generate_barcode_csv(dst: pathlib.Path, names: Sequence[str], sequences: Sequence[str]) → pathlib.Path[source]

Write a name,sequence barcode CSV.

This is the format spacr.sequencing.map_sequences_to_names reads — it requires both columns by name and rejects duplicate sequences. (The demo used to ship a FASTA, which that function cannot read at all.)

Parameters:
  • dst – output .csv path.

  • names – barcode names, aligned with sequences.

  • sequences – barcode sequences.

Returns:

the resolved dst path.

Raises:

ValueError – when names and sequences do not hold the same number of entries. This is a count of rows, not a comparison of base counts — the previous wording said the opposite, in a module whose subject is DNA of a declared size, and the two CSVs this writes for a demo folder deliberately carry different barcode lengths (21 for gRNAs, 8 for wells), so a base-length check would be wrong as well as unimplemented. The count check earns its stop because the zip below halts at the shorter list: an unequal pair would silently drop the tail and hand map_sequences_to_names a table missing barcodes the reads actually carry, and every read of those guides would then map to nothing with no line in the log to say why.

spacr.qt.synthetic.generate_classify_demo(dst: pathlib.Path, n_crops: int = 64, plate: str = 'plate1', wells: Sequence[str] = ('A01', 'A02', 'A03', 'A04')) → DemoLayout[source]

Classify wants PNG single-object crops + a measurements.db with a png_list table + an annotate column carrying class labels for training/testing.

This is a hand-built stand-in for a measured plate, not a replica of one. Two of the three things that matter match measure_crop; the third does not, and the docstring used to claim all three did.

Crop names match. A real crop is <file_name>_<cell_id>.png where file_name is the merged stack’s <plate>_<well>_<field>_<time> (spacr.utils._generate_names()) — e.g. plate1_A01_1_1_1.png. That is exactly what this writes, and exactly what spacr.utils._map_wells_png parses plate/row/column/field back out of.

The cell_png leaf matches. measure_crop appends f"{crop_mode}_png/" to the folder, so a real cell crop does live under cell_png/ — which is what spacr.io.generate_training_dataset()’s png_path.str.contains(png_type) filter (png_type='cell_png') needs to see. Crops written flat as data/crop_000.png, the layout this replaced, were all filtered away and the run died on “got 0 classes”.

The folder above it does not match. measure_crop buckets every crop by what it contains first: data/<single|multiple|no>_nucleus/<single_pathogen|multiple_pathogens|uninfected>/<plate>_<well>/cell_png/. This demo writes data/<plate>_<well>/cell_png/ with no bucket folders — nothing downstream of png_type reads them, and inventing an infection status per synthetic crop would be a fiction the pixels do not support.

The png_list columns do not match either. spacr.utils.filepaths_to_database() writes png_path, file_name, plateID, rowID, columnID, fieldID, prcfo, cell_id, with the tokenised values rowID='r1'/columnID='c1'/fieldID='f1'. This table carries png_path, plateID, wellID, rowID, columnID, fieldID, timeID, label in plain form, plus the annotate column — which is the point: annotate is what dataset_mode='annotation' selects classes on, and a measure run never writes one. A human does, in the Annotate screen.

Parameters:
  • dst – destination folder.

  • n_crops – total number of crops, spread evenly over the wells.

  • plate – plate ID baked into the crop names and png_list.

  • wells – well IDs to spread the crops over; at least two, so the classifier can hold a whole well out.

Returns:

DemoLayout.

spacr.qt.synthetic.generate_crop_demo(dst: pathlib.Path, plate: str = 'plate1', wells: Iterable[str] = ('A01', 'A02'), fields: int = 2, channels: Iterable[int] = (0, 1, 2, 3)) → DemoLayout[source]

Same dataset as measure — Crop is measure with save_png on, and writes PNG crops into per-object folders alongside the DB.

Builds the dataset directly rather than by calling generate_measure_demo(). Chaining them wrote settings_measure.csv first and then only reassigned layout.settings_csv, leaving two settings files in one folder — and the one named after the folder’s own demo was the one that turns PNG crops off. A folder holds exactly one settings_*.csv now, so “Import settings…” cannot pick the wrong run.

Parameters:
  • dst – destination folder.

  • plate – plate ID baked into every filename.

  • wells – well IDs to emit.

  • fields – fields per well.

  • channels – acquisition channels to render.

Returns:

DemoLayout whose settings_csv is settings_crop.csv.

spacr.qt.synthetic.generate_map_barcodes_demo(dst: pathlib.Path, n_barcodes: int = 12, n_reads: int = 5000, seed: int = 0, n_rows: int = 4, n_columns: int = 6) → DemoLayout[source]

Populate dst with a self-contained map_barcodes demo:

dst/
  barcodes/
    grna.csv            # ← N gRNA barcodes, name,sequence
    row.csv             # ← row (plate-row) barcodes
    column.csv          # ← column barcodes
  demo_R1_001.fastq.gz  # ← reads carrying those barcodes
  demo_R2_001.fastq.gz
  settings_map_barcodes.csv

The FASTQs sit in dst itself, not a fastq/ subfolder: the pipeline’s src is listed flat for *.fastq.gz (spacr.io.parse_gz_files), so a subfolder means zero samples found and a run that exits having written nothing.

Parameters:
  • dst – destination folder.

  • n_barcodes – number of unique gRNA barcodes to plant.

  • n_reads – approximate total number of reads to emit.

  • seed – RNG seed for reproducibility.

  • n_rows – number of plate-row barcodes.

  • n_columns – number of plate-column barcodes.

Returns:

DemoLayout describing the emitted files.

spacr.qt.synthetic.generate_mask_demo(dst: pathlib.Path, plate: str = 'plate1', wells: Iterable[str] = ('A01', 'A02'), fields: int = 2, channels: Iterable[int] = (0, 1, 2, 3)) → DemoLayout[source]

Populate dst with a folder that runs cleanly through the Mask app. Layout:

dst/
  <plateID>_<wellID>_T01F<field>L01A01Z01C<chan>.tif
  settings_mask.csv
Parameters:

dst – destination folder; made absolute and created if absent.

spacr.qt.synthetic.generate_measure_demo(dst: pathlib.Path, plate: str = 'plate1', wells: Iterable[str] = ('A01', 'A02'), fields: int = 2, channels: Iterable[int] = (0, 1, 2, 3)) → DemoLayout[source]

Measure consumes what Mask produces: a merged/ folder of .npy stacks whose trailing planes are the label masks.

We pre-build those stacks so a user can jump straight into Measure without a GPU. Before this, the demo wrote a masks/ folder of per-file tiffs and an empty measurements.db — neither of which any pipeline reads — and measure’s pre-flight rejected the folder outright with “no merged folder for measure”.

Note

The organelle plane this writes is measured into nothing, and the defect is not in this module. Plane organelle_mask_dim of every merged stack carries 64 real labels, but a measure run over this folder writes cell/nucleus/pathogen/cytoplasm and no organelle table at all (verified: 4 fields → cell 64, nucleus 64, pathogen 71, cytoplasm 64, organelle absent).

All four organelle writes in spacr.measure._measure_crop_core are gated on settings.get('summarize_organelles_by') is not None, and spacr.settings.get_measure_crop_settings — the defaults every measure run is canonicalised through — never sets that key.

That much was already known. What was wrong was the remedy: “default it for measure the way the Mask app does” does not fix this, and cannot, for two reasons that have to be fixed together.

  1. set_default_settings_preprocess_generate_masks defaults it to the string 'cell', and measure.py tests it with "organelle" in settings['summarize_organelles_by'] — a substring test when the value is a str. Running this demo with summarize_organelles_by='cell' gives cell_organelle_summary (16 rows/field) and still no organelle table. Only a value containing 'organelle' writes the per-organelle table (['cell', 'organelle'] → organelle 64 rows/field, verified).

  2. A list cannot be shipped today: spacr.settings.expected_types declares 'summarize_organelles_by': str, so spacr.validate.validate_settings rejects ['cell', 'organelle'] with “is a list, but str is expected” — a hard pre-flight error on a demo that must load clean. The tooltip and spacr.gui_utils both describe it as a list, and spacr.external_masks builds one; only the type table disagrees.

So the demo deliberately omits the key, and the wiring needed elsewhere is: widen expected_types['summarize_organelles_by'] to (str, list, type(None)); make get_measure_crop_settings default it to ['cell', 'organelle'] (safe — every write is separately gated on organelle_mask_dim is not None, which defaults to None); and add it to the measure section of spacr/qt/screens/settings_model.py so the Measure form can hold it — without a widget, apply_settings_dict drops it and collect() never emits it, so a CSV key would change what a CLI run measures and nothing about a GUI run.

Parameters:

dst – destination folder; made absolute and created if absent.

spacr.qt.synthetic.generate_synthetic_fastq(dst_dir: pathlib.Path, grnas: Sequence[str], rows: Sequence[str], columns: Sequence[str], n_reads: int = 5000, seed: int = 0, sample: str = FASTQ_SAMPLE, paired: bool = True) → List[pathlib.Path][source]

Write a gzip-compressed synthetic FASTQ pair carrying known barcodes.

Every read is one (column, gRNA, row) triplet in the frame the shipped barcode-mapping defaults parse, so unique_combinations.csv comes back with the planted wells and guides in it. Reads are spread evenly over the rows x columns wells, and within a well over the gRNAs with a skew, because a real screen has a handful of abundant guides and a long tail.

Parameters:
  • dst_dir – folder to write into.

  • grnas – gRNA barcode sequences (21 bases each).

  • rows – row barcode sequences (8 bases each).

  • columns – column barcode sequences (8 bases each).

  • n_reads – approximate total number of reads; the real total is rounded down to a whole number of reads per well.

  • seed – RNG seed for reproducible read pools.

  • sample – sample name; the files are <sample>_R1_001.fastq.gz (and _R2_).

  • paired – also write the R2 mate. R2 is the exact reverse complement of R1 — a perfectly overlapping pair — which is what spacr.sequencing’s paired path reduces to after it reverse-complements R2 and takes the per-base consensus.

Returns:

the written paths, R1 first.

Raises:

ValueError – when any barcode list is empty.

spacr.qt.synthetic.generate_timelapse_demo(dst: pathlib.Path, plate: str = 'plate1', wells: Iterable[str] = ('A01',), fields: int = 1, times: int = 8, channels: Iterable[int] = (0, 1)) → DemoLayout[source]

Timelapse needs multi-T frames per (well, field) so tracking has something to lock onto. Same cellvoyager naming, just with T01..T<N>, and every frame holds the same cells drifting a couple of pixels rather than a fresh random field.

Only the nucleus and cell channels are acquired, and the settings say so: the base settings used to advertise channels=[0,1,2,3] plus a pathogen_channel/organelle_channel this dataset never had, which pre-flight rejected with two hard errors before the run could start.

Note

The dataset and settings this writes clear pre-flight, but the Timelapse pipeline still cannot consume them, and neither defect is in this module. Both were reproduced on this demo and both fixes were proved by patching the two functions at run time; with the pair applied the demo completes and writes 8 merged stacks, per-channel movies, a track-overlay GIF and a 16-track tracks/*.csv.

  1. spacr.io._rename_and_organize_image_files names its stack files <plate>_<well>_<field>.npy when timelapse=True — dropping the timeID and max-projecting every timepoint of a field into one array — while spacr.io._generate_time_lists groups on <plate>_<well>_<field>_<time>.npy and skips anything with fewer than four underscore-separated parts. An 8-frame field becomes one stack/plate1_A01_1.npy, _generate_time_lists returns [], no *_norm_timelapse.npz is written, no masks are generated, and preprocess_generate_masks dies in _pivot_counts_table on no such table: object_counts. Emitting the timeID in both branches (the non-timelapse spelling is already exactly what _generate_time_lists parses) is the fix.

  2. Past that, spacr.object.generate_cellpose_masks_sam hands spacr.timelapse._trackpy_track_cells a list of 2-D frames, and the tracking chain indexes it as an array: _track_by_iou does masks.shape[0] and _relabel_masks_based_on_tracks does np.zeros(masks.shape, …), both AttributeError: 'list' object has no attribute 'shape'. In the timelapse_mode='iou' path this demo asks for, the first one is swallowed by the except Exception retry loop in _facilitate_trackin_with_adaptive_removal, which then shrinks the search range 100 times and reports Failed to track after 100 attempts — a message about displacement for a bug about a type. Coercing once at the top of _trackpy_track_cells (masks = np.asarray(masks)) clears both.

Parameters:

dst – destination folder; made absolute and created if absent.

spacr.qt.synthetic.main(argv: list[str] | None = None) → int[source]

Generate one (or every) demo dataset via the python -m CLI.

Parameters:

argv – optional argv list; defaults to sys.argv[1:].

Returns:

process exit code (0 on success).

spacr.qt.synthetic.save_settings_csv(path: pathlib.Path, settings: Dict[str, Any]) → pathlib.Path[source]

Write settings in the two-column Key,Value format that spacr.utils.load_settings reads.

Parameters:
  • path – CSV file to write; made absolute, its parent folder is created, and an existing file is overwritten.

  • settings – settings to write, one Key,Value row each; None is written as an empty value and everything else with str().

spacr.qt.synthetic.synthetic_read(column_barcode: str, grna: str, row_barcode: str, prefix: str = SEQ_READ_PREFIX) → str[source]

Build one 150-base read carrying a (column, gRNA, row) triplet.

The layout is the one documented above SEQ_TARGET; a read built here is recovered exactly by the shipped regex / target_sequence / offset_start / window_length defaults.

Parameters:
  • column_barcode – 8-base column barcode.

  • grna – 21-base gRNA barcode.

  • row_barcode – 8-base row barcode.

  • prefix – stagger placed before the anchor window.

Returns:

a 150-base read.

Raises:

ValueError – when a barcode is the wrong length — a silently mis-sized barcode would shift every downstream field by that many bases and map to nothing, which is far harder to see than a stop.

Nested helpers

_synth_field._peak() → float

One object’s peak intensity, jittered around the nominal.

spacr/qt/synthetic.py:345