spacr.qt.synthetic¶
Synthetic datasets + saved settings for exercising every pipeline app.
The goal: give a developer (or a bug reporter) a one-line way to generate a demo folder that flows cleanly through every spacr pipeline — mask, measure, crop, classify, timelapse, map_barcodes — plus a matching settings CSV that plugs into the “Import settings…” button on each app screen.
Everything is reverse-engineered from what the pipelines actually consume:
Filenames match the cellvoyager regex in
spacr.utils._get_regex('.tif', 'cellvoyager'):<plateID>_<wellID>_T<timeID>F<fieldID>L<laserID>A<AID>Z<sliceID>C<chanID>.tif
Channels are laid out in the order every mask default expects:
C0 = nucleus, C1 = cell, C2 = pathogen, C3 = organelle
Images are 16-bit uint16. Every channel of a field is drawn from one shared cell layout, so the nucleus really is inside the cell and the pathogen really is inside the same cell — the relationships measure_crop goes looking for when it links objects.
Measure/crop demos ship the
merged/*.npystacks measure_crop actually reads (image planes first, then the label-mask planes), not a stand-in.Settings CSVs are written in the two-column
Key,Valueformat, the first of the header pairsspacr.qt.screens.app_screen.AppScreen ._load_settings_csvtries, so the “Import settings…” button restores every value into the form. Reading one straight from Python means naming the columns —load_settings(path, setting_key="Key", setting_value="Value")— becausespacr.utils.load_settingsdefaults to the other spelling (setting_key/setting_value, whatspacr.io.save_settings_to_dbwrites) and raises rather than guessing.
Every generator is reproducible: identical inputs give byte-identical
output on any machine. That is not decoration. Two people comparing
“the demo fails here” have to be looking at the same pixels, and the
previous seeding (hash((well, field, time, chan))) was salted by
PYTHONHASHSEED, so it changed on every interpreter start.
Public API:
generate_mask_demo(dst, ...) -> DemoLayout
generate_measure_demo(dst, ...) -> DemoLayout
generate_crop_demo(dst, ...) -> DemoLayout
generate_classify_demo(dst, ...) -> DemoLayout
generate_timelapse_demo(dst, ...) -> DemoLayout
generate_map_barcodes_demo(dst, ...) -> DemoLayout
save_settings_csv(dst, settings) -> Path
demo_settings(app_key, src, channels=None) -> Dict[str, Any]
CLI:
python -m spacr.qt.synthetic mask /tmp/demo
python -m spacr.qt.synthetic all /tmp/demo
Classes¶
What a demo generator produced. Absolute paths only. |
Functions¶
|
Return |
|
Return a filename matching: |
|
Return a spacr settings dict tailored for the demo dataset |
|
Write a |
|
Classify wants PNG single-object crops + a |
|
Same dataset as measure — Crop is measure with |
|
Populate |
|
Populate |
|
Measure consumes what Mask produces: a |
|
Write a gzip-compressed synthetic FASTQ pair carrying known barcodes. |
|
Timelapse needs multi-T frames per (well, field) so tracking |
|
Generate one (or every) demo dataset via the |
|
Write |
|
Build one 150-base read carrying a (column, gRNA, row) triplet. |
Module Contents¶
- class spacr.qt.synthetic.DemoLayout[source]¶
What a demo generator produced. Absolute paths only.
merged_filesreplaces the oldmask_files: nothing in spaCR reads a folder of standalone label tiffs, and the measure/crop demos now ship themerged/*.npystacks measure_crop actually opens — the label planes are the trailing planes of those arrays.- Parameters:
src – absolute path of the demo folder, the value the demo’s settings use as
src.image_dir – absolute path of the folder holding the generated raw images; the same as
srcexcept for demos that put them in adatasubfolder.
- spacr.qt.synthetic.barcode_pool(n: int, length: int, seed: int = 0) List[str][source]¶
Return
ndistinct synthetic barcodes oflengthbases.- Parameters:
n – how many to draw.
length – barcode length in bases.
seed – RNG seed — the same seed always returns the same pool.
- Returns:
list of
nunique uppercase DNA strings.
- spacr.qt.synthetic.cellvoyager_filename(plate: str = 'plate1', well: str = 'A01', time: int = 1, field: int = 1, laser: int = 1, a: int = 1, slice_: int = 1, chan: int = 1, ext: str = 'tif') str[source]¶
Return a filename matching: <plateID>_<wellID>_T<timeID>F<fieldID>L<laserID>A<AID>Z<sliceID>C<chanID>.<ext>
- spacr.qt.synthetic.demo_settings(app_key: str, src: str, channels: Sequence[int] | None = None) Dict[str, Any][source]¶
Return a spacr settings dict tailored for the demo dataset generated by
generate_<app>_demo.- Parameters:
app_key – which app the settings are for.
src – the demo folder (for
map_barcodes, the demo root — its barcode CSVs are resolved relative to it).channels – the acquisition channels the dataset actually holds. Defaults to all four. Every
*_channeland*_mask_dimkey is derived from it, so a two-channel demo cannot advertise a pathogen channel it never acquired.
Values are the minimum needed to make the pipeline flow — real users will tweak thresholds + channel numbers to fit their data.
- spacr.qt.synthetic.generate_barcode_csv(dst: pathlib.Path, names: Sequence[str], sequences: Sequence[str]) pathlib.Path[source]¶
Write a
name,sequencebarcode CSV.This is the format spacr.sequencing.map_sequences_to_names reads — it requires both columns by name and rejects duplicate sequences. (The demo used to ship a FASTA, which that function cannot read at all.)
- Parameters:
dst – output
.csvpath.names – barcode names, aligned with
sequences.sequences – barcode sequences.
- Returns:
the resolved
dstpath.- Raises:
ValueError – when
namesandsequencesdo not hold the same number of entries. This is a count of rows, not a comparison of base counts — the previous wording said the opposite, in a module whose subject is DNA of a declared size, and the two CSVs this writes for a demo folder deliberately carry different barcode lengths (21 for gRNAs, 8 for wells), so a base-length check would be wrong as well as unimplemented. The count check earns its stop because thezipbelow halts at the shorter list: an unequal pair would silently drop the tail and handmap_sequences_to_namesa table missing barcodes the reads actually carry, and every read of those guides would then map to nothing with no line in the log to say why.
- spacr.qt.synthetic.generate_classify_demo(dst: pathlib.Path, n_crops: int = 64, plate: str = 'plate1', wells: Sequence[str] = ('A01', 'A02', 'A03', 'A04')) DemoLayout[source]¶
Classify wants PNG single-object crops + a
measurements.dbwith apng_listtable + anannotatecolumn carrying class labels for training/testing.This is a hand-built stand-in for a measured plate, not a replica of one. Two of the three things that matter match measure_crop; the third does not, and the docstring used to claim all three did.
Crop names match. A real crop is
<file_name>_<cell_id>.pngwherefile_nameis the merged stack’s<plate>_<well>_<field>_<time>(spacr.utils._generate_names()) — e.g.plate1_A01_1_1_1.png. That is exactly what this writes, and exactly whatspacr.utils._map_wells_pngparses plate/row/column/field back out of.The
cell_pngleaf matches. measure_crop appendsf"{crop_mode}_png/"to the folder, so a real cell crop does live undercell_png/— which is whatspacr.io.generate_training_dataset()’spng_path.str.contains(png_type)filter (png_type='cell_png') needs to see. Crops written flat asdata/crop_000.png, the layout this replaced, were all filtered away and the run died on “got 0 classes”.The folder above it does not match. measure_crop buckets every crop by what it contains first:
data/<single|multiple|no>_nucleus/<single_pathogen|multiple_pathogens|uninfected>/<plate>_<well>/cell_png/. This demo writesdata/<plate>_<well>/cell_png/with no bucket folders — nothing downstream ofpng_typereads them, and inventing an infection status per synthetic crop would be a fiction the pixels do not support.The
png_listcolumns do not match either.spacr.utils.filepaths_to_database()writespng_path, file_name, plateID, rowID, columnID, fieldID, prcfo, cell_id, with the tokenised valuesrowID='r1'/columnID='c1'/fieldID='f1'. This table carriespng_path, plateID, wellID, rowID, columnID, fieldID, timeID, labelin plain form, plus theannotatecolumn — which is the point:annotateis whatdataset_mode='annotation'selects classes on, and a measure run never writes one. A human does, in the Annotate screen.- Parameters:
dst – destination folder.
n_crops – total number of crops, spread evenly over the wells.
plate – plate ID baked into the crop names and png_list.
wells – well IDs to spread the crops over; at least two, so the classifier can hold a whole well out.
- Returns:
- spacr.qt.synthetic.generate_crop_demo(dst: pathlib.Path, plate: str = 'plate1', wells: Iterable[str] = ('A01', 'A02'), fields: int = 2, channels: Iterable[int] = (0, 1, 2, 3)) DemoLayout[source]¶
Same dataset as measure — Crop is measure with
save_pngon, and writes PNG crops into per-object folders alongside the DB.Builds the dataset directly rather than by calling
generate_measure_demo(). Chaining them wrotesettings_measure.csvfirst and then only reassignedlayout.settings_csv, leaving two settings files in one folder — and the one named after the folder’s own demo was the one that turns PNG crops off. A folder holds exactly onesettings_*.csvnow, so “Import settings…” cannot pick the wrong run.- Parameters:
dst – destination folder.
plate – plate ID baked into every filename.
wells – well IDs to emit.
fields – fields per well.
channels – acquisition channels to render.
- Returns:
DemoLayoutwhosesettings_csvissettings_crop.csv.
- spacr.qt.synthetic.generate_map_barcodes_demo(dst: pathlib.Path, n_barcodes: int = 12, n_reads: int = 5000, seed: int = 0, n_rows: int = 4, n_columns: int = 6) DemoLayout[source]¶
Populate
dstwith a self-contained map_barcodes demo:dst/ barcodes/ grna.csv # ← N gRNA barcodes, name,sequence row.csv # ← row (plate-row) barcodes column.csv # ← column barcodes demo_R1_001.fastq.gz # ← reads carrying those barcodes demo_R2_001.fastq.gz settings_map_barcodes.csv
The FASTQs sit in
dstitself, not afastq/subfolder: the pipeline’ssrcis listed flat for*.fastq.gz(spacr.io.parse_gz_files), so a subfolder means zero samples found and a run that exits having written nothing.- Parameters:
dst – destination folder.
n_barcodes – number of unique gRNA barcodes to plant.
n_reads – approximate total number of reads to emit.
seed – RNG seed for reproducibility.
n_rows – number of plate-row barcodes.
n_columns – number of plate-column barcodes.
- Returns:
DemoLayoutdescribing the emitted files.
- spacr.qt.synthetic.generate_mask_demo(dst: pathlib.Path, plate: str = 'plate1', wells: Iterable[str] = ('A01', 'A02'), fields: int = 2, channels: Iterable[int] = (0, 1, 2, 3)) DemoLayout[source]¶
Populate
dstwith a folder that runs cleanly through the Mask app. Layout:dst/ <plateID>_<wellID>_T01F<field>L01A01Z01C<chan>.tif settings_mask.csv
- Parameters:
dst – destination folder; made absolute and created if absent.
- spacr.qt.synthetic.generate_measure_demo(dst: pathlib.Path, plate: str = 'plate1', wells: Iterable[str] = ('A01', 'A02'), fields: int = 2, channels: Iterable[int] = (0, 1, 2, 3)) DemoLayout[source]¶
Measure consumes what Mask produces: a
merged/folder of.npystacks whose trailing planes are the label masks.We pre-build those stacks so a user can jump straight into Measure without a GPU. Before this, the demo wrote a
masks/folder of per-file tiffs and an emptymeasurements.db— neither of which any pipeline reads — and measure’s pre-flight rejected the folder outright with “no merged folder for measure”.Note
The organelle plane this writes is measured into nothing, and the defect is not in this module. Plane
organelle_mask_dimof every merged stack carries 64 real labels, but a measure run over this folder writescell/nucleus/pathogen/cytoplasmand noorganelletable at all (verified: 4 fields → cell 64, nucleus 64, pathogen 71, cytoplasm 64, organelle absent).All four organelle writes in
spacr.measure._measure_crop_coreare gated onsettings.get('summarize_organelles_by') is not None, andspacr.settings.get_measure_crop_settings— the defaults every measure run is canonicalised through — never sets that key.That much was already known. What was wrong was the remedy: “default it for measure the way the Mask app does” does not fix this, and cannot, for two reasons that have to be fixed together.
set_default_settings_preprocess_generate_masksdefaults it to the string'cell', andmeasure.pytests it with"organelle" in settings['summarize_organelles_by']— a substring test when the value is a str. Running this demo withsummarize_organelles_by='cell'givescell_organelle_summary(16 rows/field) and still noorganelletable. Only a value containing'organelle'writes the per-organelle table (['cell', 'organelle']→ organelle 64 rows/field, verified).A list cannot be shipped today:
spacr.settings.expected_typesdeclares'summarize_organelles_by': str, sospacr.validate.validate_settingsrejects['cell', 'organelle']with “is a list, but str is expected” — a hard pre-flight error on a demo that must load clean. The tooltip andspacr.gui_utilsboth describe it as a list, andspacr.external_masksbuilds one; only the type table disagrees.
So the demo deliberately omits the key, and the wiring needed elsewhere is: widen
expected_types['summarize_organelles_by']to(str, list, type(None)); makeget_measure_crop_settingsdefault it to['cell', 'organelle'](safe — every write is separately gated onorganelle_mask_dim is not None, which defaults toNone); and add it to themeasuresection ofspacr/qt/screens/settings_model.pyso the Measure form can hold it — without a widget,apply_settings_dictdrops it andcollect()never emits it, so a CSV key would change what a CLI run measures and nothing about a GUI run.- Parameters:
dst – destination folder; made absolute and created if absent.
- spacr.qt.synthetic.generate_synthetic_fastq(dst_dir: pathlib.Path, grnas: Sequence[str], rows: Sequence[str], columns: Sequence[str], n_reads: int = 5000, seed: int = 0, sample: str = FASTQ_SAMPLE, paired: bool = True) List[pathlib.Path][source]¶
Write a gzip-compressed synthetic FASTQ pair carrying known barcodes.
Every read is one (column, gRNA, row) triplet in the frame the shipped barcode-mapping defaults parse, so
unique_combinations.csvcomes back with the planted wells and guides in it. Reads are spread evenly over therows x columnswells, and within a well over the gRNAs with a skew, because a real screen has a handful of abundant guides and a long tail.- Parameters:
dst_dir – folder to write into.
grnas – gRNA barcode sequences (21 bases each).
rows – row barcode sequences (8 bases each).
columns – column barcode sequences (8 bases each).
n_reads – approximate total number of reads; the real total is rounded down to a whole number of reads per well.
seed – RNG seed for reproducible read pools.
sample – sample name; the files are
<sample>_R1_001.fastq.gz(and_R2_).paired – also write the R2 mate. R2 is the exact reverse complement of R1 — a perfectly overlapping pair — which is what spacr.sequencing’s paired path reduces to after it reverse-complements R2 and takes the per-base consensus.
- Returns:
the written paths, R1 first.
- Raises:
ValueError – when any barcode list is empty.
- spacr.qt.synthetic.generate_timelapse_demo(dst: pathlib.Path, plate: str = 'plate1', wells: Iterable[str] = ('A01',), fields: int = 1, times: int = 8, channels: Iterable[int] = (0, 1)) DemoLayout[source]¶
Timelapse needs multi-T frames per (well, field) so tracking has something to lock onto. Same cellvoyager naming, just with T01..T<N>, and every frame holds the same cells drifting a couple of pixels rather than a fresh random field.
Only the nucleus and cell channels are acquired, and the settings say so: the base settings used to advertise
channels=[0,1,2,3]plus apathogen_channel/organelle_channelthis dataset never had, which pre-flight rejected with two hard errors before the run could start.Note
The dataset and settings this writes clear pre-flight, but the Timelapse pipeline still cannot consume them, and neither defect is in this module. Both were reproduced on this demo and both fixes were proved by patching the two functions at run time; with the pair applied the demo completes and writes 8 merged stacks, per-channel movies, a track-overlay GIF and a 16-track
tracks/*.csv.spacr.io._rename_and_organize_image_filesnames its stack files<plate>_<well>_<field>.npywhentimelapse=True— dropping the timeID and max-projecting every timepoint of a field into one array — whilespacr.io._generate_time_listsgroups on<plate>_<well>_<field>_<time>.npyand skips anything with fewer than four underscore-separated parts. An 8-frame field becomes onestack/plate1_A01_1.npy,_generate_time_listsreturns[], no*_norm_timelapse.npzis written, no masks are generated, andpreprocess_generate_masksdies in_pivot_counts_tableonno such table: object_counts. Emitting the timeID in both branches (the non-timelapse spelling is already exactly what_generate_time_listsparses) is the fix.Past that,
spacr.object.generate_cellpose_masks_samhandsspacr.timelapse._trackpy_track_cellsa list of 2-D frames, and the tracking chain indexes it as an array:_track_by_ioudoesmasks.shape[0]and_relabel_masks_based_on_tracksdoesnp.zeros(masks.shape, …), bothAttributeError: 'list' object has no attribute 'shape'. In thetimelapse_mode='iou'path this demo asks for, the first one is swallowed by theexcept Exceptionretry loop in_facilitate_trackin_with_adaptive_removal, which then shrinks the search range 100 times and reportsFailed to track after 100 attempts— a message about displacement for a bug about a type. Coercing once at the top of_trackpy_track_cells(masks = np.asarray(masks)) clears both.
- Parameters:
dst – destination folder; made absolute and created if absent.
- spacr.qt.synthetic.main(argv: list[str] | None = None) int[source]¶
Generate one (or every) demo dataset via the
python -mCLI.- Parameters:
argv – optional argv list; defaults to
sys.argv[1:].- Returns:
process exit code (0 on success).
- spacr.qt.synthetic.save_settings_csv(path: pathlib.Path, settings: Dict[str, Any]) pathlib.Path[source]¶
Write
settingsin the two-column Key,Value format thatspacr.utils.load_settingsreads.- Parameters:
path – CSV file to write; made absolute, its parent folder is created, and an existing file is overwritten.
settings – settings to write, one
Key,Valuerow each;Noneis written as an empty value and everything else withstr().
- spacr.qt.synthetic.synthetic_read(column_barcode: str, grna: str, row_barcode: str, prefix: str = SEQ_READ_PREFIX) str[source]¶
Build one 150-base read carrying a (column, gRNA, row) triplet.
The layout is the one documented above
SEQ_TARGET; a read built here is recovered exactly by the shippedregex/target_sequence/offset_start/window_lengthdefaults.- Parameters:
column_barcode – 8-base column barcode.
grna – 21-base gRNA barcode.
row_barcode – 8-base row barcode.
prefix – stagger placed before the anchor window.
- Returns:
a 150-base read.
- Raises:
ValueError – when a barcode is the wrong length — a silently mis-sized barcode would shift every downstream field by that many bases and map to nothing, which is far harder to see than a stop.