spacr.predictions¶
Write per-object classification scores back into a spaCR measurements.db.
Two classifiers in this package score every object of a whole database:
the convolutional one —
spacr.deep_spacr.apply_model_to_tar(), driven byspacr.deep_spacr.deep_spacr(); andthe classical-ML one —
spacr.ml.generate_ml_scores().
Both produce one row per crop, and both used to leave those rows in a CSV next
to the model. A CSV is not where the rest of spaCR looks: the Annotate app, the
active-learning queue, generate_image_umap, the plate heatmaps and every
GUI table read the png_list table of measurements.db. So the scores have
to land there, on the row they belong to.
Why this is its own module¶
The merge used to live in spacr.deep_spacr, which is an odd home for
something spacr.ml needs just as much — importing deep_spacr pulls
in torch and torchvision, which the classical-ML path has no use for. The
obvious alternative, spacr.utils, already owns the database write
helpers (_append_to_measurements_db, rename_columns_in_db), but it is
also 8000 lines and imports most of the package; adding a third database
concern there would make it harder, not easier, to see that CV and ML share one
code path. This module is small, imports nothing from spaCR at module scope
except spacr.utils helpers pulled in lazily, and is what both callers
import.
The join key¶
prcfo — plate_row_column_field_object — is the canonical per-object
identity in this codebase. spacr.utils.filepaths_to_database() writes it
onto every png_list row, spacr.io._read_and_merge_data() indexes the
merged feature frame by it, and spacr.ml.generate_ml_scores() already
joins annotations onto features with it. Keying the merge on it means the CV
and the ML stage land on the same row, which is the whole point of letting
them coexist.
The previous implementation matched on os.path.basename(png_path). A
basename is not an identity:
the tar handed to the model is built with
arcname=os.path.basename(img_path)(spacr.utils.add_images_to_tar()), so a run over two source folders whose plates are both calledplate1— which is what happens whenever the plate name comes from the source folder name, seespacr.io_rename_and_organize_image_files— puts two different crops in the archive under one member name; andthe old lookup was a plain
dictassignment, so the second of those two crops silently overwrote the first and one of the two plates was scored with the other plate’s predictions.
The second half of that is the real defect, and it is not fixed by changing the
key: two crops that share a basename share a prcfo too. So the fix is to
detect the collision. A key that arrives twice with two different values is
recorded as ambiguous, written nowhere, and counted in the report. A wrong score
is worse than a missing one.
Key selection is measured, not assumed. Candidate keys are built for both sides
(prcfo, then the full png_path, then file_name — never a basename
computed behind the caller’s back when a real column exists) and the one that
actually matches the most rows wins, ties going to the earliest in
KEY_PRIORITY. The chosen key and every count are printed, because a
merge that matched 3 of 40000 rows used to look exactly like one that matched
all of them.
Atomicity¶
Python’s sqlite3 opens an implicit transaction for DML only, so an
ALTER TABLE runs in autocommit and lands immediately — the same trap
spacr.utils.rename_columns_in_db() was fixed for. The transaction here is
opened explicitly and rolled back on any error, so an interrupted merge leaves
the table exactly as it was: no half-added column, no half-scored rows.
Classes¶
What one merge did, in numbers. |
Functions¶
|
Join prediction columns onto an object frame IN MEMORY. |
|
Parse spaCR crop file names into the metadata |
|
The first of |
|
Merge |
|
Merge |
|
Write a classifier's per-object results onto the rows of |
|
Put a legacy prediction column back into the encoding it claims to be in. |
Module Contents¶
- class spacr.predictions.MergeReport[source]¶
What one merge did, in numbers.
- Parameters:
table – database table into which prediction results were merged.
key – join-key strategy selected for the merge.
columns – prediction columns requested for insertion or update.
db_rows – target-table rows considered by the merge.
result_rows – incoming prediction rows considered by the merge.
matched_rows – target rows that received at least one prediction.
matched_keys – distinct incoming identities found in the target table.
unmatched_db_rows – target rows left unchanged because no result carried their identity.
unmatched_result_rows – parseable result rows whose identity was not present in the target table.
unparsed_result_rows – result rows from which no join identity could be constructed.
ambiguous_keys – identities repeated with conflicting prediction values and therefore deliberately not written.
ambiguous_result_rows – incoming rows involved in those conflicts.
fanout_rows – additional target rows sharing a matched identity and receiving the same value, such as alternate crops of one object.
repaired – legacy prediction columns repaired before this merge, as
(table, column, rows_repaired)records.added_columns – prediction columns newly created in the target table.
Returned by
merge_prediction_results()and printed by it. Every count is here because a merge that matched three rows of forty thousand used to be indistinguishable from one that matched all of them.
- spacr.predictions.attach_predictions(objects, results, *, score_source: str = 'pred', class_source: str = 'cv_predictions', score_col: str = CV_SCORE_COLUMN, class_col: str = CV_CLASS_COLUMN, timelapse: bool = False)[source]¶
Join prediction columns onto an object frame IN MEMORY.
THE SAME JOIN AS
merge_prediction_results(), AND NOTHING WRITTEN._choose_keypicks the key by MEASURING which one lands on the most rows, so a montage reading scores out of a score CSV and a database that had the same CSV merged into it cannot disagree about which object got which number – which two separate join implementations eventually would.This supports projects whose
png_listtable has no prediction column but whose regression inputs already contain one score row per cell. The join uses those scores without changing the database.- Parameters:
objects – the per-object frame, e.g.
png_listread back.results – the score table –
path,pred,cv_predictions, asprocess_vision_resultsand the regression module’s score CSVs carry.
- Returns:
(frame, matched)– a COPY ofobjectswith the score and class columns added where they joined, and how many rows matched.matchedis 0 when nothing lined up, and the frame comes back without the columns, so a caller can refuse with a real number.
- spacr.predictions.crop_name_metadata(names, timelapse: bool = False) pandas.DataFrame[source]¶
Parse spaCR crop file names into the metadata
prcfois built from.Every consumer of a classifier’s results has to answer the same question – which object is this crop? – and there is exactly one right way to answer it:
spacr.utils._map_wells_png(), the parserspacr.utils.filepaths_to_database()used to write these very columns ontopng_list. Re-deriving them with the writer’s own parser is what guarantees the two sides cannot drift apart.It also recovers what a positional guess gets wrong.
spacr.utils.process_vision_results()takes the object id aspath.split('_')[3], which on a timelapse crop (plate_well_field_time_object) is the timepoint, not the object.- Parameters:
names – crop names or paths; only the basename is parsed.
timelapse – whether the names carry a timepoint component.
- Returns:
a DataFrame aligned to
nameswithplateID,rowID,columnID,fieldID,timeID(only whentimelapse),object_label(the bare integer id, no'o'prefix) andprcfo. A name that cannot be parsed givesNonethroughout.
- spacr.predictions.first_present(frame, names) str | None[source]¶
The first of
namesthatframeactually has, else None.
- spacr.predictions.merge_cv_predictions(df, db_path, table: str = PNG_TABLE, score_col: str = CV_SCORE_COLUMN, class_col: str = CV_CLASS_COLUMN, score_source: str = 'pred', class_source: str = 'cv_predictions', verbose: bool = True) MergeReport | None[source]¶
Merge
spacr.deep_spacr.apply_model_to_tar()results intotable.- Parameters:
df – frame from
apply_model_to_tar->process_vision_results(path,pred,cv_predictions).db_path – SQLite database to write into.
table – target table. Default
'png_list'.score_col – database column for the probability.
class_col – database column for the thresholded class.
score_source – column of
dfthe probability comes from.class_source – column of
dfthe class comes from.verbose – print the report.
- Returns:
a
MergeReport, orNoneif the database is missing.
- spacr.predictions.merge_ml_predictions(df, db_path, table: str = PNG_TABLE, score_col: str = ML_SCORE_COLUMN, class_col: str = ML_CLASS_COLUMN, verbose: bool = True) MergeReport | None[source]¶
Merge
spacr.ml.ml_analysis()results intotable.ml_analysisreturns the predicted class inpredictionsand the per-class probabilities inprediction_probability_class_<i>; the positive-class probability is taken from class 1 when the model produced one, and skipped when it did not (a one-class fit) rather than inventing a column.- Parameters:
df – the scored frame –
ml_analysisoutput[0].db_path – SQLite database to write into.
table – target table. Default
'png_list'.score_col – database column for the probability.
class_col – database column for the predicted class.
verbose – print the report.
- Returns:
a
MergeReport, orNoneif the database is missing, or if the frame carries no prediction column at all.
- spacr.predictions.merge_prediction_results(results, db_path, columns, table: str = PNG_TABLE, key: str = 'auto', timelapse: bool | None = None, verbose: bool = True) MergeReport | None[source]¶
Write a classifier’s per-object results onto the rows of
table.Shared by Classify (CV) and Classify (ML) — see the module docstring for why the key is
prcfo, why an ambiguous key is refused rather than guessed at, and why the whole thing is one transaction.- Parameters:
results – DataFrame of per-object results. Must carry the source columns named in
columns, plus something to key on: aprcfocolumn, an index namedprcfo, or apath/png_path/file_namecolumn holding spaCR crop names.db_path – SQLite or DuckDB file, Parquet store directory, or PostgreSQL connection string. A missing local store is reported and skipped.
columns – mapping of database column name to
(results column, 'REAL' | 'INTEGER').table – target table. Default
'png_list'.key –
'auto'(measure every candidate and use the best) or one ofKEY_PRIORITYto force it.timelapse – whether crop names carry a timepoint.
Nonedetects it from the presence of a time column on the table.verbose – print the report.
- Returns:
a
MergeReport, orNoneif a local store is missing.- Raises:
KeyError – if
resultslacks one of the source columns.ValueError – if a Parquet table is missing or prediction output would replace a join-key column in an alternate store.
- spacr.predictions.migrate_prediction_columns(db_path, table: str = PNG_TABLE, verbose: bool = True) List[Tuple[str, str, int]][source]¶
Put a legacy prediction column back into the encoding it claims to be in.
spacr.utils.add_column_to_database()— the ML stage’s old write path — replaced every0with a2on the way into the database, because the Annotate app labels classes 1 and 2. Nothing else did that, sopng_list.predictionssaid2whereresults.csvfrom the very same run said0: one number, two meanings, depending on which file you opened. This repairs it, so a database written before the change reads correctly with no manual action — the same repair-on-read contractspacr.utils.rename_columns_in_db()has, and the same properties:Idempotent. After the repair the column holds 0s and 1s, which is not the encoding this looks for, so a second pass does nothing.
Never destructive. The substitution is only reversed when the column holds nothing but 1s and 2s. A three-class model’s genuine class 2 is indistinguishable from a mangled 0, and guessing would be worse than leaving it exactly as it is.
All or nothing. SQLite, DuckDB and PostgreSQL repair in a transaction; Parquet publishes a complete replacement part snapshot. A half-repaired column would be a column in two encodings at once.
- Parameters:
db_path – SQLite or DuckDB file, Parquet store directory, or PostgreSQL connection string. A missing local store is a no-op.
table – table to migrate. Default
'png_list'.verbose – print each repair.
- Returns:
list of
(table, column, rows_repaired).
Nested helpers¶
- crop_name_metadata.bare_object_label(value)¶
Normalize a parsed object label and remove one leading
o.spacr/predictions.py:455
- crop_name_metadata.convert(value)¶
Parse and cache one crop basename, or return all-missing metadata.
spacr/predictions.py:441